Prompt word expansion method and device
By generating multiple question sentences and response keywords associated with the initial prompt word, the problem of unreasonable filling of prompt word in the prior art is solved, and the effective expansion of prompt word and the rationality of generated content is achieved.
Patent Information
- Application Number
- CN202510299873.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-01
AI Technical Summary
When generating content, the prior art usually uses default prompt words to fill in short prompt words provided by the user, resulting in the generated content being unreasonable enough and cannot be effectively expanded.
By generating multiple question statements associated with the initial prompt words to be expanded, a response keyword corresponding to the question statement is generated, and an extended prompt word is generated based on the response keyword and the initial prompt word.
The effective extension of prompt words is realized, and the generated extended prompt words are more reasonable, improving the accuracy and rationality of the generated content.
Smart Images

Figure CN120235149A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to the field of artificial intelligence technology, and particularly to a method and device for prompt expansion. Background Art
[0002] With the prosperous development of AI technology, various applications for generating content through AI have emerged. For example, applications for generating images based on users' voice or text input; applications for generating images based on text + images; applications for generating videos based on text; etc.; applications for generating music based on text.
[0003] When generating the content desired by users, these applications all require a detailed prompt. When the prompt provided by the user is short, default prompts are usually used to fill in the prompt provided by the user. Summary of the Invention
[0004] Embodiments of this application provide a method, device, equipment, and storage medium for prompt expansion.
[0005] According to a first aspect, embodiments of this application provide a method for prompt expansion. The method includes: generating a plurality of question statements associated with an initial prompt to be expanded; generating response keywords corresponding to the plurality of question statements; and generating an expanded prompt based on the response keywords and the initial prompt.
[0006] According to a second aspect, embodiments of this application provide a prompt expansion device, including: a generation module, a response module, and an expansion module. Among them, the generation module is configured to generate a plurality of question statements associated with an initial prompt to be expanded; the response module is configured to generate response keywords corresponding to the plurality of question statements; and the expansion module is configured to generate an expanded prompt based on the response keywords and the initial prompt.
[0007] According to a third aspect, embodiments of this application provide an electronic device, which includes one or more processors; a storage device on which one or more programs are stored. When the one or more programs are executed by the one or more processors, the one or more processors implement the prompt expansion method according to any one of the embodiments of the first aspect.
[0008] According to a fourth aspect, embodiments of this application provide a computer-readable medium on which a computer program is stored. When the program is executed by a processor, it implements the prompt expansion method according to any one of the embodiments of the first aspect.
[0009] This application generates multiple question statements associated with the initial prompt word based on the initial prompt word to be expanded; generates response keywords corresponding to the multiple question statements; and generates an expanded prompt word based on the response keywords and the initial prompt word, avoiding directly using the default prompt word to fill in the shorter prompt word provided by the user, making the generated expanded prompt word more reasonable and achieving the effective expansion of the prompt word.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is an exemplary system architecture diagram in which this application can be applied;
[0012] Figure 2 is a flowchart of an embodiment of the prompt word expansion method according to this application;
[0013] Figure 3 is a schematic diagram of an application scenario of the prompt word expansion method according to this application;
[0014] Figure 4 is a flowchart of another embodiment of the prompt word expansion method according to this application;
[0015] Figure 5 is a schematic diagram of an embodiment of the prompt word expansion device according to this application;
[0016] Figure 6 is a schematic diagram of the structure of a computer system of a server suitable for implementing the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following describes exemplary embodiments of this application with reference to the accompanying drawings. Various details of the embodiments of this application are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0018] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will describe this application in detail with reference to the drawings and in combination with the embodiments.
[0019] Figure 1 Illustrates an exemplary system architecture 100 of an embodiment in which the prompt word expansion method of this application can be applied.
[0020] As Figure 1 shown in Figure 1 , the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105, and between the terminal devices. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0021] Users can use the terminal devices 101, 102, 103 to interact with other terminal devices or the server 105 through the network 104 to receive or send messages, etc. Client application software, such as video playback application software, communication application software, etc., may be installed on the terminal devices 101, 102, 103.
[0022] The terminal devices 101, 102, 103 can be hardware or software. When the terminal devices 101, 102, 103 are hardware, they can be various electronic devices, including but not limited to mobile phones, laptop computers, AR glasses, VR headsets, etc. When the terminal devices 101, 102, 103 are software, they can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules, or as a single software or software module. No specific limitation is made here.
[0023] The server 105 can be a server that provides various services. For example, based on an initial prompt word to be extended, generate multiple question statements associated with the initial prompt word; generate response keywords corresponding to the multiple question statements; and generate an extended prompt word according to the response keywords and the initial prompt word.
[0024] It should be noted that the server 105 can be hardware or software. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules (such as those used to provide prompt word extension services), or as a single software or software module. No specific limitation is made here.
[0025] It should be pointed out that the prompt word extension method provided by the embodiments of the present disclosure can be executed by the server 105, or by the terminal devices 101, 102, 103, or by the server 105 and the terminal devices 101, 102, 103 in cooperation with each other. Correspondingly, each part (such as each unit, subunit, module, submodule) included in the prompt word extension device can be all set in the server 105, or all set in the terminal devices 101, 102, 103, or can be respectively set in the server 105 and the terminal devices 101, 102, 103.
[0026] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in [[ ]] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0027] Figure 2 FIG. 200 is a schematic flowchart of an embodiment of a prompting word expansion method that can be applied to the present application. The prompting word expansion method includes the following steps:
[0028] Step 201, based on an initial prompting word to be expanded, generate a plurality of question statements associated with the initial prompting word.
[0029] In this embodiment, the execution entity (e.g., Figure 1 the server 105 or the terminal devices 101, 102, 103 in [[ ]]) may first perform semantic analysis on the prompting word (i.e., prompt) input by the user to determine the initial prompting word to be expanded, and then generate a plurality of question statements associated with the initial prompting word according to the initial prompting word to be expanded.
[0030] Among them, the prompting word is the prompting information input to the generation model, which is used to guide the model to generate specific types of content. The prompting word can be a single-modal prompting word or a multi-modal prompting word. For example, the prompting word can include at least one of information such as text, image, music, video, etc.
[0031] For a multi-modal prompting word, the execution entity may separately perform semantic analysis on each modal information in the multi-modal prompting word to generate descriptive information, perform semantic fusion and distance matching on the descriptive information, and perform semantic analysis on the results of the semantic fusion and distance matching to obtain the initial prompting word to be expanded.
[0032] Specifically, if the prompting word input by the user is a multi-modal prompting word, such as including an image, music, and video, for the image, the execution entity may identify the elements in the image and perform semantic analysis based on the elements in the image (e.g., determining element positions, relationships, environments, actions, etc.) to obtain image descriptive information; for the music, the execution entity may extract the audio features and background information of the music (e.g., album, singer, time, etc.) and perform semantic analysis based on the audio features and background information to obtain music descriptive information. For the video, the execution entity may extract the visual content, music content, text content, and background information of the video (e.g., video author, video classification, etc.) and perform semantic analysis based on the visual content, music content, text content, and background information respectively to obtain video descriptive information.
[0033] Further, the execution entity may perform semantic fusion and distance matching on the image description information, music description information, and video description information, where the semantic fusion is used to improve the semantic fusion degree of the associated data, and the distance matching is used to filter out semantic information that does not meet the set threshold or rules.
[0034] Further, the execution entity may perform semantic analysis on the description information after semantic fusion and distance matching to extract keywords, and determine the extracted keywords as the initial prompt words to be expanded.
[0035] Here, the ways for the execution entity to generate multiple question statements associated with the initial prompt words to be expanded may include multiple types. For example, according to the initial prompt words and the preset mapping relationship table between the initial prompt words and multiple associated question statements, determine multiple question statements associated with the initial prompt words; input the initial prompt words into the preset second language model to generate multiple question statements associated with the initial prompt words, etc.
[0036] Step 202, generate response keywords corresponding to the multiple question statements.
[0037] In this embodiment, the execution entity may determine the response keywords corresponding to each question statement according to the multiple question statements and the preset mapping relationship table between the question statements and the response keywords; or input the multiple question statements into the preset first language model to generate the response keywords corresponding to the multiple question statements. This application does not make any limitations on this.
[0038] Specifically, the multiple question statements are "a. What color is this beautiful lady’s hair? b. What color are this beautiful lady’s eyes?", and the response keywords corresponding to the multiple question statements are "a. black hair b. black eyes".
[0039] In some alternative ways, generating the response keywords corresponding to the multiple question statements includes: generating the response keywords corresponding to the multiple question statements based on the knowledge graph and the preset first language model.
[0040] In this implementation manner, there are various ways for the execution entity to generate response keywords corresponding to multiple question statements based on the knowledge graph and a preset first language model. For example, based on multiple question statements, the retriever (such as RAG (Retrieval-Augmented Generation) technology, Scopus, etc.) is used to find relevant entities and relationships in the knowledge graph, and the found entities and relationships are input into the preset first language model to generate response keywords corresponding to the multiple question statements; the first language model is fine-tuned based on the knowledge graph to obtain a fine-tuned first language model, and the multiple question statements are input into the fine-tuned first language model to generate response keywords corresponding to the multiple question statements; the multiple question statements are input into the preset first language model to obtain answers, and the answers are verified and corrected based on the knowledge graph to obtain response keywords corresponding to the multiple question statements, etc.
[0041] Here, the knowledge graph is a technical method that uses a graph model to describe knowledge and model the association relationships between all things in the world. It consists of three key components: entities, attributes, and relationships. Entities are the main objects or concepts in the graph, which can be people, places, things, or abstract concepts; attributes are the characteristics or properties of these entities; relationships are the connections or associations between different entities, defining the relationships between entities.
[0042] The knowledge graph can be constructed based on the user information corresponding to the initial prompt (i.e., the user information of the user who submitted the prompt).
[0043] Among them, the user information can include the environmental information where the user is located (such as user location, time, electronic device used, scene, etc.), the attribute information of the user (such as gender, age, address, preference, etc.), the operation behavior information of the user (such as browsing, forwarding, collecting, liking, etc.), etc.
[0044] Here, the first language model can be any deep learning model. For example, BERT (Bidirectional Encoder Representations from Transformers), LLM (Large Language Model), etc. This application does not make any limitations in this regard.
[0045] In addition, to ensure the accuracy of the generated response keywords, the execution entity can update the knowledge graph before generating response keywords each time. It can update the knowledge graph when it detects that the user information has changed or when the change meets the specified conditions, or it can also update the knowledge graph when it detects that the evaluation information of the user on the generation result obtained based on the extended prompt does not meet the preset conditions.
[0046] Specifically, the user information for constructing the knowledge graph includes: the user owns mobile phone A, the user's preferred music genre: classic, the user's preferred video genre: realism, the user's favorite female celebrity: YM, the user's favorite male celebrity: SJK, the user's residential address: CC, the user travels to S Island, the user's family member: husband, the user's pet: cat, the user's preferred sport: Yoga, the user's physical condition: health, the user's preferred color: Gray, the user's favorite game: null, the user's location: China, the time: afternoon, the user's surrounding environment: Quiet, User gender: female, the most frequently liked app: LRB. The aspect ratio of mobile phone A is 9:16.
[0047] If an update message is detected, such as mobile phone A connecting to S TV, and S TV: supports 4K, supports.mp4, aspect ratio 16:9, update the knowledge graph.
[0048] This implementation method generates response keywords corresponding to multiple question statements based on the knowledge graph and a preset first language model, which helps to generate personalized response keywords, thereby improving the accuracy of the generated extended prompt words.
[0049] In some optional ways, the method further includes: determining a generation result based on the extended prompt word and a preset generation model; generating evaluation information for the generation result based on the user's operation behavior regarding the generation result; and updating the knowledge graph in response to determining that the evaluation information does not meet the preset conditions.
[0050] In this implementation method, after obtaining the extended prompt word, the execution entity can directly input the extended prompt word into the preset generation model to obtain the generation result.
[0051] Among them, the generation model can be any deep learning model that can generate different modal information based on the extended prompt word. For example, the CLIP (Contrastive Language-Image Pre-training) model, the Multimodal GPT (Multimodal Generative Pre-trained Transformer) model, etc.
[0052] Furthermore, the execution entity can generate evaluation information for the generation result according to the user's operation behavior regarding the generation result and a preset evaluation rule.
[0053] Here, the evaluation information is usually used to characterize the user's satisfaction with the generated result.
[0054] Among them, the operation behaviors can include various ones. For example, whether to share the generated result, the number of times of sharing the generated result, the browsing time of the generated result, whether to edit the generated result, whether to collect the generated result, whether to use the generated result, etc.
[0055] The preset evaluation rules can include various ones. For example, determine the user's score for the generated result according to the score corresponding to the user's operation behavior; determine the user's score for the generated result (i.e., the evaluation information) according to the score corresponding to the user's operation behavior and the weight.
[0056] If the evaluation information does not meet the preset conditions, for example, the score is less than or equal to the preset score threshold, then update the knowledge graph.
[0057] Among them, the preset score threshold can be set according to experience and actual needs.
[0058] Specifically, the generated result is a picture. The user's operation behaviors for the generated result include: sharing the picture 5 times, browsing the picture for 10 minutes, editing the picture 1 time, collecting the picture, setting the picture as the wallpaper, and applying the picture to other connected devices. The scores corresponding to each operation behavior are 4 points, 12 points, 1 point, 1 point, 6 points, and 14 points respectively. The weights corresponding to each operation behavior are 0.3, 0.2, -0.1, 0.3, 0.1, and 0.2 respectively. The user's score for the picture is 4 * 0.3 + 12 * 0.2 + 1 * (-0.1) + 1 * 0.3 + 6 * 0.1 + 14 * 0.2 = 7.2. If the preset score threshold is 5 points and 7.5 points is greater than 5 points, then there is no need to update the knowledge graph.
[0059] This implementation method determines the generated result based on the extended prompt and the preset generation model; generates the evaluation information for the generated result based on the user's operation behavior for the generated result; in response to determining that the evaluation information does not meet the preset conditions, updates the knowledge graph, improving the timeliness and effectiveness of updating the knowledge graph.
[0060] Step 203, generate an extended prompt according to the response keyword and the initial prompt.
[0061] In this embodiment, after obtaining the response keyword, the execution subject can insert the response keyword into the initial prompt according to the semantic association between the response keyword and the initial prompt to obtain the extended prompt, or directly splice the response keyword after the initial prompt to obtain the extended prompt. This application does not make any limitations in this regard.
[0062] Specifically, the initial prompt is "Generate a beautiful lady image". The multiple question statements generated based on the initial prompt are "a. What color is this beautiful lady’s hair? b. What color are this beautiful lady’s eyes? c. What is the face shape of this beautiful lady? d. What style of images do user prefer? e. Does the device support 4K? f. What is the optimal aspect ratio for displaying image?". The corresponding response keywords for each question statement are "a. black hair b. black eyes c. oval face d. classical art e. yes f. aspect:16:9". Based on the response keywords and the initial prompt, the execution entity generates an extended prompt: "Generate a beautiful lady image, black hair, black eyes, oval face, classical art style, 4K, aspect:16:9".
[0063] Here, the extended prompt may include a content prompt (e.g., the theme, key elements of the generation result, etc.) and a parameter prompt (e.g., the overall ratio, style, resolution of the generation result, etc.). The execution entity can input the content prompt into the generation model to obtain an initial result, and then process the initial result according to the parameter prompt to obtain the generation result.
[0064] In some alternative ways, the method further includes: in response to determining that the extended prompt does not meet the preset extension conditions, using the extended prompt as a new initial prompt, and continuing to execute to generate multiple question statements associated with the initial prompt based on the initial prompt to be extended.
[0065] In this implementation, after obtaining the extended prompt, the execution entity can determine whether the extended prompt meets the preset extension conditions. If not, it uses the extended prompt as a new initial prompt and continues to execute to generate multiple question statements associated with the initial prompt based on the initial prompt to be extended, as well as generate response keywords corresponding to the multiple question statements; generate an extended prompt based on the response keywords and the initial prompt, that is, continue to execute the above steps 201, 202, and 203.
[0066] Among them, the preset expansion conditions can be set according to actual needs. For example, whether the number of times of generating expansion prompts reaches the preset number threshold, whether the number of keywords included in the expansion prompts reaches the preset number threshold, whether the accuracy of the expansion prompts is greater than or equal to the preset accuracy threshold, etc.
[0067] Specifically, the preset expansion condition is that the number of times of generating expansion prompts reaches the preset number threshold. The execution entity counts the number of times of generating expansion prompts and compares the statistical result with the preset number threshold (i.e., 5 times). If the statistical result is less than the number threshold, the execution entity can use the expansion prompt as a new initial prompt and continue to execute the above steps 201, 202, and 203 until the statistical result reaches the number threshold.
[0068] This implementation method, by responding to the determination that the expansion prompt does not meet the preset expansion conditions, uses the expansion prompt as a new initial prompt and continues to execute the generation of multiple question statements associated with the initial prompt based on the initial prompt to be expanded, effectively improving the accuracy of generating expansion prompts.
[0069] Continue to refer to Figure 3 , Figure 3 which is a schematic diagram of an application scenario of the prompt expansion method according to this embodiment.
[0070] In Figure 3In the application scenario, user 301 is using mobile phone 302, and the mobile phone is connected to TV 303. User 301 inputs a prompt word through the voice assistant, such as the voice 304 of "Generate a picture of girls traveling on a train" and a picture 305 of four girls chatting together. The execution entity can perform semantic analysis based on picture 304 and voice 305 to obtain the initial prompt word 306 to be expanded. At the same time, the execution entity can construct a knowledge graph according to the user information, where the user information can include the electronic device (mobile phone) the user is using, the device connected to the electronic device (TV), and the display scenario (the TV presents pictures). The execution entity can input the initial prompt word 306 to be expanded into a preset second language model 307 to generate multiple question statements associated with the initial prompt word. Based on the first language model 308 and the knowledge graph, generate response keywords corresponding to the multiple question statements, and generate an expanded prompt word 309 according to the response keywords and the initial prompt word. Among them, the expanded prompt word 309 can include a content prompt word 310 (such as Four girls traveling on train, chats and laugh, girls are black hair, brown eyes) and a parameter prompt word 311 (such as aspect: 16:9; file format.jpeg; 4K; realism).
[0071] Here, if the expanded prompt word does not meet the preset expansion conditions, the expanded prompt word is used as the new initial prompt word (i.e., update the initial prompt word), and continue to execute the above steps of generating the expanded prompt word until the generated expanded prompt word meets the preset expansion conditions.
[0072] Further, the execution entity can input the content prompt word 310 into a generation model 311 (such as a large generation model) to obtain an initial result, and process the initial result according to the parameter prompt word 312 to obtain a generation result, that is, generate a picture.
[0073] The prompt word expansion method provided by the embodiments of the present disclosure realizes the effective expansion of the prompt word by generating multiple question statements associated with the initial prompt word based on the initial prompt word to be expanded; generating response keywords corresponding to the multiple question statements; and generating an expanded prompt word according to the response keywords and the initial prompt word.
[0074] Further reference Figure 4 , which shows the flow 400 of another embodiment of the prompt word expansion method. In this embodiment, the flow 400 of the prompt word expansion method may include the following steps:
[0075] Step 401, input the initial prompt word to be expanded into a preset second language model to generate multiple question statements associated with the initial prompt word.
[0076] In this embodiment, the execution entity may input the initial prompt word to be expanded into a preset second language model to generate multiple question statements associated with the initial prompt word.
[0077] Among them, the second language model can be any deep learning model. For example, BERT (Bidirectional Encoder Representations from Transformers), LLM (Large Language Model), etc. This application does not make any limitations in this regard.
[0078] Here, the second language model can be trained based on the initial prompt word samples annotated with multiple question statements.
[0079] In some alternative ways, the second language model can be trained as follows: input the sample prompt word into the preset second language model to generate multiple actual question statements associated with the sample prompt word; determine the actual response keywords corresponding to the multiple actual question statements; construct a first loss function based on the actual expansion words and the expected expansion words corresponding to the sample prompt word; update the parameters of the second language model based on the first loss function.
[0080] In this implementation, the execution entity may collect non-text modal data (such as pictures, videos, etc.) and input the non-text modal data into a preset vision-language model to obtain the corresponding text data. Further, sample data is obtained by extracting prompt words of different levels (such as basic level, intermediate level, advanced level, context level, intention level, etc.) based on the text data.
[0081] Among them, the sample data may include sample prompt words and expected expansion words corresponding to the sample prompt words, that is, expected expanded prompt words.
[0082] Here, before extracting the prompt words from the text data, the execution entity may screen the generated text data based on the matching degree between the non-text modal data and the corresponding text data, and generate sample data based on the screened text data.
[0083] Among them, the matching degree between the non-text modal data and the corresponding text data can be determined based on the score given by the technical personnel on the matching degree of the two.
[0084] Specifically, the data in non-text modality is image data. The executing entity can collect an image data set, input the image data into a preset vision-language model to generate corresponding captions, that is, text data, score the matching degree between the image data and the corresponding text data, determine the text data with a score greater than or equal to the preset score threshold as the filtered text data, and extract prompt words of different levels based on the filtered text data to obtain sample data.
[0085] Furthermore, the executing entity can input multiple actual problem statements into a preset first language model to obtain actual response keywords corresponding to the multiple actual problem statements, determine actual expansion words according to the actual response keywords and sample prompt words, construct a first loss function according to the actual expansion words and expected expansion words, and update the parameters of the second language model based on the first loss function.
[0086] This implementation method realizes the training of the second language model in the accurate direction of the generated extended prompt words, and improves the accuracy and reliability of the trained second language model.
[0087] In some optional ways, constructing a first loss function based on the actual expansion words and the expected expansion words corresponding to the sample prompt words includes: extracting target expansion words from the actual expansion words, and constructing a first loss function based on the target expansion words and the expected expansion words corresponding to the sample prompt words.
[0088] In this implementation method, after obtaining the actual expansion words, the executing entity can extract target expansion words from the actual expansion words. The target expansion words can include nouns, verbs, and adjectives, and construct a first loss function according to the extracted target expansion words and the expected expansion words corresponding to the sample prompt words.
[0089] Specifically, the actual expansion word is “A little girl is posing for a photo. She has long blonde hair and is wearing a light blue dress.” Among them, the nouns, verbs, and adjectives included in the actual expansion word are girl, dress, hair, posing, wearing, light blue, long blonde respectively. Therefore, the target expansion words are “girl, dress, hair, posing, wearing, light blue, long blonde”.
[0090] This implementation extracts the target extended words from the actual extended words, constructs a first loss function based on the target extended words and the expected extended words corresponding to the sample prompt words, that is, constructs a loss function based on the prompt words of multiple parts of speech, and trains the second language model based on the prompt words of multiple parts of speech, improving the diversity and accuracy of the prompt words generated by the second language model obtained through training.
[0091] In some alternative ways, the method further includes: in response to determining that during the parameter update process, the second language model meets the specified conditions, inputting the actual extended words into a preset generation model to obtain an actual generation result; constructing a second loss function based on the expected generation result corresponding to the sample prompt words and the actual generation result; and continuing to update the parameters of the second language model based on the first loss function and the second loss function.
[0092] In this implementation, the sample data further includes the expected generation result corresponding to the sample prompt words. The execution entity updates the parameters of the second language model based on the first loss function, and determines whether the second language model during the update process meets the specified conditions. For example, the number of updates reaches a preset value, the update duration reaches a preset duration, the accuracy rate of the question sentences generated by the second language model reaches a preset value, etc. If so, the actual extended words are input into the preset generation model to obtain an actual generation result; a second loss function is constructed based on the expected generation result corresponding to the sample prompt words and the actual generation result.
[0093] Furthermore, the execution entity can continue to update the parameters of the second language model according to the first loss function, the first weight corresponding to the first loss function, the second loss function, and the second weight corresponding to the second loss function until the training is completed.
[0094] Among them, the value of the second weight is usually greater than the first weight.
[0095] Specifically, the execution entity can collect an image data set, input the image data into a preset vision-language model to generate corresponding captions, that is, text data, determine the sample prompt words and expected extended words included in the sample data based on the text data, and determine the image data for generating the text data as the expected generation result included in the sample data.
[0096] It should be noted that when the generation result is image data, the second loss function can be determined according to the clip loss (i.e., Clip Loss), the weight of the clip loss, the aesthetic loss (i.e., Aesthetic Loss), and the weight of the aesthetic loss.
[0097] Among them, the clipping loss is a loss function based on content recognition and understanding, aiming to narrow the difference between the generation result of the generation model and the expected target; the aesthetic loss is a loss function that improves the visual attractiveness of the generated content, aiming to ensure that the generated content is not only technically correct but also visually pleasing.
[0098] This implementation method further trains the second language model by combining the generation model when the second language model meets the specified conditions, realizing the training of the second language model in the direction of accurate generated extended words and reasonable generation results, and improving the accuracy and reliability of the trained second language model.
[0099] Step 402: Generate response keywords corresponding to multiple question statements.
[0100] In this embodiment, for the implementation details and technical effects of step 402, reference can be made to the description of step 202, which will not be elaborated here.
[0101] Step 403: Generate an extended prompt word based on the response keyword and the initial prompt word.
[0102] In this embodiment, for the implementation details and technical effects of step 403, reference can be made to the description of step 203, which will not be elaborated here.
[0103] The above embodiments of the present application, compared with Figure 2 the corresponding embodiments, in this embodiment, the process 400 of the prompt word extension method reflects that the initial prompt word to be extended is input into a preset second language model to generate multiple question statements associated with the initial prompt word; an extended prompt word is generated according to the response keyword and the initial prompt word, further improving the accuracy of the generated extended prompt word.
[0104] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, the present application provides an embodiment of a prompt word extension device. This device embodiment corresponds to Figure 1 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0105] As shown in Figure 5 , the prompt word extension device 500 in this embodiment includes: a generation module 501, a response module 502, and an extension module 503.
[0106] Among them, the generation module 501 can be configured to generate multiple question statements associated with the initial prompt word based on the initial prompt word to be extended.
[0107] The response module 502 can be configured to generate response keywords corresponding to multiple question statements.
[0108] An expansion module 503, which can be configured to generate an expansion prompt word according to a response keyword and the initial prompt word.
[0109] In some optional ways of this embodiment, the generation module is further configured to input the initial prompt word to be expanded into a preset second language model to generate a plurality of question statements associated with the initial prompt word.
[0110] In some optional ways of this embodiment, the second language model is trained as follows: inputting a sample prompt word into a preset second language model to generate a plurality of actual question statements associated with the sample prompt word; determining actual response keywords corresponding to the plurality of actual question statements; constructing a first loss function based on the actual expansion word and the expected expansion word corresponding to the sample prompt word, where the actual expansion word is determined according to the actual response keyword and the sample prompt word; and updating the parameters of the second language model based on the first loss function.
[0111] In some optional ways of this embodiment, in response to determining that the second language model meets a specified condition during the parameter update process, inputting the actual expansion word into a preset generation model to obtain an actual generation result; constructing a second loss function based on the expected generation result and the actual generation result corresponding to the sample prompt word; and continuing to update the parameters of the second language model based on the first loss function and the second loss function.
[0112] In some optional ways of this embodiment, constructing a first loss function based on the actual expansion word and the expected expansion word corresponding to the sample prompt word includes: extracting a target expansion word from the actual expansion word, and constructing a first loss function based on the target expansion word and the expected expansion word corresponding to the sample prompt word.
[0113] In some optional ways of this embodiment, the response module is further configured to: generate response keywords corresponding to the plurality of question statements based on a knowledge graph and a preset first language model.
[0114] In some optional ways of this embodiment, the device further includes an update module, and the update module is configured to determine a generation result based on the expansion prompt word and a preset generation model; generate evaluation information for the generation result based on the user's operation behavior for the generation result; and update the knowledge graph in response to determining that the evaluation information does not meet a preset condition.
[0115] In some optional ways of this embodiment, the device further includes an execution module, and the execution module is configured to, in response to determining that the expansion prompt word does not meet a preset expansion condition, use the expansion prompt word as a new initial prompt word and continue to execute generating a plurality of question statements associated with the initial prompt word based on the initial prompt word to be expanded.
[0116] It should be noted that in the technical solution of the present disclosure, in aspects such as the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information, it complies with the provisions of relevant laws and regulations, is used for legal purposes, and does not violate public order and good customs. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and to safeguard the security of user personal information, network security, and national security.
[0117] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.
[0118] As Figure 6 shown, it is a block diagram of an electronic device for a prompt word expansion method according to an embodiment of the present application.
[0119] 600 is a block diagram of an electronic device for a prompt word expansion method according to an embodiment of the present application. As Figure 6 shown, the electronic device includes: one or more processors 601, a memory 602, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component is interconnected using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 6 In
[0120]
[0121] Figure 5 The memory 602 is the non-transitory computer-readable storage medium provided by the present application. Among them, the memory stores instructions executable by at least one processor, so that the at least one processor executes the prompt word expansion method provided by the present application. The non-transitory computer-readable storage medium of the present application stores computer instructions, and these computer instructions are used to cause a computer to execute the prompt word expansion method provided by the present application. Figure 5The shown generation module 501, response module 502, and extension module 503). The processor 601 executes various functional applications and data processing of the server by running non-transitory software programs, instructions, and modules stored in the memory 602, that is, implements the prompt word extension method in the above method embodiments.
[0122] The memory 602 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the use of the electronic device for prompt word extension, etc. In addition, the memory 602 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 602 optionally includes a memory remotely set relative to the processor 601, and these remote memories can be connected to the electronic device for prompt word extension through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0123] The electronic device for the prompt word extension method may further include: an input device 603 and an output device 604. The processor 601, the memory 602, the input device 603, and the output device 804 can be connected through a bus or other means. Figure 6 Taking the connection through the bus as an example.
[0124] The input device 603 can receive input digital or character information, and generate key signal inputs related to user settings and function controls of the electronic device for quality monitoring of the live video stream, such as input devices like touchscreens, keypads, mice, trackpads, touchpads, joysticks, one or more mouse buttons, trackballs, joysticks, etc. The output device 604 may include a display device, an auxiliary lighting device (e.g., LED), and a tactile feedback device (e.g., vibration motor), etc. The display device may include but is not limited to a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touchscreen.
[0125] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuit systems, application specific ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0126] These computational programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0127] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0128] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0129] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client - server relationship is created by computer programs running on respective computers and having a client - server relationship with each other.
[0130] According to the technical solution of the embodiment of the present application, the effective expansion of the prompt word is achieved.
[0131] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present application can be achieved, and no limitation is made herein.
[0132] The above - mentioned specific implementation manners do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A prompt word expansion method, the method comprising: Based on the initial prompt word to be expanded, generating a plurality of question sentences associated with the initial prompt word; generating answer keywords corresponding to the plurality of question statements; An extended prompt word is generated according to the response keyword and the initial prompt word.
2. The method according to claim 1, wherein: The generating of answer keywords corresponding to the plurality of question statements includes: Based on a knowledge graph and a preset first language model, answer keywords corresponding to the multiple question sentences are generated, and the knowledge graph is constructed based on user information corresponding to the initial prompt word.
3. The method according to claim 2, further comprising: Determining a generation result based on the extended prompt word and a preset generation model; generating evaluation information for the generated result based on the user's operation behavior for the generated result; In response to determining that the evaluation information does not meet a preset condition, the knowledge graph is updated.
4. The method according to claim 1, wherein: The step of generating a plurality of question statements associated with the initial prompt word based on the initial prompt word to be expanded includes: The initial prompt word to be expanded is input into a preset second language model to generate multiple question sentences associated with the initial prompt word, wherein the second language model is trained in the following manner: Inputting the sample prompt words into a preset second language model to generate a plurality of actual question sentences associated with the sample prompt words; Determining actual answer keywords corresponding to the plurality of actual question statements; constructing a first loss function based on an actual expansion word and an expected expansion word corresponding to the sample prompt word, wherein the actual expansion word is determined according to the actual answer keyword and the sample prompt word; The parameters of the second language model are updated based on the first loss function.
5. The method according to claim 4, further comprising: In response to determining that the second language model meets the specified condition during the parameter update process, inputting the actual expanded word into a preset generation model to obtain an actual generation result; Constructing a second loss function based on the expected generation result corresponding to the sample prompt word and the actual generation result; Based on the first loss function and the second loss function, continue to update the parameters of the second language model.
6. The method according to claim 4, wherein: The constructing a first loss function based on the actual expansion word and the expected expansion word corresponding to the sample prompt word includes: Target expansion words are extracted from actual expansion words, and a first loss function is constructed based on the target expansion words and expected expansion words corresponding to the sample prompt words, wherein the target expansion words include nouns, verbs, and adjectives.
7. The method according to any one of claims 1 to 6, further comprising: In response to determining that the extended prompt word does not meet the preset extension condition, the extended prompt word is used as a new initial prompt word, and the initial prompt word to be expanded is continued to be executed to generate multiple question sentences associated with the initial prompt word.
8. A prompt word expansion device, the device comprising: A generating module, configured to generate a plurality of question sentences associated with the initial prompt word based on the initial prompt word to be expanded; A response module, configured to generate response keywords corresponding to the plurality of question statements; The expansion module is configured to generate an extended prompt word according to the response keyword and the initial prompt word.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.