Sample generation method and device based on large model and intelligent agent

By generating training samples through targeted modification of abnormal query results, the problem of insufficient accuracy of existing retrieval models in abnormal query results is solved, and the recognition ability of the retrieval model and user experience are improved.

CN120705584AActive Publication Date: 2025-09-26BAIDU COM TIMES TECH (BEIJING) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510866603.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-26
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

When existing retrieval models use random sampling or sampling based on sampling rules to increase training samples, they are difficult to meet users' performance requirements for retrieval models. In particular, they are unable to accurately target specific abnormal samples, resulting in mismatches between query statements and query results, affecting user retrieval experience and model quality.

Method used

The initial query statement in the abnormal query results and the target query results obtained by the retrieval model query are modified through the large model to generate targeted training samples, and the retrieval model is optimized to improve its ability to recognize abnormal query results.

Benefits of technology

The retrieval model's ability to identify abnormal query results has been improved, thereby enhancing the accuracy of query results and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705584A_ABST
    Figure CN120705584A_ABST
Patent Text Reader

Abstract

The invention provides a sample generation method and device based on a large model and an intelligent agent, and relates to the technical field of artificial intelligence, in particular to the technical field of large language models, AI intelligent assistants and machine learning. According to the specific implementation scheme, at least one problem selected by an object on a problem analysis interface is received; the at least one problem is obtained by performing reason analysis on an abnormal query result output by the retrieval model; and based on the at least one problem, modifying the target query statement and the target query result by using the large model to generate a target sample, so as to optimize the retrieval model for the at least one problem based on the target sample, wherein the target query statement is associated with an initial query statement in the abnormal query result; the target query result is obtained by querying based on the target query statement by using the retrieval model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to large language models, AI intelligent assistants, and machine learning technology, and specifically to sample generation methods, devices, intelligent agents, electronic devices, storage media, and program products based on large models. Background Art

[0002] In the field of machine learning technology, for retrieval models, random sampling or sampling based on sampling rules is usually used to increase training samples so as to optimize the search model.

[0003] However, optimizing the retrieval model by increasing training samples through random sampling or sampling based on sampling rules is difficult to meet the performance requirements of users for the retrieval model. Summary of the Invention

[0004] The present disclosure provides a sample generation method, device, intelligent agent, electronic device, storage medium and program product based on a large model.

[0005] According to one aspect of the present disclosure, a sample generation method based on a big model is provided, including: receiving at least one question selected by an object on a question analysis interface; the at least one question is obtained by performing a cause analysis on an abnormal query result output by a retrieval model; and based on the at least one question, using the big model to modify a target query statement and a target query result to generate a target sample, so as to optimize the retrieval model for the at least one question based on the target sample; wherein the target query statement is associated with an initial query statement in the abnormal query result; and the target query result is obtained by querying the target query statement using the retrieval model.

[0006] According to another aspect of the present disclosure, a sample generation device based on a large model is provided, including: a receiving module for receiving at least one question selected by an object on a question analysis interface; the at least one question is obtained by performing a cause analysis on an abnormal query result output by a retrieval model; and a generation module for modifying a target query statement and a target query result using a large model based on the at least one question to generate a target sample, so as to optimize the retrieval model for the at least one question based on the target sample; wherein the target query statement is associated with an initial query statement in the abnormal query result; and the target query result is obtained by querying the target query statement using the retrieval model.

[0007] According to another aspect of the present disclosure, an intelligent agent for sample generation is provided, comprising: an input module for receiving at least one question selected by an object on a question analysis interface; a processing module for determining a target task based on at least one question information received by the input module, determining a target macro model based on the target task, and obtaining a target sample for optimizing a retrieval model by calling the target macro model to execute the above method; and an output module for outputting the target sample obtained by the processing module.

[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the above method.

[0010] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the above method when executed by a processor.

[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0013] Figure 1 Schematically illustrates an exemplary system architecture to which a large model-based sample generation method and apparatus can be applied according to an embodiment of the present disclosure;

[0014] Figure 2 The flowchart of the sample generation method based on the large model according to the embodiment of the present disclosure is schematically shown;

[0015] Figure 3 A schematic diagram schematically illustrates target query sentence generation according to an embodiment of the present disclosure;

[0016] Figure 4 The following schematically illustrates a target strategy determination according to an embodiment of the present disclosure;

[0017] Figure 5 A schematic diagram schematically illustrates an example of a specific target strategy according to an embodiment of the present disclosure;

[0018] Figure 6 Schematically shows a schematic diagram of target sample generation according to an embodiment of the present disclosure;

[0019] Figure 7 Schematically shows a schematic diagram of target text generation according to an embodiment of the present disclosure;

[0020] Figure 8 A schematic diagram schematically illustrates a prompt design according to an embodiment of the present disclosure;

[0021] Figure 9 Schematically shows a schematic diagram of target sample generation according to another embodiment of the present disclosure;

[0022] Figure 10 Schematically shows a schematic diagram of training sample generation according to an embodiment of the present disclosure;

[0023] Figure 11 Schematically shows a block diagram of a sample generation device based on a large model according to an embodiment of the present disclosure;

[0024] Figure 12 Schematically shows a block diagram of an intelligent agent for sample generation according to an embodiment of the present disclosure; and

[0025] Figure 13 A block diagram of an electronic device suitable for implementing a sample generation method based on a large model according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0026] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0027] For retrieval models, training samples are increased through random sampling or sampling based on sampling rules. However, because retrieval models cannot accurately target specific abnormal samples when using random sampling or sampling based on sampling rules, the retrieval model lacks accuracy in meeting user query requirements, affecting the user search experience and the quality of the retrieval model. For example, there may be a mismatch between the query statement and the query results of abnormal samples.

[0028] Although abnormal samples can be manually labeled, manual labeling is inefficient and difficult to cover all abnormal samples.

[0029] In addition, although performance can be improved by improving the model architecture of the retrieval model, the improved retrieval model still lacks the ability to perform targeted optimization for complex semantic understanding problems.

[0030] In light of this, the disclosed embodiments address the issue of abnormal query results output by the retrieval model by utilizing a large model to modify the target query statements associated with the initial query statements in the abnormal query results and the target query results obtained by the retrieval model query, thereby generating targeted training samples. Because the training samples are generated by targeted modifications targeting the issue that causes the abnormal query results output by the retrieval model, optimizing the retrieval model based on these training samples can improve the retrieval model's ability to identify abnormal query results, enhance the accuracy of the query results output by the retrieval model, and improve the user experience.

[0031] Figure 1 An exemplary system architecture to which the large model-based sample generation method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.

[0032] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not imply that the embodiments of the present disclosure may not be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the large-model-based sample generation method and apparatus may be applied may include a terminal device, but the terminal device may implement the large-model-based sample generation method and apparatus provided in the embodiments of the present disclosure without interacting with a server.

[0033] like Figure 1 As shown, the exemplary architecture 100 may include a terminal device 101 , an agent 102 , a server 103 and a model library 104 .

[0034] Various communication client applications can be installed on the terminal device 101, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients and / or social platform software, AI smart assistants, etc. (only as examples).

[0035] The terminal device 101 may be any electronic device having a display screen and supporting web browsing, including but not limited to a smart phone, a tablet computer, a laptop computer, a desktop computer, and the like.

[0036] The intelligent agent 102 can identify user needs based on a large model, such as a large language model, and output information that meets the user needs.

[0037] The model library 104 may, for example, include but is not limited to multiple models trained based on various deep learning algorithms.

[0038] Server 103 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by a user on terminal device 101. The background management server may analyze and process received data such as user requests, and provide feedback (e.g., web pages, information, or data obtained or generated based on user requests) to terminal device 101.

[0039] For example, a user can input at least one question selected on the question analysis interface into terminal device 101. Terminal device 101 can then invoke agent 102 to determine a target task based on the at least one question. Based on the target task, the agent can then invoke a target macromodel from model library 104 and use the target macromodel to modify the target query statement and query results to generate a target text. The target text generated by the target macromodel can then be used to optimize the search model, addressing the aforementioned difficulty in meeting the user's performance requirements for the search model.

[0040] It should be noted that the sample generation method based on the large model provided in the embodiment of the present disclosure can generally be executed by the terminal device 101. Accordingly, the sample generation apparatus based on the large model provided in the embodiment of the present disclosure can also be set in the terminal device 101.

[0041] Alternatively, the sample generation method based on the large model provided in the embodiment of the present disclosure may also be generally executed by the server 103. Accordingly, the sample generation device based on the large model provided in the embodiment of the present disclosure may generally be set in the server 103. The sample generation method based on the large model provided in the embodiment of the present disclosure may also be executed by a server or server cluster that is different from the server 103 and can communicate with the terminal device 101 and / or the server 103. Accordingly, the sample generation device based on the large model provided in the embodiment of the present disclosure may also be set in a server or server cluster that is different from the server 103 and can communicate with the terminal device 101 and / or the server 103.

[0042] For example, terminal device 101 can send at least one question selected by a user on a question analysis interface to server 103. After receiving the at least one question, server 103 can invoke agent 102 to generate a target sample by executing the large model-based sample generation method of an embodiment of the present disclosure. The target sample is then fed back to terminal device 101 so that the agent can use the target sample for model optimization.

[0043] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0044] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0045] Figure 2 The flowchart of the sample generation method based on the large model according to the embodiment of the present disclosure is schematically shown.

[0046] like Figure 2 As shown, the method 200 includes operations S210 to S220.

[0047] In operation S210 , at least one question selected by a subject on a question analysis interface is received.

[0048] In operation S220, based on at least one question, the target query statement and the target query result are modified using the large model to generate a target sample, so as to optimize the retrieval model for the at least one question based on the target sample.

[0049] In an embodiment of the present disclosure, at least one question is obtained by analyzing the cause of an abnormal query result output by the retrieval model.

[0050] For example, the subject may include, but is not limited to, a user who trains or optimizes a retrieval model. The question analysis interface may be an interactive interface used by the agent to display analysis results. The retrieval model may be any model trained using methods such as machine learning and deep learning.

[0051] The at least one problem may include, but is not limited to, problems indicating missing keywords, abnormal entity matching, abnormal word order, text errors, abnormal keyword importance, abnormal synonyms, etc.

[0052] Keywords may include, for example, at least one of the following: words in a query statement, words in a text title, important words in a text abstract, important words in a text paragraph, and the like.

[0053] An entity may include, for example, at least one of the following: a person's name, a place name, an organization, a subject entity, etc. in the text.

[0054] Word order anomalies may include at least one of the following: different semantics or the same semantics caused by disordered order of words in the text, etc.

[0055] Text errors may include, for example, at least one of the following: text that does not conform to language standards, text with logical errors, text with inaccurate information, etc. Keyword importance may include, for example, the importance of the keyword to the search requirements, etc. Synonym anomalies may include, for example, the failure to identify or associate a word with the same or similar meaning as a word in the text.

[0056] For example, an abnormal query result indicates that at least one text output by the model does not match the initial query. For example, if the initial query is "Can physical methods be used to reduce fever in children?", an abnormal query result could be at least one of the following: "Texts about physical methods for reducing fever in adults," "Texts discussing the health effects of colds on children," "Texts about drug treatments for fever in children," etc.

[0057] For example, for an abnormal query result such as "text about physical cooling methods for adults with fever," the agent invokes the large model to analyze the cause and output an indication of missing keywords. For an abnormal query result such as "text about the impact of colds on children's health," the agent invokes a large model, such as a large language model, leveraging its deep understanding capabilities to analyze the cause of the abnormal query result and output an indication of abnormal word order. For an abnormal query result such as "text about drug treatment methods for children with fever," the agent outputs an indication of a mismatch in subject entities.

[0058] In the embodiment of the present disclosure, the target query statement is associated with the initial query statement in the abnormal query result. The target query result is obtained by querying the target query statement using a retrieval model.

[0059] For example, a target query sentence with similar semantic features can be determined based on the semantic features of the initial query sentence, and the target query sentence is input into the retrieval model to output the target query result.

[0060] For example, if the initial query is "Can children's fevers be reduced physically?", the target query can be "Can children's fevers be reduced physically?" and / or "Can children's fevers be reduced physically? and / or Can children's fevers be reduced with a wet towel?" The target query results can be "Text about children's fevers being reduced physically" and / or "Text about children's fevers being reduced with a wet towel."

[0061] For example, to address the problem of missing indicator keywords, for the target query sentence "Can children's fevers be cooled down physically?", the keyword "child" can be modified from "text about whether children's fevers can be cooled down physically" to "text about whether adults' fevers can be cooled down physically".

[0062] The disclosed embodiments address the issue of abnormal query results output by the retrieval model by utilizing a large model to modify the target query statements associated with the initial query statements in the abnormal query results and the target query results obtained by the retrieval model, thereby generating targeted training samples. Because the training samples are generated by targeted modifications targeting the issue that causes the abnormal query results output by the retrieval model, optimizing the retrieval model based on these training samples can improve the retrieval model's ability to identify abnormal query results, enhance the accuracy of the query results output by the retrieval model, and improve the user experience.

[0063] Reference below Figures 3 to 10 , combined with specific embodiments Figure 2 The method shown is further explained.

[0064] In an embodiment of the present disclosure, the cause analysis of the abnormal query results output by the retrieval model may specifically include the following operations: generating a question prompt based on the abnormal query result and the initial query statement of the abnormal query result, inputting the question prompt into the intelligent agent, and the intelligent agent calling the large language model to output at least one question.

[0065] Through question prompts, the large language model can understand the task requirements more clearly and the questions it asks are more accurate.

[0066] In some embodiments of the present disclosure, in order to improve the recognition ability of the retrieval model for the same type of query statements, the sample generation method based on the large model is as follows: Figure 2 In addition to the operations S210 to S220 shown, the process further includes performing content enhancement on the initial query statement in the abnormal query result to generate a target query statement.

[0067] Exemplarily, content enhancement may include, but is not limited to: logic enhancement, detail richness enhancement, synonym generalization, same-domain generalization, and cross-domain generalization.

[0068] The generalization of synonymous expressions can include maintaining complete semantic elements and changing the expression method.

[0069] Same-domain generalization can include expanding query scenarios, query roles, query time, spatial dimensions, etc. within the same domain.

[0070] Cross-domain generalization can include extending queries to other domains through structural mapping.

[0071] By enhancing the content of the initial query statement, the retrieval model's ability to recognize the same type of query statements can be improved.

[0072] Figure 3 A schematic diagram of target query sentence generation according to an embodiment of the present disclosure is schematically shown.

[0073] like Figure 3 As shown, in this embodiment, enhancing the content of the initial query statement in the abnormal query result to generate a target query statement may include the following operations: determining query elements 302 in the initial query statement 301. Based on the query elements 302, the initial query statement 301 is enhanced using the large model M to generate a target query statement 303.

[0074] Exemplarily, the query elements are used to indicate key conditions used in the initial query statement 301 to locate and filter query results.

[0075] For example, when the initial query statement is "Can physical cooling be used to treat a child's fever?", the query elements may be "child," "fever," "physical cooling," etc.

[0076] When the query element is "children", the content is enhanced to "children, infants", etc. When the query element is "fever", the content is enhanced to "fever, elevated body temperature, high fever", etc. When the query element is "physical cooling", the content is enhanced to "warm water wipe, ice compress", etc.

[0077] By determining the query elements of the initial query statement and using the large model to enhance the content of the query elements, the target query statement can maintain similar semantic features with the initial query statement, achieve generalization of the initial query statement, and improve the coverage of samples in the same type of queries, so as to improve the generalization ability of the model.

[0078] Exemplarily, the query elements may include at least one of the following: query scenario, query role, query dimension, query field, etc.

[0079] For example, the query elements may include query scenarios. According to the query elements, using the big model to enhance the content of the initial query statement to generate the target query statement may include the operation of: according to the query scenario, using the big model to expand the query scenario of the initial query statement to generate the target query statement.

[0080] The query context may be used to indicate the specific context in which the query is being performed.

[0081] Expanding a query scenario can include replacing it with a scenario of the same category but with different specific entities, for example, expanding "Brand A mobile phone repair" to "Brand B mobile phone repair." Expanding a query scenario can also include adjusting the specific parameters of the qualifying conditions, for example, expanding "mobile phones under 3,000 yuan" to "mobile phones under 2,000 yuan." Expanding a query scenario can also include replacing a request with a different expression of the same function, for example, expanding "weight loss recipes" to "calorie-controlled diet plans."

[0082] By expanding the query scenarios, more queries with similar semantic features to the initial query statements can be generated, thereby improving the coverage of specific question types from the query scenarios.

[0083] For example, the query elements may also include query roles. According to the query elements, the initial query statement is enhanced using the big model to generate a target query statement. The operation may also include: according to the query roles, the query roles of the initial query statement are expanded using the big model to generate a target query statement.

[0084] Query roles can be used to indicate entities involved in the query process. Different query domain types have different query roles. The query role extension can be to change the query role within the query domain type. For example, professional consulting: changing the specific consulting object in professional consulting fields such as medical care, law, and education; commodity and service: changing the specific commodity or service in e-commerce, local life and other fields; content consumption: changing the specific content type in media, entertainment and other fields; tool query: changing the specific query type in information query, online tools and other fields; life service: changing the specific service type in transportation, life payment and other fields.

[0085] By further expanding the query roles, more queries with similar semantic features to the initial query statements can be generated, improving content enhancement capabilities and increasing the coverage of specific question types from the query roles.

[0086] For example, the query elements may also include query dimensions. According to the query elements, the initial query statement is enhanced using the big model to generate a target query statement. The operation may also include: according to the query dimensions, the query dimensions of the initial query statement are expanded using the big model to generate a target query statement.

[0087] Query dimensions may include query time and query space dimensions. For example, the latest product of a certain brand now is expanded to the latest product of a certain brand at the current query time; nearby restaurants are expanded to restaurants near the user's location, etc.

[0088] By further expanding the query time and space dimensions, the coverage of specific question types is improved from the query time and space dimensions.

[0089] For example, the query elements may also include a query domain. According to the query elements, using the big model to enhance the content of the initial query statement to generate a target query statement may include the following operations: according to the query domain, using the big model to expand the query domain of the initial query statement to generate a target query statement.

[0090] Expanding the query domain can include extending the query to different domains through structural mapping and element migration. Structural mapping involves generalizing the semantic structure of the initial query and the relationships between its elements to ensure that the expanded target query retains all query elements from the initial query. Element migration involves converting the subject object, constraints, requirements, and scenarios of the initial query into corresponding entities, constraints, demands, and scenarios across domains.

[0091] For example, the medical consultation field has expanded to the legal resources field. The product purchase field has expanded to the service reservation field. The education and training field has expanded to the skill learning field. The job search field has expanded to the school selection field.

[0092] By further expanding the query domain, the coverage of specific question types from the query domain is improved.

[0093] Therefore, the initial query statement can be diversified and expanded through the various query elements in the initial query statement to achieve generalization of the initial query statement. The resulting multiple target query statements can cover the query situations of specific question types and realize the retrieval model's ability to recognize the same type of query statements.

[0094] According to the embodiments of the present disclosure, for the above Figure 2 Operation S220, in which a target query statement and target query results are modified using the large model based on at least one question to generate target samples, so as to optimize the retrieval model for the at least one question based on the target samples, may include the following operations: determining a target strategy based on the at least one question; and modifying the target query statement and target query results using the large model based on the target strategy to generate target samples.

[0095] In the embodiment of the present disclosure, the target strategy indicates the distinguishing factor between positive samples and negative samples for optimizing the retrieval model.

[0096] For example, positive samples may include text and query statements from among multiple texts obtained using a query statement that meet the query requirements. Negative samples may include text and query statements from among multiple texts obtained using a query statement that do not meet the query requirements. Meeting the query requirements can be understood as the credibility of the text obtained using the query statement being greater than a credibility threshold. The credibility threshold may be an empirical value.

[0097] Figure 4 The following schematically illustrates a target strategy determination according to an embodiment of the present disclosure; Figure 5 A schematic diagram schematically shows a specific target strategy example according to an embodiment of the present disclosure.

[0098] For example, a target strategy matching the type of at least one question can be determined based on a preset mapping relationship according to the type of at least one question. The preset mapping relationship represents the mapping relationship between the type of question and the target strategy, such as Figure 4 As shown, the first preset mapping relationship 410 can represent a mapping relationship between a first type of problem and a first target strategy. The second preset mapping relationship 420 can represent a mapping relationship between a second type of problem and a second target strategy. Similarly, the Nth preset mapping relationship 4N0 can represent a mapping relationship between an Nth type of problem and an Nth target strategy. N is an integer greater than or equal to 1.

[0099] For example, question types may include: missing keywords, abnormal entity matching, abnormal word order, abnormal synonyms, etc. The corresponding target strategy can be configured in advance for each question type, such as Figure 5 As shown, a first mapping relationship 541 between "First Type of Problem: Missing Keywords" and "First Target Strategy: Highlighting the Advantages of the Target Text over Texts to be Corrected with Missing Keywords"; a second mapping relationship 542 between "Second Type of Problem: Entity Matching Anomalies" and "Second Target Strategy: Highlighting the Advantages of the Target Text over Texts to be Corrected with Entity Matching Anomalies"; a third mapping relationship 543 between "Third Type of Problem: Word Order Anomalies" and "Third Target Strategy: Highlighting the Advantages of the Target Text and Texts with Standard Word Order over Texts to be Corrected with Word Order Anomalies"; and a fourth mapping relationship 544 between "Fourth Type of Problem: Synonym Anomalies" and "Fourth Target Strategy: Highlighting the Advantages of the Target Text and Texts with Synonyms over Texts to be Corrected without Synonym Anomalies." The distinguishing elements in each target strategy are: keywords, entities, word order, and synonyms.

[0100] For example, in order to solve the problem of keywords being located in different text fields, the advantages of the target text over the text to be corrected in which keywords are missing can be highlighted. For example, the target text in which keywords are located in different text fields is better than text without key information or irrelevant text.

[0101] For example, in response to the problem of keywords being located at different positions in a text field, the advantages of the target text over the text to be corrected in which keywords are missing can be highlighted. For example, the target text in which keywords are located at different positions in the text field is better than text without key information or irrelevant text.

[0102] For example, in order to solve the problem of keywords being located in different text fields and keywords being located in different positions in a text field, the above two situations can be combined to highlight the advantages of the target text over the text to be corrected in which keywords are missing.

[0103] For example, if the target query result is a text with low confidence, the low-confidence text can be repaired to obtain a repaired positive sample. For the positive sample text, the target sample can be obtained by modifying the keywords, thereby further increasing the base number of sample amplification.

[0104] Since a target strategy for training samples is generated for the specific type of problem in which the retrieval model outputs abnormal query results, and the difference between positive samples and negative samples for this specific type of problem is indicated in the target strategy, the retrieval model can learn how to reduce the output of abnormal query results based on these training samples.

[0105] Since multiple target query statements can be obtained by performing content enhancement based on the initial query statement, and multiple texts can also be obtained by querying using one target query statement, the target query result may include multiple initial texts.

[0106] Figure 6 The diagram schematically shows a target sample generation according to an embodiment of the present disclosure.

[0107] like Figure 6 As shown, in this embodiment, based on the target strategy, the target query statement and target query result are modified using the large model to generate the target sample, which may include the following operations: determining a target text 602 from multiple initial texts 601. Based on the target strategy, the target query statement 303 and target text 602 are modified using the large model to generate the target sample.

[0108] In the embodiment of the present disclosure, the matching degree between the target text 602 and the target query statement 303 is greater than or equal to a first predetermined threshold.

[0109] For example, the first predetermined threshold can be determined based on the degree of match between the text that meets the query requirements and the query statement. Since the degree of match between the target text 602 and the target query statement 303 is greater than or equal to the first predetermined threshold, it can be determined that the target text 602 is a positive sample text.

[0110] For example, the content in the target text 602 can be modified so that it has the same problems as those indicating missing keywords and / or abnormal entity matching and / or abnormal word order and / or text errors and / or abnormal keyword importance and / or abnormal synonyms, so that the target sample obtained belongs to the text of the negative sample.

[0111] For example, the text to be revised about "discussing physical cooling methods for children or young children with fever" is modified into the target text about "discussing drug treatment methods for adults with fever" using the large language model, and is determined as the target sample together with the target query sentence.

[0112] Since the target text is the text that matches the target query statement, the training samples generated based on the target text modification can construct a variety of negative samples.

[0113] In another embodiment of the present disclosure, based on a target strategy, a target query statement and target query result are modified using a large model to generate a target sample, which may further include determining a text to be modified from multiple initial texts, and modifying the text to be modified using the large model to generate a target text.

[0114] In the embodiment of the present disclosure, the matching degree between the text to be corrected and the target query is less than the first predetermined threshold. Since the matching degree between the text to be corrected and the target query is less than the first predetermined threshold, it can be determined that the text to be corrected is a negative sample text.

[0115] Exemplarily, the text to be corrected belonging to the negative sample can be repaired into the target text belonging to the positive sample.

[0116] For example, the text to be corrected about "discussing physical cooling methods for adults with fever" is repaired into the target text about "discussing physical cooling methods for children or young children with fever" using the large language model.

[0117] Since the text to be corrected whose matching degree with the target query statement is less than the first predetermined threshold is corrected using a large model, the matching degree with the target query statement can be improved, so that the negative sample can be converted into a positive sample after modification. Therefore, the number of positive samples can be increased, which is conducive to generating more target samples.

[0118] Figure 7 The following schematically illustrates a schematic diagram of target text generation according to an embodiment of the present disclosure.

[0119] like Figure 7 As shown, in this embodiment, using the large model to correct the text to be corrected and generate the target text may include the following operations: constructing prompt text 704 based on the target query sentence 303, the text to be corrected 701, the matching degree 702 between the text to be corrected 701 and the target query sentence 303, and the reference example 703. The prompt text 704 is input into the large model M, and the target text 602 is output.

[0120] In the embodiment of the present disclosure, the reference examples 703 include reference texts whose matching degree with the reference query statement is greater than a first predetermined threshold.

[0121] Figure 8 The figure schematically shows a prompt design according to an embodiment of the present disclosure.

[0122] For example, the prompt text 704 can be as follows Figure 8The template of the prompt design 800 is shown as being generated. The template of the prompt design 800 may include an environment description 810, a task definition 820, constraints 830, an output format 840, and input data 850.

[0123] For example, the environmental description 810 can be "You are a search engine optimization expert and need assistance in fixing matching problems in the retrieval model. These problems mainly manifest as inaccurate matching between query statements and query results, specifically including: inverse matching: high-matching query results have lower matching scores than low-matching query results, and single query result anomalies: the similarity of query results does not match the actual situation."

[0124] Task definition 820 may be “1. Analyze the problem: Carefully analyze the core intent and requirements of the user’s query generalization problem, and determine the specific manifestations of similar abnormal query results. 2. Determine key features… 3. Clearly list the conditions required for positive samples…”.

[0125] Constraint 830 may be “1. Authenticity: in line with the user's actual search habits. 2. Professionalism…”.

[0126] The output format 840 may describe the output format, etc.

[0127] Input data 850 may include a description of the input data, etc.

[0128] For example, the template of the prompt design 800 can be filled in according to the target query statement 303, the text to be corrected 701, the matching degree 702 between the text to be corrected 701 and the target query statement 303, and the reference example 703 to obtain the prompt text 704.

[0129] By constructing prompt text, the large model can understand the task requirements more clearly and accurately output the target text belonging to the positive sample.

[0130] According to another embodiment of the present disclosure, based on a target strategy, a target query statement and a target query result are modified using a large model to generate a target sample. The generation of a target sample may include the following operations: in response to determining that at least one question indicates a missing keyword, determining that the distinguishing factor indicated by the target strategy is a keyword. Based on the target strategy, a field in a target text corresponding to the keyword in the target query statement is modified using the large model to generate a first text. The target query statement, the target text, and the first text are determined as a target sample.

[0131] For example, keywords may include words in a query sentence.

[0132] For example, for target query statements such as "Can children's fever be cooled down physically?" and / or "Can children's fever be cooled down physically?" and / or "Can children's fever be cooled down physically?" and / or "Can children's fever be cooled down with a wet towel?", the target text about "Can children or children or children's fever be cooled down physically" is obtained. For the problem of missing keywords "children and / or children and / or children", the target strategy can completely remove all content about "children and / or children and / or children", and change the content to only discuss physical cooling methods for adults with fever, without mentioning children and / or children and / or children, and ensure that any child-related words such as "children and / or children and / or children" do not appear in the text title, text summary, and text paragraphs to ensure the professionalism of the content. At the same time, it is ensured that it only involves physical cooling of adult fever. It is not a simple replacement of words, but an overall adjustment to an adult perspective to obtain the first text. The target strategy for the problem of missing the keyword "fever" is to completely remove all content related to "fever" and instead discuss physical care methods for other health problems of children (such as colds, coughs, etc.), but not involving fever or body temperature. Ensure that any fever-related words such as "fever and / or increased body temperature and / or high fever" do not appear in the text title, text summary, and text paragraphs. Keep the content focused on children, but change the health problem. Do not simply delete the words, but adjust the content direction as a whole to obtain the first text.

[0133] Since the first text is the text of the negative sample obtained by replacing the target text in the positive sample, in order to solve the problem of missing keywords, using a large model to replace the first text constructed by keywords such as the title, abstract, paragraph, etc. in the text of the positive sample can improve the model's ability to recognize keywords.

[0134] According to another embodiment of the present disclosure, based on a target policy, a target query statement and target query results are modified using a large model to generate a target sample, which may include the following operations: in response to determining that at least one question indicates an abnormal entity match, determining that the distinguishing factor indicated by the target policy is an entity. Based on the target policy, a field in the target text corresponding to the entity in the target query statement is modified using the large model to generate a second text. The target query statement, target text, and the second text are determined as a target sample.

[0135] For example, an entity matching anomaly may include an entity having a high degree of distinction or an entity being completely irrelevant. An entity may include a subject entity, a person entity, etc. A modification operation may, for example, replace a field corresponding to an entity with a completely dissimilar word.

[0136] For example, the target strategy for the entity matching anomaly problem can be to completely remove all content about the entity "physical cooling" and instead discuss other treatment methods for children's fever (such as drug treatment) without mentioning any content related to physical cooling. Ensure that any physical cooling-related words such as "physical cooling and / or warm water bath and / or ice compress" do not appear in the text title, text summary, and text paragraphs. Maintain the focus on children's fever, but only discuss other treatment methods. Do not simply replace words, but change the direction of the treatment method as a whole to obtain the second text.

[0137] To address the problem of abnormal entity matching, a large model is used to replace multiple similar expressions of entities in the target text of positive samples to construct negative samples. Therefore, the model's ability to recognize similar expressions of entities can be improved.

[0138] According to another embodiment of the present disclosure, based on a target strategy, a target query statement and a target query result are modified using a large model to generate a target sample. The generation of a target sample may include the following operations: in response to determining that at least one question indicates a word order anomaly, determining that the distinguishing factor indicated by the target strategy is word order. Based on the target strategy, the large model is used to modify the word order of content fields associated with the target query statement in the target text to generate a third text. The target query statement, the target text, and the third text are determined as target samples.

[0139] For example, the word order anomaly may include an anomaly in which the word order is disordered but the semantics are the same as those in the original order, or an anomaly in which the word order is disordered but the semantics are different from those in the original order.

[0140] Word order modification operations may include changing the position order between words, etc.

[0141] For example, the content field associated with the target query statement in the target text, "When the child's body temperature does not exceed 38.5°C, physical cooling can be given priority, which is a safe and effective auxiliary means," is modified in word order to "When the child's body temperature does not exceed 38.5°C, auxiliary means can be given priority, which is a safe and effective physical cooling" and / or "When the child's body temperature exceeds 38.5°C, physical cooling is not given priority, but it is a safe and effective auxiliary means," etc.

[0142] To address the problem of abnormal word order, a large model is used to modify the word order of the target text in the positive sample without changing the semantics, or to modify the word order of the target text in the positive sample and then change the original semantics, thereby constructing difficult samples that are more difficult to recognize, which can improve the retrieval model's ability to recognize word order.

[0143] Since for each target text in the initial text or the target text obtained after repair, if the entire text content is modified each time, it will result in the consumption of more computing resources, based on this, the present disclosure minimizes the amount of modification and saves computing resources.

[0144] Figure 9 The figure schematically shows a schematic diagram of target sample generation according to another embodiment of the present disclosure.

[0145] like Figure 9 As shown, based on the target strategy for at least one question, the target query statement and target query result associated with the initial query statement in the abnormal query result are modified using a large model to generate a target sample, which may include the following operations: based on the target strategy for at least one question, determining a field to be modified 901. Based on the target strategy, the target query statement and target query result are modified using a large model with respect to the field to be modified 901 to generate an intermediate sample 902. In response to determining that the degree of match 904 between the intermediate sample 902 and the modification requirement 903 indicated by the target strategy is less than a second predetermined threshold, returning to perform the determination operation on the field to be modified 901. In response to determining that the degree of match 904 between the intermediate sample 902 and the modification requirement 903 indicated by the target strategy is greater than or equal to the second predetermined threshold, determining that the intermediate sample 902 is the target sample 603.

[0146] Exemplarily, the target strategy may further include a modification requirement 903 . The modification requirement 903 may be, for example, that the modified text complies with semantic logic and that the text is modified as little as possible.

[0147] The fields to be modified 901 may be necessary fields to be modified determined from the target query result based on the target strategy, the target to be modified and the modification requirement.

[0148] The second predetermined threshold may be determined in advance based on the degree of matching between the text meeting the modification requirement and the modification requirement.

[0149] For example, while meeting the modification requirements, determine the minimum necessary modifications to the original text in the target query results. These modifications might include retaining all original paragraph structures unrelated to the modification objective and / or preserving the original sentence structure as much as possible, replacing only keywords or phrases, and / or retaining subject entities, place names, facility descriptions, etc. in the original text if there are no conflicts.

[0150] Since the fields to be modified are determined based on the target strategy, the intermediate text is modified. Therefore, the target sample can be generated in an iterative form with as few modifications as possible while meeting the modification requirements, thereby reducing the consumption of computing resources.

[0151] In other embodiments of the present disclosure, the sample generation method based on the large model may include the following: Figure 2 In addition to the operations S210 to S220 shown, the method may further include: based on at least one problem, using a large model to perform sample enhancement on the target sample to generate an augmented sample, so as to optimize the retrieval model using the target sample and the augmented sample.

[0152] Exemplarily, performing sample enhancement on the target sample may include performing synonym replacement on the target text and the target query sentence in the target sample or adjusting the word order of the target text and the target query sentence in the target sample.

[0153] Figure 10 The diagram schematically shows the generation of training samples according to an embodiment of the present disclosure.

[0154] like Figure 10 As shown, the initial query statement 301 corresponding to the abnormal query result 1001 is enhanced with content such as synonymous expressions, same-domain generalization, and cross-domain generalization to obtain a target query statement 303. The target query statement 303 is then input into the retrieval model, and a target query result 1003 is output. The target query result 1003 may include multiple initial texts. From the multiple initial texts in the target query result 1003, a target text 602 is determined whose matching degree with the target query statement is greater than or equal to a first preset threshold. Analysis of the cause of the abnormal query result 1001 can identify at least one issue, and a target strategy 1002 can be determined based on the at least one issue. Based on the target strategy 1002, the target text 602 can be modified and combined with the target query statement to obtain a target sample 603. The target sample 603 can be enhanced to obtain an augmented sample 1004. The augmented sample 1004 and the target sample 603 are used together as training samples 1005. The training samples 1005 can be used to optimize the retrieval model and enhance its performance.

[0155] Since sample enhancement is performed on all modified target samples, the number of training samples can be increased, which is beneficial to the optimization of the retrieval model, improving the model retrieval efficiency and retrieval accuracy, etc.

[0156] For example, based on at least one question, using a large model to perform sample enhancement on a target sample to generate an augmented sample may include the following operations: in response to determining that at least one question indicates word order abnormality, using the large model to adjust the word order of the target sample to generate an augmented sample.

[0157] In the disclosed embodiment, the amplified sample and the target sample have different word orders and similar semantics in their contents.

[0158] For example, the word order of the text in the target sample can be adjusted so that the word order of the content expressed by the text in the augmented sample is different from that of the target sample, but the semantics are similar.

[0159] Since the target sample is further expanded based on word order adjustment to achieve sample enhancement, the retrieval model's ability to recognize word order changes can be improved.

[0160] For example, based on at least one question, using a large model to perform sample enhancement on a target sample to generate an augmented sample may include the following operations: in response to determining that at least one question indicates a matching anomaly, using the large model to perform entity adjustment on the target sample to generate an augmented sample.

[0161] In the embodiment of the present disclosure, the amplified sample and the target sample express contents in different ways but have similar semantics; the matching anomaly includes at least one of the following: entity matching anomaly and keyword matching anomaly.

[0162] For example, entity matching anomalies may include entity matching degree anomalies, and keyword matching anomalies may include keyword missing, etc.

[0163] The entity adjustment operation may include, for example, replacing words in the target sample text with synonyms, or replacing words indicating entities with different expressions.

[0164] Since the target sample is adjusted again based on the entity, synonym expansion is performed by changing the way synonyms or entity expressions are expressed, and sample enhancement is achieved, the retrieval model's ability to recognize changes in word order can be improved.

[0165] Figure 11 The block diagram of the apparatus for generating samples based on a large model according to an embodiment of the present disclosure is schematically shown.

[0166] like Figure 11 As shown, the sample generating device 1100 includes: a receiving module 1110 and a generating module 1120 .

[0167] The receiving module 1110 is configured to receive at least one question selected by the subject on the question analysis interface; the at least one question is obtained by performing a cause analysis on an abnormal query result output by the retrieval model.

[0168] The generation module 1120 is used to modify the target query statement and the target query result based on at least one question using the large model to generate a target sample, so as to optimize the retrieval model for at least one question based on the target sample.

[0169] According to an embodiment of the present disclosure, the target query statement is associated with the initial query statement in the abnormal query result; the target query result is obtained by querying the target query statement using a retrieval model.

[0170] According to an embodiment of the present disclosure, generation module 1120 includes a first determination unit and a first modification unit. The first determination unit is configured to determine a target strategy based on at least one question. The target strategy indicates the distinguishing factors between positive samples and negative text samples used to optimize the retrieval model. The first modification unit is configured to modify the target query statement and target query results using the large model based on the target strategy to generate a target sample.

[0171] According to an embodiment of the present disclosure, a target query result includes: multiple initial texts. A first determination unit includes a first determination subunit and a first modification subunit. The first determination subunit is used to determine a target text from the multiple initial texts. The degree of match between the target text and the target query statement is greater than or equal to a first predetermined threshold. The first modification subunit is used to modify the target query statement and target text using a large model based on a target strategy to generate a target sample.

[0172] According to an embodiment of the present disclosure, the first determination unit further includes a second determination subunit and a second modification subunit. The second determination subunit is configured to determine a text to be modified from a plurality of initial texts. The degree of match between the text to be modified and the target query statement is less than a first predetermined threshold. The second modification subunit is configured to modify the text to be modified using the large model to generate a target text.

[0173] According to an embodiment of the present disclosure, a large model is used to correct a text to be corrected to generate a target text. The method includes constructing a prompt text based on a target query, the text to be corrected, the degree of match between the text to be corrected and the target query, and reference examples. The reference examples include reference text whose degree of match with the reference query exceeds a first predetermined threshold. The prompt text is input into the large model, and the target text is output.

[0174] According to an embodiment of the present disclosure, the first determining unit includes a third determining subunit, a third modifying subunit, and a fourth determining subunit.

[0175] The third determining subunit is configured to, in response to determining that at least one question indicates a missing keyword, determine that the distinguishing factor indicated by the target policy is a keyword. The third modifying subunit is configured to, based on the target policy, utilize the large model to modify fields in the target text corresponding to the keywords in the target query to generate a first text. The fourth determining subunit is configured to determine the target query, the target text, and the first text as a target sample.

[0176] According to an embodiment of the present disclosure, the first determining unit includes a fifth determining subunit, a fourth modifying subunit, and a sixth determining subunit.

[0177] The fifth determination subunit is configured to, in response to determining that at least one question indicates an abnormal entity match, determine that the distinguishing factor indicated by the target policy is an entity. The fourth modification subunit is configured to, based on the target policy, utilize the large model to modify the fields in the target text corresponding to the entity in the target query to generate a second text. The sixth determination subunit is configured to determine the target query, the target text, and the second text as a target sample.

[0178] According to an embodiment of the present disclosure, the first determination unit includes a seventh determination subunit, a fifth modification subunit, and an eighth determination subunit. The seventh determination subunit is configured to, in response to determining that at least one question indicates a word order anomaly, determine that the distinguishing factor indicated by the target policy is word order. The fifth modification subunit is configured to, based on the target policy, utilize a large model to modify the word order of content fields associated with the target query statement in the target text to generate a third text. The eighth determination subunit is configured to determine the target query statement, the target text, and the third text as target samples.

[0179] According to an embodiment of the present disclosure, the generation module 1120 includes a second determination unit, a second modification unit, a return unit, and a third determination unit. The second determination unit is used to determine the field to be modified based on the target strategy for at least one problem. The second modification unit is used to modify the target query statement and the target query result for the field to be modified using the large model based on the target strategy to generate an intermediate sample. The return unit is used to return to perform the determination operation on the field to be modified in response to determining that the degree of match between the intermediate sample and the modification requirement indicated by the target strategy is less than a second predetermined threshold. The third determination unit is used to determine that the intermediate sample is a target sample in response to determining that the degree of match between the intermediate sample and the modification requirement indicated by the target strategy is greater than or equal to the second predetermined threshold.

[0180] According to an embodiment of the present disclosure, the sample generation device 1100 further includes a content enhancement module. The content enhancement module is configured to enhance the content of the initial query statement in the abnormal query result to generate a target query statement.

[0181] According to an embodiment of the present disclosure, the content enhancement module includes an element determination unit and a sentence generation unit. The element determination unit is used to determine query elements in an initial query sentence. The sentence generation unit is used to enhance the content of the initial query sentence using a large model based on the query elements to generate a target query sentence.

[0182] According to an embodiment of the present disclosure, the query elements include query scenarios. According to the query elements, the initial query statement is enhanced using the big model to generate a target query statement, including: according to the query scenario, the query scenario of the initial query statement is expanded using the big model to generate the target query statement.

[0183] According to an embodiment of the present disclosure, the query elements also include query roles. According to the query elements, the initial query statement is enhanced using the big model to generate a target query statement, including: according to the query roles, the query roles of the initial query statement are expanded using the big model to generate the target query statement.

[0184] According to an embodiment of the present disclosure, the query elements also include query dimensions. According to the query elements, the initial query statement is enhanced using the big model to generate a target query statement, including: according to the query dimensions, the query time of the initial query statement is expanded using the big model to generate the target query statement.

[0185] According to an embodiment of the present disclosure, the query elements also include query domains. According to the query elements, the initial query statement is enhanced using the big model to generate a target query statement, including: according to the query space, the query domain of the initial query statement is expanded using the big model to generate the target query statement.

[0186] According to an embodiment of the present disclosure, the sample generation apparatus 1100 further includes an amplification module. The amplification module is configured to enhance the target sample using the large model to generate an amplified sample based on at least one question, so as to optimize the retrieval model using the target sample and the amplified sample.

[0187] According to an embodiment of the present disclosure, the amplification module includes a first adjustment unit. The first adjustment unit is configured to, in response to determining that at least one question indicates a word order anomaly, use the large model to adjust the word order of a target sample to generate an amplified sample. The amplified sample and the target sample express content in a different word order but with similar semantics.

[0188] According to an embodiment of the present disclosure, the amplification module includes a second adjustment unit. The second adjustment unit is configured to, in response to determining that at least one problem indicates a matching anomaly, perform entity adjustments on the target sample using the large model to generate an amplified sample. The amplified sample and the target sample express content in different ways but with similar semantics. Matching anomalies include at least one of the following: entity matching anomalies and keyword matching anomalies.

[0189] According to an embodiment of the present disclosure, the present disclosure also provides an intelligent agent for sample optimization, an electronic device, a readable storage medium and a computer program product.

[0190] Figure 12 A block diagram of an intelligent agent for sample generation according to an embodiment of the present disclosure is schematically shown.

[0191] In the embodiments of the present disclosure, inspired by the von Neumann structure in modern computer theory, such as Figure 12As shown, the AI ​​agent 1200 may include three core modules: an input module 1210, an output module 1220, and a processing module 1230. The processing module 1230 may include a control unit 1231, a storage unit 1232, and an operation unit 1233.

[0192] Input module 1210 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment) and converting it into a format that AI agent 1200 can understand and process. Input module 1210 is the primary link for AI agent 1200 to interact with the outside world. It enables AI agent 1200 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.

[0193] In an example, the input information received by the input module 1210 may be at least one optimization problem described above.

[0194] In this example, processing module 1230 is the core support for AI agent 1200's ability to handle complex tasks. Processing module 1230 can determine a target task based on the input information received by input module 1210, determine a large model based on the target task, and execute the large model-based sample generation method described above by calling the large model to output the target sample.

[0195] In the example, the control unit 1231 in the processing module 1230 will continuously interact with the storage unit 1232, the computing unit 1233, and / or the output module 1220 during operation. However, it should be noted that in the embodiment of the present disclosure, the control unit 1231 acts as a single initiator to initiate communication with the storage unit 1232, the computing unit 1233, and / or the output module 1220, and there is no communication coupling between the storage unit 1232, the computing unit 1233, and the output module 1220.

[0196] In this example, the performance of control unit 1231 can be closely related to the large model on which AI agent 1200 is based. To fully utilize the capabilities of the large language model, the internal structure of control unit 1231 can be designed to be highly configurable and scalable to cope with various types of tasks and requirements in real-world scenarios.

[0197] The storage unit 1232 may be responsible for memorizing information such as historical conversations, event flows, etc. The configuration information, target text, and data resources generated in each round as described above may be included in the storage unit 1232 .

[0198] In the example, after the AI ​​agent 1200 obtains the configuration generation request, the AI ​​agent 1200 can use the intent recognition model to determine the configuration intent from the initial text. The configuration intent can be stored in the storage unit 1232. The AI ​​agent 1200 can retrieve the relevant data resources from the storage unit 1232 and feed it back to the control unit 1231. Then, the control unit 1231 can use the fed-back data resources to obtain the configuration data corresponding to the initial text. It can also retrieve relevant text data from the storage unit 1232 and feed it back to the control unit 1231. Then, the control unit 1231 can use the returned text data to obtain the target text. And pass the target text and configuration data to the output module 1220.

[0199] The operation unit 1233 can be regarded as a predefined tool library, and the renderer and display controls mentioned above can be included in the operation unit 1233.

[0200] In the example, when the AI ​​agent 1200 needs to render multiple output data, it can call the relevant renderer and display control from the operation unit 1233 and feed it back to the control unit 1232. Then, the control unit 1232 can use the feedback renderer and display control to render the first search result and pass the first search result to the output module 1220. It can be understood that although the large language model has excellent language understanding and generation capabilities, it is the same as a human. Without the help of any tools, the tasks that can be solved are very limited. When the AI ​​agent 1200 is given the ability to call tools, it can achieve tasks such as completing mathematical operations with the help of a calculator, completing data analysis with the help of Python, and completing prediction tasks with the help of a search engine.

[0201] In an example, the output module 1220 may output the target sample described above.

[0202] The AI ​​agent 1200 according to the embodiment of the present disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.

[0203] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0204] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above method.

[0205] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the above method when executed by a processor.

[0206] Figure 13 A block diagram of an electronic device suitable for implementing a large model-based sample generation method according to an embodiment of the present disclosure is schematically shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0207] like Figure 13 As shown, device 1300 includes a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1302 or a computer program loaded from a storage unit 1308 into a random access memory (RAM) 1303. RAM 1303 may also store various programs and data required for the operation of device 1300. Computing unit 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to bus 1304.

[0208] Various components in device 1300 are connected to I / O interface 1305, including an input unit 1306, such as a keyboard and mouse; an output unit 1307, such as various types of displays and speakers; a storage unit 1308, such as a magnetic disk and optical disk; and a communication unit 1309, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1309 allows device 1300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0209] Computing unit 1301 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1301 performs the various methods and processes described above, such as the large-model-based sample optimization method. For example, in some embodiments, the large-model-based sample optimization method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1300 via ROM 1302 and / or communication unit 1309. When the computer program is loaded into RAM 1303 and executed by computing unit 1301, one or more steps of the large-model-based sample optimization method described above can be performed. Alternatively, in other embodiments, the computing unit 1301 may be configured to execute the large model-based sample optimization method in any other appropriate manner (for example, by means of firmware).

[0210] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0211] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0212] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0213] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0214] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0215] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0216] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0217] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A sample generation method based on a large model, comprising: receiving at least one problem selected by a subject on a problem analysis interface; The at least one problem is obtained by analyzing the cause of an abnormal query result output by the retrieval model; as well as Based on the at least one question, modify the target query statement and the target query result using the large model to generate a target sample, so as to optimize the retrieval model for the at least one question based on the target sample; Wherein, the target query statement is associated with the initial query statement in the abnormal query result; The target query result is obtained by querying the target query statement using the retrieval model.

2. The method according to claim 1, wherein The step of modifying the target query statement and the target query result based on the at least one question using the large model to generate a target sample includes: Determining a target strategy based on the at least one question; wherein the target strategy indicates a distinguishing factor between positive samples and negative samples for optimizing the retrieval model; Based on the target strategy, the target query statement and the target query result are modified using the large model to generate the target sample.

3. The method according to claim 2, wherein: The target query result includes: multiple initial texts; The step of modifying the target query statement and the target query result based on the target strategy and using the large model to generate the target sample includes: Determining a target text from the multiple initial texts; wherein the matching degree between the target text and the target query statement is greater than or equal to a first predetermined threshold; and Based on the target strategy, the target query statement and the target text are modified using the large model to generate the target sample.

4. The method according to claim 2 or 3, wherein: The method further comprises: modifying the target query statement and the target query result using the large model based on the target strategy to generate the target sample; Determining a text to be revised from the multiple initial texts; wherein the matching degree between the text to be revised and the target query statement is less than the first predetermined threshold; The large model is used to correct the text to be corrected to generate the target text.

5. The method according to claim 4, wherein The method of using the large model to correct the text to be corrected to generate the target text includes: Constructing a prompt text based on the target query, the text to be revised, the degree of matching between the text to be revised and the target query, and reference examples; wherein the reference examples include reference texts whose degree of matching with the reference query is greater than the first predetermined threshold; and The prompt text is input into the large model, and the target text is output.

6. The method according to any one of claims 3 to 5, wherein: The step of modifying the target query statement and the target text using the large model based on the target strategy to generate the target sample includes: In response to determining that the at least one problem indication keyword is missing, determining that the distinguishing element of the target policy indication is a keyword; Based on the target strategy, using the large model, modifying the fields in the target text corresponding to the keywords in the target query sentence to generate a first text; and The target query statement, the target text, and the first text are determined as the target sample.

7. The method according to any one of claims 3 to 6, wherein The step of modifying the target query statement and the target text using the large model based on the target strategy to generate the target sample includes: In response to determining that the at least one problem indicates an entity match anomaly, determining that the distinguishing element indicated by the target policy is an entity; Based on the target strategy, using the large model, modifying the fields in the target text corresponding to the entities in the target query statement to generate a second text; and The target query statement, the target text, and the second text are determined as the target sample.

8. The method according to any one of claims 3 to 7, wherein: The step of modifying the target query statement and the target text using the large model based on the target strategy to generate the target sample includes: In response to determining that the at least one problem indication has an abnormal word order, determining that the distinguishing element of the target policy indication is word order; Based on the target strategy, using the large model, modifying the word order of content fields associated with the target query in the target text to generate a third text; and The target query statement, the target text, and the third text are determined as the target sample.

9. The method according to any one of claims 1 to 8, wherein The method of modifying the target query statement and target query result associated with the initial query statement in the abnormal query result by using the large model based on the target strategy for the at least one question to generate a target sample includes: determining a field to be modified based on a target strategy for the at least one problem; Based on the target strategy, the target query statement and target query result are modified with the large model for the field to be modified to generate an intermediate sample; In response to determining that the degree of match between the intermediate sample and the modification requirement indicated by the target policy is less than a second predetermined threshold, returning to perform a determination operation on the field to be modified; In response to determining that the degree of matching between the intermediate sample and the modification requirement indicated by the target policy is greater than or equal to the second predetermined threshold, the intermediate sample is determined to be the target sample.

10. The method according to any one of claims 1 to 9, wherein The method further comprises: The initial query statement in the abnormal query result is enhanced to generate the target query statement.

11. The method according to claim 10, wherein: The step of enhancing the content of the initial query statement in the abnormal query result to generate the target query statement includes: Determining query elements in the initial query statement; and According to the query elements, the initial query statement is enhanced using the large model to generate the target query statement.

12. The method according to claim 11, wherein The query elements include query scenarios; The step of enhancing the content of the initial query statement using the large model according to the query elements to generate the target query statement includes: According to the query scenario, the query scenario of the initial query statement is expanded using the large model to generate the target query statement.

13. The method according to claim 11 or 12, wherein: The query elements also include: query roles; The step of enhancing the content of the initial query statement using the large model according to the query elements to generate the target query statement includes: According to the query role, the query role of the initial query statement is expanded using the large model to generate the target query statement.

14. The method according to any one of claims 11 to 13, wherein The query elements also include: query dimensions; The step of enhancing the content of the initial query statement using the large model according to the query elements to generate the target query statement includes: According to the query dimension, the query dimension of the initial query statement is expanded using the large model to generate the target query statement.

15. The method according to any one of claims 11 to 14, wherein The query elements also include: query field; The step of enhancing the content of the initial query statement using the large model according to the query elements to generate the target query statement includes: According to the query domain, the query domain of the initial query statement is expanded using the large model to generate the target query statement.

16. The method according to any one of claims 1 to 15, further comprising: Based on the at least one problem, the target sample is enhanced using a large model to generate an augmented sample, so as to optimize the retrieval model using the target sample and the augmented sample.

17. The method according to claim 16, wherein The method of performing sample enhancement on the target sample using the large model to generate an augmented sample based on the at least one problem includes: In response to determining that the at least one question indicates a word order anomaly, the word order of the target sample is adjusted using the large model to generate the augmented sample; wherein the word order of the content expressed by the augmented sample is different from that of the target sample and the semantics are similar.

18. The method according to claim 16 or 17, wherein The method of performing sample enhancement on the target sample using the large model to generate an augmented sample based on the at least one problem includes: In response to determining that the at least one question indicates a matching anomaly, the target sample is entity-adjusted using the large model to generate the augmented sample; wherein the augmented sample and the target sample express content in a different manner but have similar semantics; the matching anomaly includes at least one of the following: entity matching anomaly and keyword matching anomaly.

19. A sample generation device based on a large model, comprising: A receiving module, configured to receive at least one question selected by a subject on a question analysis interface; The at least one problem is obtained by analyzing the cause of an abnormal query result output by the retrieval model; as well as a generation module, configured to modify a target query statement and a target query result using a large model based on the at least one question to generate a target sample, so as to optimize the retrieval model for the at least one question based on the target sample; Wherein, the target query statement is associated with the initial query statement in the abnormal query result; The target query result is obtained by querying the target query statement using the retrieval model.

20. An intelligent agent for sample generation, comprising: An input module, configured to receive at least one question selected by a subject on a question analysis interface; a processing module, configured to determine a target task based on the at least one question received by the input module, determine a target macro model based on the target task, and obtain a target sample for optimizing a retrieval model by executing the method according to any one of claims 1 to 19 by calling the target macro model; as well as An output module is used to output the target sample obtained by the processing module.

21. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 19.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-19.

23. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 19.

Citation Information

Patent Citations

  • Semantic understanding method and device, electronic equipment and storage medium

    CN114548110A

  • Data enhancement method and device, electronic equipment and storage medium

    CN119293160A

  • Knowledge base information processing method and device, storage medium and computer equipment

    CN119396991A

  • Domain knowledge fusion data enhancement method for text classification

    CN119646223A