Domain data recall method, domain model training method, domain question and answer method, domain model information processing method, cloud training platform, computing device and computer readable storage medium

By generating semantically similar target key information in the information processing model, the accuracy and comprehensiveness issues of vertical field data recall are solved, and more efficient data recall effects are achieved.

CN120687542APending Publication Date: 2025-09-23HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410318244.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies have difficulty in ensuring the accuracy and comprehensiveness of data recall in vertical fields. Expert data recall methods are inefficient and highly subjective, while traditional methods lack semantic relevance.

Method used

By obtaining the initial key information of the target domain, the information processing model is used to generate target key information in the target domain that is semantically similar to the initial key information under the prompt of the generated prompt information, and the target domain data is recalled based on the initial and target key information.

Benefits of technology

The accuracy and comprehensiveness of data recall in the target field are improved. The generated target key information is used as index information to provide more accurate and comprehensive data recall, thereby improving the quality of the recall results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687542A_ABST
    Figure CN120687542A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a domain data recall method, a domain model training method, a domain question and answer method, an information processing method of a domain model, a cloud training platform, computing equipment and a computer readable storage medium, and is applied to the technical field of machine learning. The domain data recall method comprises the following steps: acquiring initial key information of a target domain; inputting the initial key information into an information processing model, and generating target key information which is similar to the initial key information in semantics in a target field under the prompt of the generated prompt information; and recalling the target domain data based on the initial key information and the target key information. Based on a small amount of initial key information, the target key information with semantics similar to that of the initial key information in the target field is generated under the prompt of the generation prompt information, data recall is completed based on the initial key information and the target key information, and the accuracy and comprehensiveness of data recall in the target field are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of machine learning technology, and in particular to a domain data recall method, a domain model training method, a domain question-answering method, an information processing method for domain models, a cloud training platform, a computing device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the development of computer information technology, the scale of data is becoming larger and larger. How to recall vertical field data from massive data has become a problem to be solved.

[0003] At present, expert data recall methods represented by knowledge graphs and traditional data recall methods represented by TF-IDF (Term Frequency-Inverse Document Frequency) algorithms can achieve high-accuracy data recall based on key information.

[0004] However, expert data recall methods, which construct knowledge graphs, are inefficient and highly subjective, making it difficult to guarantee the accuracy and comprehensiveness of recall results within specific vertical domains. Traditional data recall methods, while quantifying key information to a certain extent, lack semantic relevance and are limited by the available key information, making it equally difficult to guarantee the accuracy and comprehensiveness of recall results within specific vertical domains. Therefore, a domain data retrieval method is urgently needed. Summary of the Invention

[0005] In view of this, embodiments of this specification provide a domain data recall method. One or more embodiments of this specification also involve a domain model training method, a domain question-answering method, a domain model information processing method, a cloud training platform, a domain data recall device, a domain model training device, a domain question-answering device, a domain model information processing device, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.

[0006] According to a first aspect of an embodiment of this specification, a domain data recall method is provided, comprising:

[0007] Obtain initial key information in the target area;

[0008] Inputting the initial key information into the information processing model, and generating target key information in the target domain that is semantically similar to the initial key information under the prompt of the generated prompt information;

[0009] Recall target domain data based on initial key information and target key information.

[0010] According to a second aspect of the embodiments of this specification, a domain model training method is provided, including:

[0011] Obtain a model training request for a target domain, where the model training request includes initial key information of the target domain;

[0012] Inputting the initial key information into the information processing model, and generating target key information in the target domain that is semantically similar to the initial key information under the prompt of the generated prompt information;

[0013] Recall target domain data based on initial key information and target key information;

[0014] Build a training dataset based on target domain data;

[0015] Based on the training data set, the initial domain model is trained to obtain the target domain model corresponding to the target domain.

[0016] According to a third aspect of the embodiments of this specification, a domain question answering method is provided, including:

[0017] Obtaining a target domain question text and a target domain question-answering model corresponding to the target domain, wherein the target domain question-answering model is obtained by training an initial domain question-answering model based on a training data set, the training data set is constructed based on the target domain question-answering text, the target domain question-answering text is recalled based on the initial keywords of the target domain and the target keywords of the target domain, and the target keywords are keywords in the target domain that are semantically similar to the initial keywords and are generated by inputting the initial keywords into the information processing model and under the prompt of generating prompt information;

[0018] Input the question text into the target domain question answering model to obtain the answer text;

[0019] Feedback the reply text to the front-end user.

[0020] According to a fourth aspect of the embodiments of this specification, a method for processing information of a domain model is provided, which is applied to a cloud training platform, including:

[0021] receiving a task generation request for a target domain sent by a terminal device, wherein the task generation request includes request information;

[0022] Based on the request information, a target domain model applied to the target domain is obtained, wherein the target domain question-answering model is obtained by training the initial domain model based on the training dataset, the training dataset is constructed based on the target domain data, the target domain data is recalled based on the initial key information of the target domain and the target key information of the target domain, and the target key information is key information in the target domain that is semantically similar to the initial key information and is generated under the prompt of the generated prompt information by inputting the initial key information into the information processing model;

[0023] Based on the target domain model, task information is generated, wherein the task information is used by the terminal device to perform the target domain task.

[0024] According to a fifth aspect of the embodiments of this specification, there is provided a cloud training platform, comprising a request interface and a response unit;

[0025] A request interface, configured to receive a task generation request for a target domain sent by a terminal device, wherein the task generation request includes request information;

[0026] The response unit is used to obtain a target domain model applied to the target domain based on the request information, wherein the target domain question-answering model is obtained by training the initial domain model based on the training data set, the training data set is constructed based on the target domain data, the target domain data is based on the initial key information of the target domain and the target key information of the target domain is recalled, and the target key information is the key information in the target domain that is semantically similar to the initial key information and is generated under the prompt of generating prompt information by inputting the initial key information into the information processing model; based on the target domain model, task information is generated, wherein the task information is used for the terminal device to perform the target domain task.

[0027] According to a sixth aspect of the embodiments of this specification, a domain data recall device is provided, including:

[0028] A first acquisition module is configured to acquire initial key information of a target domain;

[0029] a first generating module configured to input the initial key information into the information processing model and, under the prompt of the generation prompt information, generate target key information in the target domain that is semantically similar to the initial key information;

[0030] The first recall module is configured to recall target domain data based on the initial key information and the target key information.

[0031] According to a seventh aspect of the embodiments of this specification, a domain model training device is provided, including:

[0032] A second acquisition module is configured to acquire a model training request of a target domain, wherein the model training request includes initial key information of the target domain;

[0033] a second generating module configured to input the initial key information into the information processing model and, under the prompt of the generation prompt information, generate target key information in the target domain that is semantically similar to the initial key information;

[0034] a second recall module configured to recall target domain data based on the initial key information and the target key information;

[0035] The second building module is configured to build a training dataset based on the target domain data;

[0036] The second training module is configured to train the initial domain model based on the training data set to obtain a target domain model corresponding to the target domain.

[0037] According to an eighth aspect of the embodiments of this specification, a domain question-answering device is provided, including:

[0038] a text model acquisition module configured to acquire a question text of a target domain and a target domain question-answering model corresponding to the target domain, wherein the target domain question-answering model is obtained by training an initial domain question-answering model based on a training data set, the training data set is constructed based on the question-answering text of the target domain, the question-answering text of the target domain is recalled based on the initial keywords of the target domain and the target keywords of the target domain, the target keywords are keywords in the target domain that are semantically similar to the initial keywords and are generated by inputting the initial keywords into the information processing model and under the prompt of generating prompt information;

[0039] A question answering module is configured to input a question text into a target domain question answering model and obtain an answer text;

[0040] The reply feedback module is configured to feed back the reply text to the front-end user.

[0041] According to a ninth aspect of the embodiments of this specification, there is provided a domain model information processing device, which is applied to a cloud training platform, including:

[0042] a request receiving module configured to receive a task generation request for a target domain sent by a terminal device, wherein the task generation request includes request information;

[0043] a model acquisition module configured to acquire a target domain model applied to a target domain based on the request information, wherein the target domain question-answering model is obtained by training an initial domain model based on a training data set, the training data set is constructed based on target domain data, the target domain data is recalled based on initial key information of the target domain and target key information of the target domain, and the target key information is key information in the target domain that is semantically similar to the initial key information and is generated by inputting the initial key information into the information processing model under the prompt of generating prompt information;

[0044] The information generation module is configured to generate task information based on the target domain model, wherein the task information is used for the terminal device to perform the target domain task.

[0045] According to a tenth aspect of an embodiment of this specification, a computing device is provided, including:

[0046] memory and processor;

[0047] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above method are implemented.

[0048] According to an eleventh aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and the steps of the above method are implemented when the computer program / instruction is executed by a processor.

[0049] According to a twelfth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.

[0050] In one embodiment of this specification, initial key information for a target domain is obtained; the initial key information is input into an information processing model, and target key information in the target domain that is semantically similar to the initial key information is generated, prompted by generation prompt information; and target domain data is recalled based on the initial key information and the target key information. Based on a small amount of initial key information, target key information in the target domain that is semantically similar to the initial key information is generated, prompted by generation prompt information, and data recall is completed based on the initial key information and the target key information, resulting in more accurate and comprehensive target domain data in the target domain, thereby improving the accuracy and comprehensiveness of target domain data recall. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a flowchart of a field data recall method provided by one embodiment of this specification;

[0052] Figure 2This is a schematic diagram of extracting and generating key information in a domain data recall method provided in one embodiment of this specification;

[0053] Figure 3 This is a schematic diagram of target domain data recall in a domain data recall method provided by one embodiment of this specification;

[0054] Figure 4 This is a flowchart of constructing a training data set in a domain data recall method provided in one embodiment of this specification;

[0055] Figure 5 This is a flowchart of a domain model training method provided by one embodiment of this specification;

[0056] Figure 6 This is a flowchart of a domain question-answering method provided by one embodiment of this specification;

[0057] Figure 7 This is a flowchart of a method for processing information of a domain model provided by one embodiment of this specification;

[0058] Figure 8 This is a front-end schematic diagram of a domain question-answering method applied to the financial field provided by an embodiment of this specification;

[0059] Figure 9 This is a schematic diagram of the structure of a cloud training platform provided by one embodiment of this specification;

[0060] Figure 10 This is a structural diagram of a field data recall device provided by an embodiment of this specification;

[0061] Figure 11 This is a schematic diagram of the structure of a domain model training device provided by one embodiment of this specification;

[0062] Figure 12 This is a schematic diagram of the structure of a domain question-answering device provided by one embodiment of this specification;

[0063] Figure 13 This is a schematic diagram of the structure of an information processing device for a domain model provided by an embodiment of this specification;

[0064] Figure 14 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION

[0065] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0066] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0067] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0068] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0069] In one or more embodiments of this specification, a large model refers to a machine learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model. By pre-training a large model with large-scale unlabeled corpus, a pre-trained model with more than 100 million parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large language model (LLM) and a multi-modal pre-training model.

[0070] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as information-based sentiment classification, information summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0071] First, the terms involved in one or more embodiments of this specification are explained.

[0072] Machine Learning: A computer science technique that studies how computer systems can use empirical data to improve their performance, perform tasks, or make decisions without being explicitly programmed. In machine learning, algorithms "learn" patterns, general rules, and general principles by studying and analyzing large amounts of input data. Based on these learnings, they build predictive models or behavioral strategies.

[0073] Deep Learning: A branch of machine learning, it is an algorithm that uses artificial neural networks as its architecture to represent and learn data.

[0074] Domain Adaptation: This generally refers to the process of adapting a general large model pre-trained on a large-scale dataset to the language characteristics and task requirements of a specific vertical field (such as healthcare, law, finance, etc.) through further pre-training or fine-tuning.

[0075] Vector Retrieval: An information retrieval method based on vector space that uses machine learning techniques to represent information (such as text) as mathematical vectors, and then retrieves and recalls relevant information by calculating the similarity between vectors.

[0076] Key-Phrase Extraction: A text analysis process that involves identifying and extracting words or phrases that summarize the content and theme of a document. These keywords are usually words that reflect the core concept, topic, or most important information in the document.

[0077] Deep Self-Attention Model (Transformer Model): A deep learning model architecture based on the attention mechanism for processing sequential data such as natural language.

[0078] Bidirectional Encoder Representations from Transformers (BERT): A special Transformer model trained using a bidirectional Transformer encoder and large-scale unlabeled text data.

[0079] Pre-training: refers to the pre-training stage conducted on large-scale pre-training data, which can be used to generate processing results, can also be used for fine-tuning, or directly applied to specific tasks in various fields.

[0080] Fine-tuning: refers to the stage of adjusting and optimizing the pre-trained model on a specific task to improve the performance of the model on that task.

[0081] Transfer Learning: A machine learning method that allows the knowledge and patterns learned by a model on one task to be applied to another related but not identical task.

[0082] Reinforcement Learning: A machine learning method that enables a model to learn how to make optimal decisions through repeated trial and error and feedback.

[0083] Domain knowledge: refers to the process of adding pre-training data from a specific field during the pre-training phase, so that the model can better understand the data characteristics of the industry.

[0084] In this specification, a domain data recall method is provided. This specification also involves a domain model training method, a domain question and answer method, a domain model information processing method, a cloud training platform, a domain data recall device, a domain model training device, a domain question and answer device, a domain model information processing device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.

[0085] See also Figure 1 , Figure 1 A flowchart of a domain data recall method provided by an embodiment of this specification is shown, including the following specific steps:

[0086] Step 102: Obtain initial key information of the target domain.

[0087] The embodiments of this specification are applied to applications, websites or platforms with data recall functions. For example, a search engine website can realize accurate and comprehensive data recall in different vertical fields. For example, an application or website deployed with a large model, or an application or website that calls a large model through an application programming interface (Application Programming Interface, referred to as API), introduces "external knowledge" for the data processing of the large model through accurate and comprehensive data recall in vertical fields. For example, a training platform for training deep learning models, by recalling data in vertical fields, constructs a training data set to train machine learning models applied to the target field.

[0088] A domain is a specific area of ​​knowledge, such as science, finance, or e-commerce. Domains can have multiple levels. The more detailed the domain level, the more accurate the recall data will be in line with the recall requirements.

[0089] Key information is information that characterizes the domain characteristics of a domain. Key information is used as index information for data recall. Key information is information of at least one modality, including but not limited to: text, image, audio, video, and structured information. For example, key information is keywords, key words, and key sentences of text modality. For another example, key information is partial images and index images of image modality. For another example, key information is song clips of audio modality. For another example, key information is key frames (for example, I frames) and non-key frames (for example, B frames and P frames) of video modality. For another example, key information is row numbers, column numbers, row names, and column names of structured information modality.

[0090] The target domain is the knowledge domain of the data to be recalled. The target domain limits the vertical domain of the recalled target domain data to a certain extent. For example, the target domain can be the finance field or the mathematics field.

[0091] The initial key information of the target domain is the key information that characterizes the domain characteristics of the target domain. The initial key information is not only used as index information for recalling data in the target domain, but also used as seed information for generating the target initial information in step 104. For example, if the target domain is the financial domain, the initial key information of the target domain is "financial management environment", "profit distribution", "fund investment" and "sales activities". For another example, if the target domain is the mathematical domain, the initial key information of the target domain is "Fourier transform" and "Cauchy inequality". The target domain and the initial key information of the target domain constitute multi-level key information. The target domain can be understood as the first-level key information, and the initial key information of the target domain is the second-level key information.

[0092] The initial key information can be obtained directly. Taking the front-end user interaction as an example, the initial key information of the target field input by the user is received, and the initial key information of multiple fields is provided for the user to select to obtain the initial key information of the target field. Alternatively, the field information of the target field input by the user is received, and the initial key information of the target field is determined based on the field information. The initial key information can also be extracted from the initial field data, which is not limited here.

[0093] Obtaining the initial key information of the target domain provides seed information for the subsequent generation of target key information and index information for the subsequent recall of target domain data.

[0094] Step 104: Input the initial key information into the information processing model, and generate target key information in the target domain that is semantically similar to the initial key information under the prompt of generating prompt information.

[0095] The information processing model is a deep learning model that has the function of understanding input information and performing corresponding operations. The information processing model can perform corresponding operations under the prompt of prompt information. Information processing models include but are not limited to: Transformer model, BERT model and large model. For example, after inputting the seed information of key information, the information processing model uses its inherent knowledge to understand its semantic features under the prompt of generating prompt information and generate key information with similar semantics. For another example, after inputting domain data, the information processing model uses its inherent knowledge to understand its semantic features under the prompt of extracting prompt information and extract key information that conforms to the domain knowledge of the target domain from the input domain data.

[0096] The generated prompt information is the operation description information used to prompt the information processing model to perform the key information generation operation. The generated prompt information can be a pre-built template prompt information, or it can be adaptively generated based on the domain information of the target domain, which is not limited here. Generally, the generated prompt information is composed of a fixed paradigm of instruction information, input information (initial key information) and output information (target key information). For example, a generated prompt information Prompt_Gen is:

[0097] {

[0098] Instructions: You are an expert in the {} field. You need to generate similar keywords in the field based on the provided keywords. Please add line breaks after the generated keywords to separate them.

[0099] It is required to generate keywords within the field, and harmful keywords or complete sentences cannot be generated.

[0100] Input information Input (initial key information):

[0101] {}

[0102] Output information Output (target key information):

[0103] }

[0104] Target key information is key information generated by the information processing model that characterizes the domain characteristics of the target domain. The target key information is semantically similar to the initial key information and enriches the index information for data recall in the target domain. Semantic similarity is not limited to textual semantic similarity but also includes audio semantic similarity, visual semantic similarity (image semantics and video semantics), and structural semantic similarity. This is determined based on the data modality of the key information and the domain data. For example, if the target domain is finance, the initial key information is "financial management environment," "profit distribution," "fund investment," and "sales activities," and the generated target key information is "balance sheet," "portfolio management," and "marketing strategy." For another example, if the target domain is mathematics, the initial key information is "Fourier transform" and "Cauchy inequality," and the generated target key information is "Fourier series," "Laplace transform," "Z transform," and "Cauchy integral theorem."

[0105] Input the initial key information into the information processing model, and under the prompt of generating prompt information, generate target key information in the target domain that is semantically similar to the initial key information. An optional method is: input the initial key information into the information processing model, and under the prompt of generating prompt information, generate target key information in the target domain that is semantically similar to the initial key information based on the semantic features of the initial key information.

[0106] It should be noted that since the target key information and the initial key information are semantically similar, the target key information has higher accuracy than the expert data recall method and the traditional data recall method. Since the target key information is generated by the information processing model, the initial key information and the target key information as index information are more comprehensive than the expert data recall method and the traditional data recall method.

[0107] The initial key information is input into the information processing model, and under the prompt of the generated prompt information, the target key information in the target domain with similar semantics to the initial key information is generated. Based on a small amount of initial key information, the target key information in the target domain with similar semantics to the initial key information is generated under the prompt of the generated prompt information, providing more accurate and comprehensive index information for subsequent target domain data recall.

[0108] Step 106: Recall target domain data based on the initial key information and the target key information.

[0109] Domain data is data generated within a specific knowledge domain. Domain data typically includes domain features such as professional terminology, rules, patterns, and examples in the knowledge domain. Domain data is data of at least one modality, including but not limited to text, images, audio, video, and structured data. For example, domain data is text-based papers, documents, contracts, meeting minutes, and textbooks. Another example is image-based medical images, trademarks, and illustrations in instructions. Another example is audio-based conference recordings, songs, and course recordings. Another example is video-based conference recordings, course recordings, and song music videos. Another example is structured data such as price trends, drug use records, and learning behavior logs.

[0110] The target domain data is the domain data obtained by recalling the initial key information and the target key information. The target domain data is data of at least one modality, including but not limited to: text, image, audio, video and structured data. It should be noted that the target domain data is directly recalled, and it cannot be ensured that the target domain data includes the domain data of the target domain. The target domain data may only include the domain data of the target domain, or may include the domain data of other domains in addition to the target domain, or may include the domain data of the target domain and the domain data of other domains. Therefore, the target domain is a vertical domain that limits the recalled target domain data to a certain extent. Based on the initial key information and the target key information, the target domain data is recalled. An optional method is to perform vector recall based on the initial key information and the target key information to obtain the target domain data.

[0111] In the embodiments of this specification, based on a small amount of initial key information, target key information in the target domain that is semantically similar to the initial key information is generated under the prompt of generated prompt information, and data recall is completed based on the initial key information and the target key information, thereby obtaining more accurate and comprehensive target domain data in the target domain, thereby improving the accuracy and comprehensiveness of target domain data recall.

[0112] In an optional embodiment of this specification, step 102 includes the following specific steps:

[0113] Get initial domain data;

[0114] Extract the initial key information of the target domain from the initial domain data.

[0115] Initial domain data is data generated within the target domain. It contains initial key information about the target domain and is data of at least one modality, including but not limited to text, images, audio, video, and structured data. Initial domain data necessarily includes domain data within the target domain; that is, it is strictly limited to domain data within the vertical domain of the target domain.

[0116] The initial domain data can be directly obtained. Taking front-end user interaction as an example, the initial domain data input by the user is received, or the domain information of the target domain input by the user is received, and the initial domain data is determined based on the domain information. This is not limited here.

[0117] Extracting the initial key information of the target domain from the initial domain data. One optional method is to extract the initial key information of the target domain from the initial domain data based on the information features of the initial domain data, such as the TF-IDF algorithm. Another optional method is to extract the initial key information of the target domain from the initial domain data based on the domain knowledge of the target domain, such as a deep learning algorithm.

[0118] In the embodiments of this specification, initial key information representing domain characteristics of the target domain is extracted from the acquired initial domain data, providing seed information for subsequent generation of target key information and index information for subsequent recall of target domain data.

[0119] In an optional embodiment of the present specification, extracting initial key information of the target domain from the initial domain data includes the following specific steps:

[0120] The initial domain data is input into the information processing model, and under the prompt of the extraction prompt information, the initial key information of the target domain is extracted from the initial domain data, wherein the extraction prompt information is used to prompt the information processing model to extract key information that conforms to the domain knowledge of the target domain.

[0121] The extraction prompt information is the operation description information used to prompt the information processing model to perform the key information extraction operation. The extraction prompt information can be a pre-built template prompt information, or it can be adaptively generated based on the domain information of the target domain, which is not limited here. Generally, the extraction prompt information consists of a fixed paradigm of instruction information, input data (initial domain data) and output information (initial key information). For example, an extraction prompt information Prompt_Xtr is:

[0122] {

[0123] Instruction: You are an expert in the {} field. You need to extract keywords from the text. Please add line breaks after the extracted keywords to separate them.

[0124] What is required is to extract keywords instead of a sentence, and harmful keywords cannot be extracted.

[0125] Input data Input (initial field data):

[0126] {}

[0127] Output information Output (initial key information):

[0128] }

[0129] Domain knowledge in the target domain is an abstract representation of key information attributes in the target domain. This knowledge determines whether the model can accurately understand and extract key information from the target domain. Domain knowledge in the target domain is the inherent knowledge pre-learned by the information processing model and serves as a priori information for extracting key information. For example, if the target domain is science, domain knowledge in the target domain would include axioms, theorems, corollaries, and standardized terminology.

[0130] Figure 2 FIG. 1 shows a schematic diagram of extracting and generating key information in a domain data recall method provided by an embodiment of this specification, such as Figure 2 As shown:

[0131] Obtain initial domain data. Input the initial domain data into the information processing model. Following the prompts for extracting information, extract the initial key information for the target domain: "Financial Management Environment," "Profit Distribution," "Fund Investment," and "Sales Activities." Input the initial key information into the information processing model. Following the prompts for generating information, generate the target key information: "Balance Sheet," "Portfolio Management," and "Marketing Strategy." This yields the initial and target key information for the target domain, improving the coverage of the target domain's key information as index information and the comprehensiveness of data recall in the target domain.

[0132] In the embodiments of this specification, through the information processing model, which is a deep learning model, the initial key information of the target domain that conforms to the domain knowledge of the target domain is extracted from the initial domain data under the prompt of extraction prompt information, providing more accurate seed information for the subsequent generation of target key information, and providing more accurate index information for the subsequent recall of target domain data, thereby improving the accuracy of target domain data recall.

[0133] In an optional embodiment of this specification, step 106 includes the following specific steps:

[0134] Based on the initial key information and the target key information, target domain data of related domains that are semantically similar to the initial key information and the target key information are recalled, wherein the related domains include the target domain.

[0135] The related domain is the knowledge domain to which the directly recalled target domain data belongs. The related domain includes the target domain. Furthermore, the related domain may include only the target domain, other domains in addition to the target domain, or both the target domain and other domains.

[0136] Based on the initial key information and the target key information, target domain data in related fields that are semantically similar to the initial key information and the target key information are recalled. An optional method is to perform vector recall based on the initial key information and the target key information to recall target domain data in related fields that are semantically similar to the initial key information and the target key information. Semantic similarity is determined based on vector similarity of semantic feature vectors.

[0137] In the embodiments of this specification, based on the semantics of the initial key information and the target key information, data recall is completed in a semantically similar manner, and more accurate target domain data in the relevant domain is obtained, thereby improving the accuracy of target domain data recall.

[0138] In an optional embodiment of the present specification, based on the initial key information and the target key information, recalling target domain data in related fields that are semantically similar to the initial key information and the target key information includes the following specific steps:

[0139] Performing vector encoding on the initial key information, the target key information, and the data in each field stored in the preset database to obtain an initial encoding vector for the initial key information, a target encoding vector for the target key information, and encoding vectors for the data in each field;

[0140] Based on the vector semantic similarity between the initial encoding vector and the target encoding vector and the encoding vectors of the data in each domain, the target domain data in the related domain is recalled.

[0141] The preset database is a database that is pre-built and stores domain data of multiple fields, and is the recall range of the target domain data recall.

[0142] The data in each field are data generated in different knowledge fields that are pre-stored in a preset database. The data in each field usually include the professional terms, rules, patterns, examples and other field features of different knowledge fields. The domain data is data of at least one modality, including but not limited to: text, images, audio, video and structured data.

[0143] The initial encoding vector of the initial key information is a feature encoding vector of the semantic features of the initial key information, which is a high-dimensional feature encoding vector representation. It is achieved by mapping the initial key information into a high-dimensional feature vector space so that information with similar semantic features is closer in the feature vector space.

[0144] The target encoding vector of the target key information is the feature encoding vector of the semantic features of the target key information, which is a high-dimensional feature encoding vector representation. It is achieved by mapping the target key information into a high-dimensional feature vector space so that information with similar semantic features is closer in the feature vector space.

[0145] The encoding vector of each field data is the feature encoding vector of the semantic features of the data in each field, which is a high-dimensional feature encoding vector representation. By mapping the data in each field into a high-dimensional feature vector space, information with similar semantic features is made closer in the feature vector space.

[0146] Vector semantic similarity is a quantitative indicator that measures the degree of similarity between two or more feature coding vectors in the semantic space. In the embodiments of this specification, vector semantic similarity is determined by calculating the cosine similarity, Euclidean distance, Jaccard similarity coefficient and other similarity measurement methods between the initial coding vector, the target coding vector and the coding vectors of each field data to determine the degree of semantic proximity between them. When the similarity between the coding vector of a certain field data and the initial coding vector or the target coding vector exceeds a certain threshold, the data is considered to belong to the target field data of a related field that is semantically similar to the initial key information or the target key information.

[0147] In the embodiments of this specification, semantically similar data is recalled through semantic feature coding vectors, and more accurate target domain data in related fields is recalled from a preset database, further improving the accuracy of target domain data recall.

[0148] In an optional embodiment of the present specification, the target domain data of the related field includes first domain data of the target field and second domain data of the reference field, and the reference field is other fields in the related field except the target field;

[0149] After recalling target domain data in related fields that are semantically similar to the initial key information and the target key information based on the initial key information and the target key information, the following specific steps are also included:

[0150] According to a preset domain data ratio, the first domain data of the target domain and the second domain data of the reference domain are screened to obtain the screened target domain data, wherein the domain data ratio is determined based on the weights pre-set for the target domain and the reference domain.

[0151] In the embodiments of this specification, the reference field refers to other fields in the related fields except the target field. The reference field may have a certain overlapping or complementary relationship with the target field.

[0152] The first domain data of the target domain is the data generated in the target domain in the directly recalled target domain data. The first domain data is data of at least one modality, including but not limited to: text, image, audio, video and structured data.

[0153] The second domain data of the reference domain is the data generated in the reference domain in the directly recalled target domain data, and the first domain data is data of at least one modality, including but not limited to: text, image, audio, video and structured data.

[0154] The preset domain data ratio is the ratio of the data volume of the target domain and the reference domain, which is used to control the ratio of the first domain data and the second domain data in the directly recalled target domain data. The domain data ratio is determined based on the weights set in advance for the target domain and the reference domain. The weights are determined based on the assessment of the importance of data in different domains and the needs of the application scenario. If the data in the target domain is considered more critical, the ratio of the first domain data will be higher; if the cross-domain cross-influence is also taken into account, the ratio of the second domain data will be appropriately increased. For example, during the training process of the domain model in the target domain, it is necessary to construct a training data set based on the directly recalled target domain data. If the training data set only contains first domain data, it may cause the trained model to overfit. Therefore, the domain data ratio will be screened according to 70% first domain data and 30% second domain data.

[0155] Since the target domain data is directly recalled, it is necessary to determine which target domain data are first domain data and which target domain data are second domain data before screening. The specific steps for determining this include:

[0156] Based on the domain information of the target domain, first domain data of the target domain and second domain data of the reference domain are determined from the target domain data of the related domain. The domain information is an abstract representation of the domain. This can be determined based on semantic similarity between the domain information and the target domain data, or by matching key information of the target domain data based on the domain information. It can also be determined by using a classification model to classify the target domain data based on the domain information to determine the first domain data and the second domain data. Alternatively, the first domain data and the second domain data can be determined by querying a knowledge graph based on the domain information of the target domain, without limitation.

[0157] In the embodiment of this specification, the first domain data of the target domain and the second domain data of the reference domain are screened according to a preset domain data ratio, thereby improving the accuracy of recalling the target domain data and increasing the diversity of domain distribution.

[0158] In an optional embodiment of the present specification, the target domain data includes target domain data of multiple different data types; after step 106, the following specific steps are also included:

[0159] The target domain data is screened according to a preset data type distribution to obtain the screened target domain data, wherein the data type distribution is determined based on weights pre-set for a plurality of different data types.

[0160] The data type refers to the data form or structure type of the target domain data. Data types include, but are not limited to, language type (e.g., Chinese, English, etc.), text length (e.g., short text, long text, abstract, etc.), and supervised fine-tuning task type (e.g., multi-round dialogue tasks, mathematical reasoning tasks, formatted output tasks, instruction-following tasks, etc.). Regarding language type: For example, data can be divided into different types based on text length, such as news headlines, brief descriptions, and full-text reports. Regarding supervised fine-tuning task types: In intelligent question-answering systems, these tasks may involve multi-round dialogue-based answers to user questions, mathematical formula reasoning to solve financial calculation problems, and report generation according to specific template formats.

[0161] The preset data type distribution is a pre-set data type distribution parameter for the target domain data, which is used to control the data volume distribution of each data type in the directly recalled target domain data. The data type distribution is determined based on the weights pre-set for a variety of different data types. The weights are determined based on the assessment of the importance of different data types and the needs of the application scenario. For example, during the training of the domain model of the target domain, it is necessary to construct a training data set based on the directly recalled target domain data. If the training data set only contains target domain data of a certain data type, it may cause the trained model to overfit. Therefore, it is necessary to combine the degree of influence of different data types on the model capabilities to perform diversity screening on the target domain data.

[0162] In the embodiments of this specification, target domain data is screened according to a preset data type distribution, thereby improving the accuracy of target domain data recall and increasing the diversity of data type distribution.

[0163] In an optional embodiment of this specification, after step 106, the following specific steps are further included:

[0164] Build a training dataset based on target domain data;

[0165] Based on the training data set, the initial domain model is trained to obtain the target domain model applied to the target domain.

[0166] The training dataset is a dataset of sample data used to train the initial domain model. This sample data can be labeled or unlabeled. The training dataset is constructed based on the recalled target domain data and is used to enable the initial domain model (a deep learning model) to learn domain knowledge from the target domain. For example, when training a question-answering model in the financial domain, the training dataset includes sample financial question texts and their corresponding labeled answers. It may also include structured financial statements, economic indicator data, and so on.

[0167] The initial domain model is a general deep learning model to be trained for the target domain. The initial domain model has basic learning capabilities and architecture, but has not yet been domain-tuned and trained based on specific target domain data, and cannot specifically perform domain tasks in the target domain.

[0168] The target domain model is a deep learning model obtained through domain tuning training. It has been targeted at the characteristics and needs of the domain. Compared with the initial domain model, it has more accurate understanding, analysis and prediction capabilities of the target domain data, and can perform domain tasks in the target domain more specifically.

[0169] Based on the target domain data, a training dataset is constructed. The construction process includes at least one of the following: data preprocessing (data cleaning, labeling), data splitting (randomly dividing the target domain data into training set, validation set and test set) and format conversion and standardization.

[0170] Based on the training data set, the initial domain model is trained to obtain the target domain model applied to the target domain. One optional method is to fine-tune the initial domain model based on the training data set to obtain the target domain model applied to the target domain; one optional method is to perform transfer learning on the initial domain model based on the training data set to obtain the target domain model applied to the target domain; one optional method is to perform reinforcement learning on the initial domain model based on the training data set to obtain the target domain model applied to the target domain, which is not limited here. The above training is all achieved through forward propagation and backpropagation, wherein the forward propagation is specifically: inputting sample data into the initial domain model to obtain the predicted output, and determining the loss value based on the predicted output, wherein the backpropagation is specifically: based on the loss value, the model parameters of the initial domain model are updated through the gradient update method. Forward propagation and backpropagation are iteratively performed, and when the preset training end conditions are met, the target domain model applied to the target domain is obtained.

[0171] In the embodiments of this specification, based on a small amount of initial key information, target key information with similar semantics to the initial key information is generated in the target domain under the prompt of generated prompt information, data recall is completed based on the initial key information and the target key information, and more accurate and comprehensive target domain data in the target domain is obtained. Based on the target domain data, a more accurate and comprehensive training data set in the target domain is constructed to complete domain model training, thereby improving the accuracy, generalization ability and robustness of the model in the target domain.

[0172] Figure 3 A schematic diagram of target domain data recall in a domain data recall method provided by one embodiment of this specification is shown, as shown in the figure:

[0173] Obtain initial domain data. Recall target domain data from related domains that are semantically similar to the initial key information and target key information, where the related domains include the target domain and the reference domain, the related domains are primary key information, and the key information in the related domains includes the key information in the target domain and the key information in the reference domain, with the key information being secondary key information. Perform an initial screening of the target domain data in the related domains according to a preset domain data ratio to obtain the target domain data after the initial screening. Perform a second screening of the target domain data in the related domains according to a preset data type distribution to obtain the target domain data after the second screening.

[0174] Figure 4 FIG. 1 shows a flow chart of constructing a training data set in a domain data recall method provided in one embodiment of this specification, as shown in FIG. Figure 4 As shown:

[0175] Acquire initial domain data. Input the initial domain data into the information processing model. Following the extraction prompt, extract the initial key information for the target domain from the initial domain data. Input the initial key information into the information processing model. Following the generation prompt, generate the target key information. Obtain the expanded initial key information and target key information. Recall the target domain data from the pre-set database, implement key information recall and domain data / data type control, and obtain the constructed training dataset.

[0176] See also Figure 5 , Figure 5 A flowchart of a domain model training method provided by an embodiment of this specification is shown, including the following specific steps:

[0177] Step 502: Obtain a model training request for the target domain, wherein the model training request includes initial key information of the target domain.

[0178] Step 504: Input the initial key information into the information processing model, and generate target key information in the target domain that is semantically similar to the initial key information under the prompt of generating prompt information.

[0179] Step 506: Recall target domain data based on the initial key information and the target key information.

[0180] Step 508: Construct a training dataset based on the target domain data.

[0181] Step 510: Based on the training data set, the initial domain model is trained to obtain a target domain model corresponding to the target domain.

[0182] The embodiments of this specification are applied to a website or platform with deep learning model training, for example, a training platform for training deep learning models, which recalls data in a vertical field and constructs a training data set to train a machine learning model applied to the target field.

[0183] The embodiments of this specification are similar to the above Figure 1 The embodiments of the specification are based on the same inventive concept. The specific descriptions of steps 502 to 510 refer to the above embodiments and will not be repeated here.

[0184] In the embodiments of this specification, based on a small amount of initial key information, target key information with similar semantics to the initial key information is generated in the target domain under the prompt of generated prompt information, data recall is completed based on the initial key information and the target key information, and more accurate and comprehensive target domain data in the target domain is obtained. Based on the target domain data, a more accurate and comprehensive training data set in the target domain is constructed to complete domain model training, thereby improving the accuracy, generalization ability and robustness of the model in the target domain.

[0185] See also Figure 6 , Figure 6 A flowchart of a domain question-answering method provided by one embodiment of this specification is shown, including the following specific steps:

[0186] Step 602: Obtain the question text of the target domain and the target domain question-answering model corresponding to the target domain, wherein the target domain question-answering model is obtained by training the initial domain question-answering model based on the training data set, the training data set is constructed based on the question-answering text of the target domain, the question-answering text of the target domain is recalled based on the initial keywords of the target domain and the target keywords of the target domain, and the target keywords are keywords in the target domain that are semantically similar to the initial keywords and are generated by inputting the initial keywords into the information processing model and under the prompt of generating prompt information.

[0187] Step 604: Input the question text into the target domain question-answering model to obtain the answer text.

[0188] Step 606: Feedback the reply text to the front-end user.

[0189] The embodiments of this specification are applied to applications, websites or platforms with domain text question and answer capabilities, for example, an application or website that deploys a large language model, or an application or website that calls a large language model through an application programming interface (Application Programming Interface, referred to as API). The large language model is pre-trained through an accurate and comprehensive training data set in a vertical field, providing domain knowledge for text question and answering of the large language model.

[0190] The question text of the target field is the question text that needs to be answered in the target field. The question text reflects the professional problems or information query needs in the target field.

[0191] The target domain question answering model is a trained deep learning model that is optimized and tuned specifically for a specific target domain, capable of understanding and answering questions within that domain. This model is trained on a training dataset constructed from target domain data and is capable of handling natural language question answering tasks within that domain.

[0192] The question-answer texts in the target domain are sample pairs in the dataset used to train the question-answering model in the target domain, usually including paired questions and answers. These question-answer pairs come from the target domain and are obtained through the recall method based on the initial keywords and target keywords. They represent typical questions in the target domain and their corresponding correct answers.

[0193] The initial keywords of the target domain are key information used to characterize the characteristics of the target domain. They are seed information and index information in the data recall process.

[0194] The target keywords in the target field are key information with similar semantics generated based on the initial keywords under the guidance of an information processing model (such as a Transformer model, BERT, or a large model).

[0195] The response text is the answer text output by the target domain question answering model after the user's question text in the target domain is input into the target domain question answering model. The response text is generated by the target domain question answering model based on the knowledge and patterns it has learned in the target domain. It aims to accurately respond to user questions and improve the quality of question answering and user experience in the target domain.

[0196] For example, a user enters a question in the target domain (finance) into the finance Q&A module of the Big Language Model website: "How do I analyze a company's financial statements to determine its profitability?" The website has pre-trained and deployed a finance Q&A model based on a large dataset of finance Q&A pairs (e.g., "How do I calculate the price-to-earnings ratio?", "How do I assess a company's debt-paying ability?"). The user's question, "How do I analyze a company's financial statements to determine its profitability?", is input into the finance Q&A model. The model leverages its inherent financial knowledge to understand the question and generates a corresponding answer: "When analyzing a company's profitability, you can focus on the following aspects: First, examine the net profit growth rate and gross profit margin in the income statement; second, understand the asset structure, debt level, and return on equity through the balance sheet; and finally, examine the net cash flow generated by operating activities and its alignment with net profit in conjunction with the cash flow statement." The response, "When analyzing a company's profitability, you can focus on...", is rendered on the user interface, allowing users to immediately see a professional answer, thus addressing their need for financial expertise.

[0197] In the embodiments of this specification, based on a small amount of initial key information, target key information with similar semantics to the initial key information is generated in the target domain under the prompt of generated prompt information, data recall is completed based on the initial key information and the target key information, and more accurate and comprehensive target domain data in the target domain is obtained. Based on the target domain data, a more accurate and comprehensive training data set in the target domain is constructed to complete domain model training, thereby improving the accuracy, generalization ability and robustness of the model in the target domain, and utilizing this more accurate, more generalized and more stable target domain question-answering model to complete question-answering tasks in the target domain, thereby improving the quality of domain question-answering and enhancing user experience.

[0198] In an optional embodiment of this specification, after step 606, the following specific steps are further included:

[0199] Receiving a reply feedback message sent by a front-end user, wherein the reply feedback message is a message in which the front-end user provides feedback on the reply text;

[0200] Update the training dataset based on the question text, answer text, and answer feedback message;

[0201] Based on the updated training dataset, the target domain question answering model is trained.

[0202] Feedback messages are generated when a front-end user receives a response and then comments on its content, asks further questions, or makes corrections. Feedback messages directly reflect user satisfaction with the model's Q&A and can include user approval, specific reasons for dissatisfaction, additional information, or corrections.

[0203] For example, after the user obtains the answer to the financial statement analysis question in the financial question and answer module, he sends a reply feedback message: "This answer is very helpful to me, especially the two indicators mentioned: net profit growth rate and return on equity." The system updates the training data set based on this feedback, adds the question and the given answer as high-quality samples to the training data set, and combines the user's positive feedback information to strengthen the model's learning of how to answer such questions.

[0204] In the embodiments of this specification, by providing interactive feedback on the reply text with the front-end user, the model can continuously optimize its knowledge base and learning strategy, thereby improving the accuracy and pertinence of future answers to similar questions.

[0205] In an optional embodiment of this specification, step 606 includes the following specific steps:

[0206] Mark the keywords in the reply text and feed the reply text back to the front-end user;

[0207] Correspondingly, after step 606, the following specific steps are also included:

[0208] Receiving a keyword feedback message sent by a front-end user, wherein the keyword feedback message is a message in which the front-end user provides feedback on a keyword marked in a reply text;

[0209] Recall target domain data based on keyword feedback messages;

[0210] Update the training dataset based on the target domain data;

[0211] Based on the updated training dataset, the target domain question answering model is trained.

[0212] Keyword feedback messages are user responses to keywords marked in the model's response text. These messages may confirm, question, supplement, or correct the keywords. Keyword feedback messages focus on the core knowledge points in the response text, helping the system understand the user's accuracy and acceptance of specific professional terms and concepts.

[0213] For example, after tagging keywords such as "net profit growth rate," "gross profit margin," and "return on equity" in a reply, a user sends a keyword feedback message: "The explanation of 'return on equity' is a bit brief. Can you explain the calculation method in detail?" The system receives this keyword feedback and recalls more relevant field data on "return on equity," such as specific calculation formulas and case studies. It then updates the training dataset based on this new data. Based on this user feedback, the target domain question-answering model is retrained, enabling it to provide more detailed and accurate responses to similar keywords in the future.

[0214] In the embodiments of this specification, by responding to precise feedback at the keyword level, the system can refine and enrich the knowledge structure of the model, improve its expressive ability and teaching effect on specific knowledge points, thereby enhancing the user experience and deepening the professionalism of the model.

[0215] See also Figure 7 , Figure 7 A flowchart of a domain model information processing method provided by one embodiment of this specification is shown. The method is applied to a cloud training platform and includes the following specific steps:

[0216] Step 702: Receive a task generation request for a target domain sent by a terminal device, wherein the task generation request includes request information.

[0217] Step 704: Based on the request information, obtain the target domain model applied to the target domain, wherein the target domain question-answering model is obtained by training the initial domain model based on the training data set, the training data set is constructed based on the target domain data, the target domain data is recalled based on the initial key information of the target domain and the target key information of the target domain, and the target key information is the key information in the target domain that is semantically similar to the initial key information and is generated under the prompt of generating prompt information by inputting the initial key information into the information processing model.

[0218] Step 706: Generate task information based on the target domain model, wherein the task information is used for the terminal device to perform the target domain task.

[0219] The embodiments of this specification are applied to a cloud training platform with information processing capabilities.

[0220] A task generation request is a specifically formatted instruction sent by a terminal device to the cloud training platform, requesting the generation of a task solution or model application instance within a target domain. For example, an online medical consultation application might send a task generation request to the cloud training platform regarding disease diagnosis.

[0221] Request information is the specific parameters and details included in the task generation request, which guides the cloud training platform in selecting, customizing, or invoking the corresponding domain model for processing. Request information includes the target domain task scenario identifier (e.g., "heart disease case diagnosis"), the task model identifier (e.g., "cardiovascular disease diagnosis model v2.0"), or the actual training dataset provided.

[0222] Task information is a specific task execution guide or result calculated based on the received task generation request and the corresponding target domain model. It is usually in the form of structured data and can be directly used by the terminal device to complete the actual task operation. For example, in the medical field, task information may be a condition analysis report and treatment recommendations generated by a diagnostic model based on patient medical records.

[0223] Target domain tasks are specific problems or work requirements in a specific field, requiring expertise and models to solve. For example, in an intelligent customer service system, a target domain task might be answering investment and financial management questions for customers in the financial sector.

[0224] In an optional embodiment of the present specification, the request information includes a task scenario identifier of the target domain task, or a task model identifier; correspondingly, step 704 includes the following specific steps:

[0225] Based on the task scenario identifier, a target scenario template is determined from multiple preset scenario templates, and based on the target scenario template, a target domain model applied to the target domain is searched from a model library, wherein the model library stores domain models of multiple different domains; or, based on the task model identifier, a target domain model applied to the target domain is searched from a model library.

[0226] A task scenario identifier is a label or code that uniquely identifies or describes the specific application scenario in which the task occurs. For example, in the education field, the task scenario identifier might be "Solving a High School Math Problem," helping the platform find the corresponding pre-set scenario template and match it with the appropriate model.

[0227] The task model identifier is a unique identifier specific to a domain model, which is used to directly locate the model in a specific task scenario from the model library.

[0228] Preset scenario templates are standardized, pre-designed templates for different task scenarios. These templates are associated with the domain model of the task scenario and allow for rapid adaptation to different scenarios. For example, in a smart home control scenario, pre-set scenario templates might include "Away Mode" and "Returning Mode." Each template corresponds to a set of device state settings and operation rules.

[0229] The target scenario template is a specific scenario template suitable for the current task requirements, selected from a set of preset scenario templates based on the task scenario identifier. For example, if the received task scenario identifier is "emergency medical rescue," the target scenario template may include steps such as initiating the emergency response process and contacting the nearest hospital, and is integrated with relevant medical emergency domain models.

[0230] The model library stores a collection of trained models from multiple domains, which can be used to handle tasks in a variety of scenarios. The model library supports on-demand search, loading, and deployment of model resources. For example, the model library stores models for various domains, including but not limited to medical diagnosis models, financial risk control models, and educational tutoring models, enabling rapid retrieval and provision of services based on different request information.

[0231] In an optional embodiment of the present specification, the request information includes a training data set; accordingly, step 704 includes the following specific steps:

[0232] Based on the training data set, the initial domain model is trained to obtain the target domain model applied to the target domain.

[0233] For example, the online medical consultation app "Health Assistant" sends a task generation request to the cloud training platform. The request is to perform an intelligent diagnosis of a heart disease case using the specific task model identifier "Heart Disease Case Diagnosis Model v3.0." The request contains a small amount of initial key information, such as the patient's age, gender, blood pressure, and heart rate.

[0234] Obtaining the model based on the task model identifier: If the request information contains a task model identifier, the cloud training platform will directly retrieve and load the "Heart Disease Case Diagnosis Model v3.0" from the model library after receiving the task model identifier. Alternatively, determine the model based on the task scenario identifier and the preset scenario template: If the request information contains a task scenario identifier (such as "Acute Coronary Syndrome Diagnosis Process"), the platform will first match the corresponding target scenario template from multiple preset scenario templates based on this identifier. The template is associated with the "Acute Coronary Syndrome Diagnosis Model". This model is then called from the model library. Alternatively, update the model based on the training dataset: If the request information contains a specific training dataset, such as new heart disease medical record sample data. The platform will use this data to retrain the existing initial heart disease diagnosis model to adapt to the latest disease characteristics, thereby obtaining an updated target domain model for the current target domain.

[0235] The cloud training platform has successfully acquired or updated a target domain model suitable for the current heart disease diagnosis task. The platform uses this model to process patient medical records sent by end devices, analyzing and generating task information that includes a detailed explanation of the condition, possible diagnoses, and corresponding treatment recommendations. This structured task information is then returned to the end device's "health assistant" for display on the user interface to guide doctors or assist users in decision-making.

[0236] In an embodiment of the present specification, after receiving a task generation request in a target domain, target key information in the target domain that is semantically similar to the initial key information is generated based on the request information and a small amount of initial key information under the prompt of generation prompt information. Data recall is completed based on the initial key information and the target key information, and more accurate and comprehensive target domain data in the target domain is obtained. A more accurate and comprehensive training data set in the target domain is constructed based on the target domain data, and domain model training is completed, thereby improving the accuracy, generalization ability and robustness of the model in the target domain. Based on the target domain model of the target domain, task information is generated for the terminal device to perform the target domain task, thereby improving the execution quality of the target domain task and improving the user experience.

[0237] Figure 8 FIG. 1 shows a front-end schematic diagram of a domain question-answering method applied to the financial field provided by an embodiment of this specification, such as Figure 8 As shown:

[0238] The front-end interface has a dialogue display area, an input box, a send control, and an export control. After the user enters the financial question text "How to adjust the company's asset-liability structure to improve financial stability in the current economic environment?" in the input box, he clicks the send control. After the above step 818 is processed, the generated reply text is rendered in the dialogue display area: "In order to enhance financial stability in the current economic environment, it is recommended that you take the following measures: First, reasonably assess the liquidity and risks of existing assets, and dispose of non-core assets in a timely manner to optimize the asset structure; second, pay close attention to the debt maturity structure, and appropriately increase the proportion of long-term liabilities to reduce short-term debt repayment pressure; finally, based on the company's cash flow situation and future operating expectations, carefully consider capital structure adjustments and new investment projects." The user can export the reply text into a file in a specific format by clicking the export control, such as txt format, doc format, pdf format, etc.

[0239] Corresponding to the above method embodiment, this specification also provides a cloud training platform embodiment, Figure 9 FIG1 shows a schematic diagram of the structure of a cloud training platform provided by an embodiment of this specification. Figure 9 As shown, the cloud training platform 900 includes a request interface 910 and a response unit 920;

[0240] The request interface 910 is used to receive a task generation request for a target domain sent by a terminal device, wherein the task generation request includes request information;

[0241] The response unit 920 is used to obtain a target domain model applied to the target domain based on the request information, wherein the target domain question-answering model is obtained by training the initial domain model based on the training data set, the training data set is constructed based on the target domain data, the target domain data is based on the initial key information of the target domain and the target key information of the target domain is recalled, and the target key information is the key information in the target domain that is semantically similar to the initial key information and is generated under the prompt of generating prompt information by inputting the initial key information into the information processing model; based on the target domain model, task information is generated, wherein the task information is used for the terminal device to perform the target domain task.

[0242] In an embodiment of the present specification, after the request interface receives a task generation request in the target domain, the response unit generates target key information in the target domain that is semantically similar to the initial key information based on the request information and a small amount of initial key information, prompted by the generation prompt information. Data recall is completed based on the initial key information and the target key information, and more accurate and comprehensive target domain data in the target domain is obtained. A more accurate and comprehensive training data set in the target domain is constructed based on the target domain data, and domain model training is completed, thereby improving the accuracy, generalization ability and robustness of the model in the target domain. Based on the target domain model of the target domain, task information is generated for the terminal device to perform the target domain task, thereby improving the execution quality of the target domain task and improving the user experience of the cloud training platform.

[0243] In an optional embodiment of the present specification, the cloud training platform 900 further includes a model library, wherein the model library stores domain models of multiple different domains; the request information includes a task scenario identifier of the target domain task, or a task model identifier;

[0244] The response unit 920 is specifically used to determine the target scene template from multiple preset scene templates based on the task scene identifier, and based on the target scene template, search for the target domain model applied to the target domain from the model library; or, based on the task model identifier, search for the target domain model applied to the target domain from the model library.

[0245] The above is a schematic diagram of a cloud training platform according to this embodiment. It should be noted that the technical solution of this cloud training platform and the technical solution of the information processing method of the domain model described above are based on the same concept. For details not described in detail in the technical solution of the cloud training platform, please refer to the description of the technical solution of the information processing method of the domain model described above.

[0246] Corresponding to the above method embodiment, this specification also provides an embodiment of a domain data recall device, Figure 10 FIG. 1 shows a schematic diagram of the structure of a field data recall device provided by an embodiment of this specification. Figure 10 As shown, the device includes:

[0247] A first acquisition module 1002 is configured to acquire initial key information of a target domain;

[0248] The first generating module 1004 is configured to input the initial key information into the information processing model and, under the prompt of the generation prompt information, generate target key information in the target domain that is semantically similar to the initial key information;

[0249] The first recall module 1006 is configured to recall target domain data based on the initial key information and the target key information.

[0250] Optionally, the first acquisition module 1002 is further configured to: acquire initial domain data; and extract initial key information of the target domain from the initial domain data.

[0251] Optionally, the first acquisition module 1004 is further configured to: input the initial domain data into the information processing model, and extract the initial key information of the target domain from the initial domain data under the prompt of the extraction prompt information, wherein the extraction prompt information is used to prompt the information processing model to extract key information that conforms to the domain knowledge of the target domain.

[0252] Optionally, the first recall module 1006 is further configured to: based on the initial key information and the target key information, recall target domain data of related domains that are semantically similar to the initial key information and the target key information, wherein the related domains include the target domain.

[0253] Optionally, the first recall module 1006 is further configured to: perform vector encoding on the initial key information, the target key information and the data in various fields stored in the preset database to obtain the initial encoding vector of the initial key information, the target encoding vector of the target key information and the encoding vector of the data in various fields; based on the vector semantic similarity between the initial encoding vector and the target encoding vector and the encoding vector of the data in various fields, recall the target field data in the relevant fields.

[0254] Optionally, the target domain data of the related field includes first domain data of the target field and second domain data of the reference field, and the reference field is other fields in the related field except the target field; correspondingly, the device also includes: a field screening module, configured to screen the first domain data of the target field and the second domain data of the reference field according to a preset field data ratio to obtain the screened target field data, wherein the field data ratio is determined based on the weights pre-set for the target field and the reference field.

[0255] Optionally, the target domain data includes target domain data of multiple different data types; correspondingly, the device also includes: a type screening module, configured to screen the target domain data according to a preset data type distribution to obtain the screened target domain data, wherein the data type distribution is determined based on weights pre-set for multiple different data types.

[0256] Optionally, the device further includes: a first training module configured to construct a training data set based on target domain data; and train the initial domain model based on the training data set to obtain a target domain model applied to the target domain.

[0257] In the embodiments of this specification, based on a small amount of initial key information, target key information in the target domain that is semantically similar to the initial key information is generated under the prompt of generated prompt information, and data recall is completed based on the initial key information and the target key information, thereby obtaining more accurate and comprehensive target domain data in the target domain, thereby improving the accuracy and comprehensiveness of target domain data recall.

[0258] The above is a schematic diagram of a domain data recall device according to this embodiment. It should be noted that the technical solution of the domain data recall device and the technical solution of the aforementioned domain data recall method are based on the same concept. For details not described in detail in the technical solution of the domain data recall device, please refer to the description of the technical solution of the aforementioned domain data recall method.

[0259] Corresponding to the above method embodiment, this specification also provides a domain model training device embodiment, Figure 11 FIG1 shows a schematic diagram of the structure of a domain model training device provided by an embodiment of this specification. Figure 11 As shown, the device includes:

[0260] A second acquisition module 1102 is configured to acquire a model training request for a target domain, wherein the model training request includes initial key information of the target domain;

[0261] The second generating module 1104 is configured to input the initial key information into the information processing model and, under the prompt of the generation prompt information, generate target key information in the target domain that is semantically similar to the initial key information;

[0262] The second recall module 1106 is configured to recall target domain data based on the initial key information and the target key information;

[0263] The second construction module 1108 is configured to construct a training data set based on the target domain data;

[0264] The second training module 1110 is configured to train the initial domain model based on the training data set to obtain a target domain model corresponding to the target domain.

[0265] In the embodiments of this specification, based on a small amount of initial key information, target key information with similar semantics to the initial key information is generated in the target domain under the prompt of generated prompt information, data recall is completed based on the initial key information and the target key information, and more accurate and comprehensive target domain data in the target domain is obtained. Based on the target domain data, a more accurate and comprehensive training data set in the target domain is constructed to complete domain model training, thereby improving the accuracy, generalization ability and robustness of the model in the target domain.

[0266] The above is a schematic diagram of a domain model training device according to this embodiment. It should be noted that the technical solution of this domain model training device and the technical solution of the aforementioned domain model training method are based on the same concept. For details not described in detail in the technical solution of the domain model training device, please refer to the description of the technical solution of the aforementioned domain model training method.

[0267] Corresponding to the above method embodiment, this specification also provides a domain question-answering device embodiment, Figure 12 FIG1 shows a schematic diagram of the structure of a domain question-answering device provided by an embodiment of this specification. Figure 12 As shown, the device includes:

[0268] The text model acquisition module 1202 is configured to obtain a question text of a target domain and a target domain question-answering model corresponding to the target domain, wherein the target domain question-answering model is obtained by training an initial domain question-answering model based on a training data set, the training data set is constructed based on the question-answering text of the target domain, the question-answering text of the target domain is recalled based on the initial keywords of the target domain and the target keywords of the target domain, the target keywords are keywords in the target domain that are semantically similar to the initial keywords and are generated by inputting the initial keywords into the information processing model and under the prompt of generating prompt information;

[0269] The question answering module 1204 is configured to input the question text into the target domain question answering model to obtain the answer text;

[0270] The reply feedback module 1206 is configured to feed back the reply text to the front-end user.

[0271] Optionally, the device also includes: a text interaction feedback module, configured to receive a reply feedback message sent by a front-end user, wherein the reply feedback message is a message from the front-end user to provide feedback on the reply text; based on the question text, reply text and reply feedback message, updating the training data set; and training the target domain question and answer model based on the updated training data set.

[0272] Optionally, the reply feedback module 1206 is further configured to: mark keywords in the reply text and feed back the reply text to the front-end user;

[0273] Correspondingly, the device also includes: an information interaction feedback module, configured to receive keyword feedback messages sent by front-end users, wherein the keyword feedback messages are messages from front-end users to provide feedback on keywords marked in the reply text; based on the keyword feedback messages, recall target domain data; based on the target domain data, update the training data set; based on the updated training data set, train the target domain question and answer model.

[0274] In the embodiments of this specification, based on a small amount of initial key information, target key information with similar semantics to the initial key information is generated in the target domain under the prompt of generated prompt information, data recall is completed based on the initial key information and the target key information, and more accurate and comprehensive target domain data in the target domain is obtained. Based on the target domain data, a more accurate and comprehensive training data set in the target domain is constructed to complete domain model training, thereby improving the accuracy, generalization ability and robustness of the model in the target domain.

[0275] The above is a schematic diagram of a domain question-answering device according to this embodiment. It should be noted that the technical solution of this domain question-answering device is based on the same concept as the technical solution of the aforementioned domain question-answering method. For details not described in detail in the technical solution of the domain question-answering device, please refer to the description of the technical solution of the aforementioned domain question-answering method.

[0276] Corresponding to the above method embodiment, this specification also provides an information processing device embodiment of a domain model. Figure 13 FIG. 1 shows a schematic diagram of a domain model information processing device provided by an embodiment of this specification. Figure 13 As shown, the device is applied to a cloud training platform and includes:

[0277] The request receiving module 1302 is configured to receive a task generation request of a target domain sent by a terminal device, wherein the task generation request includes request information;

[0278] The model acquisition module 1304 is configured to acquire a target domain model applied to the target domain based on the request information, wherein the target domain question-answering model is obtained by training the initial domain model based on a training data set, the training data set is constructed based on the target domain data, the target domain data is recalled based on the initial key information of the target domain and the target key information of the target domain, and the target key information is key information in the target domain that is semantically similar to the initial key information and is generated under the prompt of generating prompt information by inputting the initial key information into the information processing model;

[0279] The information generation module 1306 is configured to generate task information based on the target domain model, wherein the task information is used for the terminal device to perform the target domain task.

[0280] Optionally, the request information includes a task scenario identifier of the target domain task, or a task model identifier; correspondingly, the model acquisition module 1304 is further configured to: based on the task scenario identifier, determine the target scene template from multiple preset scene templates, and based on the target scene template, search for a target domain model applied to the target domain from a model library, wherein the model library stores domain models of multiple different domains; or, based on the task model identifier, search for a target domain model applied to the target domain from the model library.

[0281] Optionally, the request information includes a training data set; correspondingly, the model acquisition module 1304 is further configured to: train the initial domain model based on the training data set to obtain a target domain model applied to the target domain.

[0282] In the embodiments of this specification, based on a small amount of initial key information, target key information with similar semantics to the initial key information is generated in the target domain under the prompt of generated prompt information, data recall is completed based on the initial key information and the target key information, and more accurate and comprehensive target domain data in the target domain is obtained. Based on the target domain data, a more accurate and comprehensive training data set in the target domain is constructed to complete domain model training, thereby improving the accuracy, generalization ability and robustness of the model in the target domain, and utilizing this more accurate, more generalized and more stable target domain question-answering model to complete question-answering tasks in the target domain, thereby improving the quality of domain question-answering and enhancing user experience.

[0283] The above is a schematic diagram of an information processing device for a domain model according to this embodiment. It should be noted that the technical solution of the information processing device for this domain model and the technical solution of the information processing method for the domain model described above are based on the same concept. For details not described in detail in the technical solution of the information processing device for the domain model, please refer to the description of the technical solution of the information processing method for the domain model described above.

[0284] Figure 14 14 shows a block diagram of a computing device according to an embodiment of the present disclosure. Components of the computing device 1400 include, but are not limited to, a memory 1410 and a processor 1420. The processor 1420 is connected to the memory 1410 via a bus 1430, and a database 1450 is used to store data.

[0285] The computing device 1400 also includes an access device 1440 that enables the computing device 1400 to communicate via one or more networks 1460. Examples of such networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1440 may include one or more of any type of network interface (e.g., a Network Interface Controller (NIC)) whether wired or wireless, such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC).

[0286] In one embodiment of the present specification, the above components of the computing device 1400 and Figure 14 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 14 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0287] Computing device 1400 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1400 can also be a mobile or stationary server.

[0288] Among them, the processor 1420 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-mentioned domain data recall method, domain model training method, domain question and answer method or domain model information processing method.

[0289] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of this computing device is based on the same concept as the technical schemes of the aforementioned domain data recall method, domain model training method, domain question-answering method, and domain model information processing method. For details not described in detail in the technical scheme of the computing device, please refer to the description of the technical schemes of the aforementioned domain data recall method, domain model training method, domain question-answering method, or domain model information processing method.

[0290] One embodiment of this specification also provides a computer-readable storage medium, which stores a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned domain data recall method, domain model training method, domain question-answering method, or domain model information processing method.

[0291] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of this storage medium is based on the same concept as the technical schemes of the aforementioned domain data recall method, domain model training method, domain question-answering method, and domain model information processing method. For details not described in detail in the technical scheme of the storage medium, please refer to the description of the technical schemes of the aforementioned domain data recall method, domain model training method, domain question-answering method, or domain model information processing method.

[0292] An embodiment of this specification also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned domain data recall method, domain model training method, domain question-answering method, or domain model information processing method.

[0293] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of this computer program product is based on the same concept as the technical schemes of the aforementioned domain data recall method, domain model training method, domain question-answering method, and domain model information processing method. For details not described in detail in the technical scheme of the computer program product, please refer to the description of the technical schemes of the aforementioned domain data recall method, domain model training method, domain question-answering method, or domain model information processing method.

[0294] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0295] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0296] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0297] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0298] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A domain data recall method, comprising: Obtain initial key information on the target area; Inputting the initial key information into an information processing model, and generating target key information in the target domain that is semantically similar to the initial key information under the prompt of generating prompt information; Target domain data is recalled based on the initial key information and the target key information.

2. The method according to claim 1, wherein obtaining initial key information of the target domain comprises: Get initial domain data; Extracting initial key information of the target domain from the initial domain data.

3. The method according to claim 2, wherein extracting the initial key information of the target domain from the initial domain data comprises: The initial domain data is input into the information processing model, and under the prompt of extraction prompt information, the initial key information of the target domain is extracted from the initial domain data, wherein the extraction prompt information is used to prompt the information processing model to extract key information that conforms to the domain knowledge of the target domain.

4. The method according to claim 1, wherein the step of recalling target domain data based on the initial key information and the target key information comprises: Based on the initial key information and the target key information, target domain data of related domains that are semantically similar to the initial key information and the target key information are recalled, wherein the related domains include the target domain.

5. The method according to claim 4, wherein the step of recalling target domain data in related domains that are semantically similar to the initial key information and the target key information based on the initial key information and the target key information comprises: Performing vector encoding on the initial key information, the target key information, and the various field data stored in a preset database to obtain an initial encoding vector for the initial key information, a target encoding vector for the target key information, and encoding vectors for the various field data; Based on the vector semantic similarity between the initial encoding vector and the target encoding vector and the encoding vectors of the data in each field, target field data in the relevant field is recalled.

6. The method according to claim 4, wherein the target domain data of the related domain comprises first domain data of the target domain and second domain data of a reference domain, wherein the reference domain is another domain in the related domain except the target domain; After recalling target domain data in related fields that are semantically similar to the initial key information and the target key information based on the initial key information and the target key information, the method further includes: According to a preset domain data ratio, the first domain data of the target domain and the second domain data of the reference domain are screened to obtain the screened target domain data, wherein the domain data ratio is determined based on the weights pre-set for the target domain and the reference domain.

7. The method according to claim 1, wherein the target domain data comprises target domain data of multiple different data types; After recalling the target domain data based on the initial key information and the target key information, the method further includes: The target domain data is screened according to a preset data type distribution to obtain screened target domain data, wherein the data type distribution is determined based on weights pre-set for the multiple different data types.

8. The method according to any one of claims 1 to 7, further comprising, after recalling target domain data based on the initial key information and the target key information: Constructing a training data set based on the target domain data; Based on the training data set, the initial domain model is trained to obtain a target domain model applied to the target domain.

9. A domain model training method, comprising: Obtaining a model training request for a target domain, wherein the model training request includes initial key information of the target domain; Inputting the initial key information into an information processing model, and generating target key information in the target domain that is semantically similar to the initial key information under the prompt of generating prompt information; Recalling target domain data based on the initial key information and the target key information; Constructing a training data set based on the target domain data; Based on the training data set, the initial domain model is trained to obtain a target domain model corresponding to the target domain.

10. A domain question answering method, comprising: Obtaining a question text in a target domain and a target domain question-answering model corresponding to the target domain, wherein the target domain question-answering model is obtained by training an initial domain question-answering model based on a training data set, the training data set is constructed based on the question-answering text in the target domain, the question-answering text in the target domain is recalled based on the initial keywords in the target domain and the target keywords in the target domain, the target keywords being keywords in the target domain that are semantically similar to the initial keywords and are generated by inputting the initial keywords into an information processing model and under the prompt of generating prompt information; Input the question text into the target domain question answering model to obtain a reply text; Feedback the reply text to the front-end user.

11. The method according to claim 10, further comprising, after feeding back the reply text to the front-end user: Receiving a reply feedback message sent by the front-end user, wherein the reply feedback message is a message in which the front-end user provides feedback on the reply text; Updating the training data set based on the question text, the answer text and the answer feedback message; Based on the updated training dataset, the target domain question answering model is trained.

12. The method according to claim 10, wherein feeding back the reply text to the front-end user comprises: Marking the keywords in the reply text and feeding back the reply text to the front-end user; After feeding back the reply text to the front-end user, the method further includes: Receiving a keyword feedback message sent by the front-end user, wherein the keyword feedback message is a message in which the front-end user provides feedback on the keyword marked in the reply text; Recalling target domain data based on the keyword feedback message; Based on the target domain data, updating the training data set; Based on the updated training dataset, the target domain question answering model is trained.

13. A domain model information processing method, applied to a cloud training platform, comprising: receiving a task generation request for a target domain sent by a terminal device, wherein the task generation request includes request information; Based on the request information, a target domain model applied to the target domain is obtained, wherein the target domain question-answering model is obtained by training an initial domain model based on a training data set, the training data set is constructed based on target domain data, the target domain data is recalled based on initial key information of the target domain and target key information of the target domain, and the target key information is key information in the target domain that is semantically similar to the initial key information and is generated by inputting the initial key information into an information processing model and under the prompt of generating prompt information; Based on the target domain model, task information is generated, wherein the task information is used for the terminal device to perform the target domain task.

14. The method according to claim 13, wherein the request information includes a task scenario identifier of the target domain task, or a task model identifier; The acquiring, based on the request information, a target domain model applied to the target domain, includes: Based on the task scenario identifier, a target scenario template is determined from a plurality of preset scenario templates, and based on the target scenario template, a target domain model applied to the target domain is searched from a model library, wherein the model library stores domain models of a plurality of different domains; or, Based on the task model identifier, a target domain model applied to the target domain is searched from the model library.

15. The method according to claim 13, wherein the request information includes a training data set; The acquiring of a target visual model based on the request information includes: Based on the training data set, the initial domain model is trained to obtain a target domain model applied to the target domain.

16. A cloud training platform, comprising a request interface and a response unit; The request interface is used to receive a task generation request for a target domain sent by a terminal device, wherein: The task generation request includes request information; The response unit is used to obtain a target domain model applied to the target domain based on the request information, wherein the target domain question-answering model is obtained by training the initial domain model based on a training data set, the training data set is constructed based on target domain data, the target domain data is recalled based on the initial key information of the target domain and the target key information of the target domain, and the target key information is key information in the target domain that is semantically similar to the initial key information and is generated under the prompt of generating prompt information by inputting the initial key information into an information processing model; based on the target domain model, task information is generated, wherein the task information is used for the terminal device to perform the target domain task.

17. The cloud training platform according to claim 16, further comprising a model library, wherein: The model library stores domain models of multiple different domains; the request information includes a task scenario identifier of a target domain task, or a task model identifier; The response unit is specifically used to determine a target scene template from multiple preset scene templates based on the task scene identifier, and based on the target scene template, search for a target domain model applied to the target domain from the model library; or, based on the task model identifier, search for a target domain model applied to the target domain from the model library.

18. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 15 are implemented.

19. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15.

20. A computer program product comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15.