Data processing method and model training method of text processing model

By introducing knowledge units and text retrieval units into the text processing model, the problems of response latency and dependence on external search engines in large language models are solved, achieving faster and more accurate text generation.

CN121660077APending Publication Date: 2026-03-13SWEET POTATO TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511776224.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing large language models suffer from high response latency and high dependence on external search engines, which affects their generation efficiency and accuracy.

Method used

By introducing knowledge units and text retrieval units into the text processing model, the generated response text information can be dynamically selected, reducing reliance on external search engines and improving generation efficiency.

Benefits of technology

It shortens the generation time of the first token, improves the user experience and generation efficiency, and provides more accurate answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121660077A_ABST
    Figure CN121660077A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and a model training method of a text processing model.The data processing method comprises the steps that a to-be-processed text is obtained, the to-be-processed text is input into the text processing model, and reply text information output by the text processing model is obtained, the text processing model comprises a knowledge unit and a text retrieval unit, the reply text information is generated according to the knowledge unit or the text retrieval unit. By means of the method, dynamic selection can be conducted according to the input to-be-processed text, and final reply text information is generated through the knowledge unit or the text retrieval unit. The knowledge updating speed and depth of the text processing model can be enriched through the text retrieval unit, an external search engine does not need to be called, and the generation efficiency of the large language model is improved. The generation time of the first token is shortened, and the use experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to data processing methods. This specification also relates to a method for training a text processing model, a computing device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] Retrieval enhancement techniques are an important aspect of large language models in application. When using large language models for knowledge reasoning and answering, external search engines can be used to retrieve information relevant to user queries from dynamic corpora, enabling the large language model to generate more accurate answers.

[0003] However, using external search engines requires steps such as query rewriting, calling the external search engine, search retrieval, and document ranking, resulting in high response latency. Furthermore, retrieval enhancement techniques heavily rely on external search engines; if the retrieved knowledge is inadequate, it will affect the final answer from the large language model. Therefore, how to shorten the response time of large language models and improve their generation efficiency has become an urgent problem for engineers to solve. Summary of the Invention

[0004] In view of this, embodiments of this specification provide a data processing method. This specification also relates to a model training method for a text processing model, a computing device, a computer-readable storage medium, and a computer program product, to address the aforementioned problems existing in the prior art.

[0005] According to a first aspect of the embodiments of this specification, a data processing method is provided, comprising: Get the text to be processed; The text to be processed is input into a text processing model to obtain the response text information output by the text processing model. The text processing model includes a knowledge unit and a text retrieval unit, and the response text information is generated based on the knowledge unit or the text retrieval unit.

[0006] According to a second aspect of the embodiments of this specification, a method for training a text processing model is provided, applied to a cloud-based device, comprising: The receiving end device sends domain knowledge text of the target business domain, and inputs the domain knowledge text into the initial text processing model to obtain a reference text processing model, wherein the initial text processing model includes knowledge units and text retrieval units; Obtain the domain sample text corresponding to the target business domain, wherein the domain sample text includes domain sample questions and domain sample answers; The domain sample question is input into the reference text processing model to obtain the predicted question response output by the reference text processing model; The model reward value is calculated based on the predicted question response and the domain sample response, and the model parameters of the reference text processing model are adjusted based on the model reward value. The reference text processing model is trained again based on the domain sample text until the model training stops, and the text processing model is obtained. The model parameters of the text processing model are then sent to the edge device.

[0007] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above method.

[0008] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0009] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0010] The data processing method provided in this specification includes acquiring a text to be processed, inputting the text to be processed into a text processing model, and obtaining response text information output by the text processing model. The text processing model includes a knowledge unit and a text retrieval unit, and the response text information is generated based on the knowledge unit or the text retrieval unit.

[0011] The method provided in the embodiments of this specification offers a novel text processing model, which includes a knowledge unit and a text retrieval unit. The knowledge unit retains the original capabilities of the large language model, storing knowledge already learned by the model. The text retrieval unit stores newly acquired knowledge, enabling the large language model to use the latest information when generating answers. This method allows for dynamic selection based on the input text to be processed, generating the final response text information through either the knowledge unit or the text retrieval unit. It enriches the knowledge update speed and depth of the text processing model through the text retrieval unit, and improves the generation efficiency of the large language model without requiring external search engines. It also shortens the generation time of the first token, enhancing the user experience. Attached Figure Description

[0012] Figure 1 This is a flowchart of a data processing method provided in one embodiment of this specification; Figure 2 This is a schematic diagram of the model structure of a text processing model provided in one embodiment of this specification; Figure 3 This is a flowchart of a data processing method provided in another embodiment of this specification; Figure 4 This is a schematic diagram of the model structure of a text processing model provided in one embodiment of this specification; Figure 5 This is an architecture diagram of a data processing system provided in one embodiment of this specification; Figure 6 This is a flowchart of a text processing model training method provided in one embodiment of this specification; Figure 7 This is a schematic diagram of the three-stage data flow of model training provided in one embodiment of this specification; Figure 8 This is a flowchart of a model training method for a text processing model applied to a cloud-based device, provided in one embodiment of this specification. Figure 9 This is a schematic diagram of the structure of a data processing device provided in one embodiment of this specification; Figure 10 This is a schematic diagram of the structure of a model training device for a text processing model provided in one embodiment of this specification; Figure 11 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0013] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0014] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0016] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0017] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0018] RAG (Retrieval-Augmented Generation) is a technique that combines language modeling and information retrieval. Specifically, when a large language model needs to generate text or answer questions, it first retrieves relevant information from a large collection of documents, and then uses this retrieved information to guide text generation, thereby improving the quality and accuracy of predictions.

[0019] TTFT (Time To First Token) is a metric for measuring the response speed of a large language model. It refers to the time elapsed from when a user sends a request to when the large language model begins generating and returning the first token (such as a Chinese character or word). TTFT reflects the initial response speed of the large language model to a request.

[0020] Traditional large language models can reason and answer user-input questions using massive amounts of knowledge acquired during training. However, this makes them inadequate for tasks requiring real-time updates or domain-specific knowledge. To address this issue, Relational Language Aggregator (RAG) technology emerged. This technology uses an external search engine to retrieve information relevant to the user query from a dynamic, real-time updated corpus. This retrieved information is then provided to the large language model as context, assisting it in providing more accurate answers.

[0021] Currently, widely used RAG (Related Aspects of Language) technology typically involves triggering a search engine to retrieve the question, identifying relevant documents, reordering these documents, and then recalling those with high relevance to the question. These recalled documents are then sent to a large language model, which uses this model to generate inferences and answers for the user's question.

[0022] While RAG technology significantly improves the accuracy of answers during the generation of reasoning and responses, it also has the following shortcomings: High response latency: The implementation of RAG technology involves multiple steps, including query rewriting, information retrieval, document sorting, and large language model generation of responses, each of which involves computation and latency. This results in a long time (TTFT) for users to see the model's response for the first time, significantly impacting the user experience.

[0023] High reliance on external search engines: Existing RAG technology heavily relies on external search engines for knowledge retrieval, which not only increases system complexity but may also reduce the accuracy of search results. If the search engine's query results do not match the question, or the ranking results are unsatisfactory, it will affect the quality of the final generated answer.

[0024] This specification provides a data processing method, and also relates to a model training method for a text processing model, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail in the following embodiments.

[0025] Figure 1 A flowchart of a data processing method according to an embodiment of this specification is shown, which specifically includes the following steps: Step 102: Obtain the text to be processed.

[0026] In this context, the text to be processed can be understood as the question text submitted by the user. In practical applications, users need to use a large language model to perform question-and-answer operations related to relevant knowledge. Users need to input the questions they want to know into the large language model. For example, "When is the Spring Festival this year?" or "When does a certain exam start every year?" The text containing the user's question is the text to be processed.

[0027] The data processing method provided in the embodiments of this specification is applied in a terminal, where a large language model or a large language model calling interface is deployed. Users can input text to be processed into the terminal, and the terminal obtains the text.

[0028] In practical applications, the terminal can display a visual interface to the user, through which it can receive text input by the user for processing. The terminal can also provide a data transmission interface to receive text sent by the user or other devices for processing via this interface.

[0029] Furthermore, the text to be processed can be plain text or multimodal data including text. For example, the text to be processed can be image and text data, audio and text data, video and text data, etc. In practical applications, users may ask questions about information in images, audio, and video. The terminal will receive data in other modalities besides the question text. In the method provided in the embodiments of this specification, the above data can be understood as the text to be processed acquired by the terminal.

[0030] Step 104: Input the text to be processed into the text processing model to obtain the response text information output by the text processing model. The text processing model includes a knowledge unit and a text retrieval unit, and the response text information is generated based on the knowledge unit or the text retrieval unit.

[0031] In this context, the text processing model can be understood as a deep learning model used to process the text to be processed; more specifically, it can be understood as a large language model. In the method provided in the embodiments of this specification, the user inputs the question to be processed into the text processing model, and the text processing model can provide corresponding response text information based on the text to be processed.

[0032] Furthermore, text processing models can also be domain-specific language models, which are trained or adjusted based on general-purpose language models and tailored to specific industries, domains, or topics using dedicated domain datasets. In other words, to better provide question-answering services for a particular target business domain, a specialized language model is typically trained for that domain. For example, in the medical field, a corresponding medical language model can be trained using medical knowledge information; similarly, in the football field, a corresponding football language model can be trained using football knowledge information, and so on.

[0033] In the methods provided in the embodiments of this specification, the text processing model, in addition to possessing the basic functions of a general-purpose large language model, can also provide text retrieval capabilities. That is, the text processing model includes knowledge units and text retrieval units.

[0034] Among them, the knowledge unit maintains the basic capabilities of the general large language model, stores the knowledge that the model has already learned, and is not affected by other information. It can generate the corresponding response text information of the text to be processed through the text processing model's own capabilities.

[0035] The text retrieval unit is used to store newly acquired knowledge, including information obtained from external corpora. This text retrieval unit enables the text processing model to retrieve relevant knowledge text based on the text to be processed and to generate response text information based on the relevant knowledge text.

[0036] Current RAG technology requires calling an external search engine to retrieve relevant text, which is then input into a large language model for processing to generate the answer. However, the text processing model provided in this specification integrates knowledge units and text retrieval units. When the text to be processed is input into the text processing model, the model can determine whether to use a knowledge unit or a text retrieval unit to process it based on the specific information of the text.

[0037] If the text to be processed is processed by a knowledge unit, the text to be processed can be input into the knowledge unit, and the knowledge unit can generate a response text information for the text to be processed based on the knowledge it has learned.

[0038] If the text to be processed is processed by a text retrieval unit, the text to be processed can be input into the text retrieval unit. The text retrieval unit can generate at least one associated text information corresponding to the text to be processed, and generate the final response text information based on the associated text information.

[0039] See Figure 2 , Figure 2 A schematic diagram of the model structure of the text processing model provided in the embodiments of this specification is shown, such as... Figure 2 As shown, the text processing model includes knowledge units and text retrieval units. When the text to be processed is input into the text processing model, the model determines whether to input it into the knowledge unit or the text retrieval unit based on the content information of the text.

[0040] After the text to be processed is processed by the knowledge unit or text retrieval unit, the corresponding response text information will be generated and output by the text processing model.

[0041] The method provided in the embodiments of this specification offers a novel text processing model, which includes a knowledge unit and a text retrieval unit. The knowledge unit retains the original capabilities of the large language model, storing knowledge already learned by the model. The text retrieval unit stores newly acquired knowledge, enabling the large language model to use the latest information when generating answers. This method allows for dynamic selection based on the input text to be processed, generating the final response text information through either the knowledge unit or the text retrieval unit. It enriches the knowledge update speed and depth of the text processing model through the text retrieval unit, and improves the generation efficiency of the large language model without requiring external search engines. It also shortens the generation time of the first token, enhancing the user experience.

[0042] Figure 3 A flowchart of a data processing method according to another embodiment of this specification is shown, which specifically includes the following steps: Step 302: Obtain the text to be processed.

[0043] In this embodiment, the specific interpretation of the text to be processed can be found in the relevant description in step 102 above, and will not be repeated here.

[0044] Step 304: Input the text to be processed into the routing unit of the text processing model, and obtain the processing unit determined by the routing unit based on the text to be processed, wherein the processing unit includes a knowledge unit or a text retrieval unit.

[0045] In this embodiment, the text to be processed is still input into the text processing model for processing. In the method provided in the embodiments of this specification, the text processing model includes a routing unit, a knowledge unit, and a text and text retrieval unit.

[0046] See Figure 4 , Figure 4 This embodiment shows a schematic diagram of the text processing model structure. Figure 4 As shown, this text processing model includes a routing unit, a knowledge unit, and a text retrieval unit. The routing unit determines the corresponding processing unit based on the input text to be processed; that is, whether the text should be processed by the knowledge unit or the text retrieval unit. After the processing unit is determined, the text to be processed can be input into the corresponding processing unit for further processing.

[0047] In practical applications, the routing unit can be trained through the model training phase of the text processing model, gaining the ability to determine the processing unit based on the text content of the text to be processed. The specific training method of the text processing model will be explained in subsequent embodiments and will not be repeated here.

[0048] In this embodiment, after the text to be processed is input into the text processing model, it is first input into the routing unit. The routing unit determines the subsequent processing unit based on the information of the text to be processed. Here, the processing unit does not refer to a specific unit in the text processing model, but is a general term for the knowledge unit or text retrieval unit used to process the text to be processed and generate the final response text information. That is, if the text to be processed is subsequently processed by a knowledge unit, then the knowledge unit is the processing unit; if the text to be processed is subsequently processed by a text retrieval unit, then the text retrieval unit is the processing unit.

[0049] Step 306: Input the text to be processed into the knowledge unit or text retrieval unit of the text processing model to obtain the response text information.

[0050] In the steps described above, after identifying the processing unit corresponding to the text to be processed as either a knowledge unit or a text retrieval unit of the text processing model, the text to be processed can be input into the knowledge unit or text retrieval unit for subsequent processing to generate the final response text information output by the text processing model.

[0051] In one specific embodiment provided in this specification, the text to be processed is input into the knowledge unit or the text retrieval unit to obtain response text information, including: The text to be processed is input into the knowledge unit to obtain the response text information output by the knowledge unit; or, The text to be processed is input into the text retrieval unit to obtain the associated text information output by the text retrieval unit, and a response text information is generated based on the associated text information.

[0052] In this embodiment, the processing of the text to be processed in the knowledge unit and the text retrieval unit will be further explained.

[0053] When the text to be processed is input into the knowledge unit, the knowledge unit retains the original capabilities of the large language model, stores the knowledge that the text processing model has already learned, and generates corresponding response text information for the text to be processed based on the knowledge that has already been learned.

[0054] When the text to be processed is input into the text retrieval unit, the unit stores the newly acquired knowledge. The text retrieval unit is pre-trained with a model, which can convert the knowledge text information into model parameters for storage. Upon receiving the text to be processed, it can generate corresponding related text information based on the model parameters corresponding to the knowledge information. Then, it generates corresponding response text information based on the related text information.

[0055] In another specific embodiment provided in this specification, the text to be processed can also be simultaneously input into the knowledge unit and the text retrieval unit. The knowledge unit generates first response information corresponding to the text to be processed, the text retrieval unit obtains associated text information corresponding to the text to be processed, the associated text information generates second response information based on the associated text information, and the first response information and the second response information are combined to generate response text information output by the text processing model.

[0056] The method provided in the embodiments of this specification incorporates a routing unit within the text processing model. This routing unit determines the appropriate processing unit to respond to the text based on its content. Once the processing unit is determined to be either a knowledge unit or a text retrieval unit, the text is forwarded to the corresponding unit for processing. This approach preserves the original capabilities of the text processing model while also providing retrieval capabilities for generating the final response text. The text processing model provided in the embodiments of this specification can provide more accurate responses to text.

[0057] See Figure 5 , Figure 5 This specification illustrates an architecture diagram of a data processing system according to one embodiment of the present specification. The data processing system may include a client 100 and a server 200. Client 100 is used to send the text to be processed to server 200; Server 200 is used to input the text to be processed into a text processing model, obtain the response text information output by the text processing model, wherein the text processing model includes a knowledge unit and a text retrieval unit, and the response text information is generated based on the knowledge unit or the text retrieval unit; and send the response text information to client 100. Client 100 is also used to receive reply text information sent by server 200.

[0058] The data processing system may include multiple clients 100 and a server 200. Clients 100 can be referred to as edge devices, and server 200 can be referred to as cloud devices. Multiple clients 100 can establish communication connections through server 200. In a text data processing scenario, server 200 is used to provide text data processing services between multiple clients 100. Each client 100 can act as a sender or receiver, communicating through server 200.

[0059] Users can interact with server 200 through client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In text data processing scenarios, users can publish data streams to server 200 through client 100, and server 200 can generate reply text information based on the data stream and push the reply text information to other clients that have established communication.

[0060] In this system, client 100 and server 200 establish a connection via a network. The network provides the medium for communication between client 100 and server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by client 100 may need to undergo encoding, transcoding, compression, or other processing before being published to server 200.

[0061] Client 100 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. Client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by server 200, such as a real-time communication (RTC) SDK. Client 100 can be deployed on a computing device and depends on the device or certain apps on the device to run. The computing device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured on the computing device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.

[0062] Server 200 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0063] It is worth noting that the data processing methods provided in the embodiments of this specification are generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the data processing methods provided in the embodiments of this specification. In other embodiments, the data processing methods provided in the embodiments of this specification may also be executed jointly by the client and the server.

[0064] The above embodiments use a specific application of a text processing model as an example for explanation, specifically including the deployment of the text processing model on a terminal or a server. The text retrieval unit of the text processing model provided in this specification has the ability to store new knowledge and provide recall information for the problem to be processed, eliminating the need to call an external search engine, thereby improving the response accuracy of the text processing model. The following embodiments will further explain the model training method of the text processing model. Specifically, Figure 6 A flowchart of a text processing model training method according to an embodiment of this specification is shown, which specifically includes the following steps: Step 602: Obtain domain knowledge text of the target business domain and input the domain knowledge text into the initial text processing model to obtain a reference text processing model, wherein the initial text processing model includes a routing unit, a knowledge unit and a text retrieval unit.

[0065] In practical applications, domain-specific large language models have been increasingly used in various scenarios. A domain-specific large language model refers to a large language model obtained by training or adjusting a general-purpose large language model for a specific industry, professional scenario, or topic range using a dedicated domain dataset. For example, in the medical field, a general-purpose large language model can be trained using medical data to generate a medical large language model, which can then be used to answer medical questions. Similarly, in the financial field, a general-purpose large language model can be trained using financial data to generate a financial large language model.

[0066] Domain-specific large language models significantly outperform general-purpose large language models in tasks within their respective professional domains, demonstrating superior performance in specific professional fields. Currently, the application of domain-specific large language models is becoming increasingly widespread.

[0067] Based on this, the text processing model provided in the embodiments of this specification can be a domain-specific large language model for the target business domain. Specifically, it can be trained by inputting domain knowledge text of the target business domain into a general large language model, so that the large language model possesses domain knowledge of the target business domain.

[0068] Specifically, the target business domain can be understood as a specific business area, such as the medical field, the financial field, the sports field, etc. Domain knowledge texts can be understood as knowledge textual information related to the target business domain, such as medical papers, pathology guidelines, and drug instructions in the medical field; financial news, policy documents, and risk control case studies in the financial field; and laws, regulations, judgments, professional articles, and textbooks in the legal field.

[0069] After obtaining the domain knowledge text of the target business domain, the domain knowledge text can be input into the initial text processing model for training to obtain the reference text processing model. It should be noted that the initial text processing model includes a routing unit, a knowledge unit, and a text retrieval unit. Detailed information on the routing unit, knowledge unit, and text retrieval unit can be found in the relevant descriptions in the above embodiments, and will not be repeated here.

[0070] Furthermore, the initial text processing model is trained using domain knowledge text, specifically referring to training the text retrieval unit within the initial text processing model. That is, in another specific embodiment provided in this specification, domain knowledge text of the target business domain is obtained, and this domain knowledge text is input into the initial text processing model to obtain a reference text processing model, including: Obtain domain knowledge text and an initial text processing model for the target business domain, wherein the initial text processing model includes knowledge units and text retrieval units; The domain knowledge text is input into the initial text processing model to train the text retrieval unit in the initial text processing model, thereby obtaining a reference text processing model.

[0071] In the specific implementation provided in this specification, the first step is to obtain the domain knowledge text of the target business domain and the initial text processing model. There are many ways to obtain the domain knowledge text; for example, it can be queried from the domain knowledge base of the target business domain, or it can be searched from the internet, etc.

[0072] After obtaining the domain knowledge text, it can be input into the initial text processing model to allow the model to learn the domain knowledge text. The text retrieval unit within the initial model is then trained, and its parameters are adjusted to equip it with the necessary parameters for recording the domain knowledge text. This process yields a reference text processing model. The reference text processing model can be understood as a text processing model trained on top of the initial model, having learned the domain knowledge text of the target business domain.

[0073] In practical applications, the training process of continuously inputting domain knowledge text into the initial text processing model to obtain the reference text processing model satisfies the Continuous Pre-Training (CPT) paradigm. This refers to continuing to train the model using new data (such as domain-specific data, new knowledge, etc.) on the basis of the model that has already completed general pre-training, so as to realize the model knowledge update and capability expansion, while reducing the high cost of training from scratch.

[0074] The goal of the continuous pre-training paradigm is to address the limitations of large models when faced with new domain knowledge and real-time information, while mitigating catastrophic forgetting in continuous learning. It aims to balance the model's retention of old knowledge with the absorption of new knowledge. The current CPT paradigm performs language modeling tasks autoregressively on new data following a pre-training paradigm, thereby training the text processing model to remember domain-specific text knowledge.

[0075] However, the autoregressive paradigm of language models can only perform unidirectional memorization of the corpus, which is inefficient, and real users' queries are often the opposite of the corpus. For example, if the corpus contains "a certain day of a certain month is holiday A", a user might typically query "which day is holiday A". Models trained using the autoregressive paradigm may experience stuttering when dealing with the above problems. To solve this problem, the method provided in the embodiments of this specification provides a bidirectional causal attention mechanism. Specifically, in another specific embodiment provided in the embodiments of this specification, the domain knowledge text includes domain knowledge statements and domain tag statements; Inputting the domain knowledge text into the initial text processing model and training the text retrieval unit in the initial text processing model includes: The domain knowledge statement and the domain tag statement are input into the text retrieval unit of the initial text processing model to obtain the domain prediction statement generated by the text retrieval unit based on the domain knowledge statement; The retrieval loss value is calculated based on the domain prediction statement and the domain tag statement, and the model parameters of the text retrieval unit are adjusted based on the retrieval loss value.

[0076] In the actual training process of the initial text processing model, the domain knowledge text is segmented into sentences, and these sentences are designated as domain knowledge statements and domain label statements. Both domain knowledge statements and domain label statements are actual statements within the domain knowledge text. For example, a domain knowledge text might include "statement 1, statement 2, statement 3, statement 4". Statements 1 and 2 can be considered domain knowledge statements, while statements 3 and 4 can be considered domain label statements. Alternatively, statements 3 and 4 can be considered domain knowledge statements, while statements 1 and 2 can be considered domain label statements.

[0077] When statements 1 and 2 are used as domain knowledge statements, statements 3 and 4 are used as domain label statements to train the model's forward memory ability. When statements 3 and 4 are used as domain knowledge statements, statements 1 and 2 are used as domain label statements to train the model's backward memory ability. This approach ensures that the initial text processing model can perform bidirectional memorization of the corpus during the CPT stage, while also adhering to the original autoregressive generation paradigm. Without affecting the capabilities of the large language model itself, bidirectional memorization of domain knowledge text can improve the model's memory efficiency and reduce its sensitivity to statement order.

[0078] The process of training an initial text processing model using domain-knowledge text specifically involves simultaneously inputting domain-knowledge statements and domain-labeled statements into the text retrieval unit of the initial text processing model. This allows the text retrieval unit to memorize the domain-knowledge statements and domain-labeled statements. The text retrieval unit can then predict and generate domain-predicted statements based on the domain-knowledge statements. The retrieval loss value is calculated by comparing the domain-predicted statements with the domain-labeled statements, and the model parameters of the text retrieval unit are adjusted based on this loss value.

[0079] In practical applications, domain knowledge text is pre-input into the text retrieval unit of the initial text processing model for memorization. Then, the beginning sentences of the domain knowledge text are masked, for example, the first quarter, first half, or first three-quarters of the text. The text retrieval unit of the initial text processing model predicts the masked portion based on the masked content, thus achieving prediction of the preceding content. The prediction result is then compared with the unmasked content to calculate the retrieval loss value, thereby adjusting the model parameters of the text retrieval unit. This enables the text retrieval unit to retrieve text information.

[0080] Step 604: Obtain the domain sample text corresponding to the target business domain, wherein the domain sample text includes domain sample questions and domain sample answers.

[0081] After the above training, the reference text processing model will possess retrieval capabilities. Further training is then needed to enable the model to retrieve information from the text to be processed and generate answers.

[0082] Specifically, the process involves acquiring domain sample text corresponding to the target business domain. This domain sample text includes domain sample questions and domain sample responses. Domain sample questions can be understood as sample questions used for model training, and domain sample responses can be understood as the answer information corresponding to the domain sample questions.

[0083] In the method provided in the embodiments of this specification, the process of continuing to train the reference text processing model is divided into the SFT stage (supervised fine-tuning stage) and the RL stage (reinforcement learning stage).

[0084] The Supervised Fine-Tuning (SFT) phase involves targeted training on a pre-trained model using manually labeled instruction-response structured data. The goal is to enable the model to learn the specific task requirements and understand the intended meaning of commands. The SFT phase is often referred to as the cold start phase. Cold start is the initialization goal of a model moving from a basic state to a state ready for the reinforcement learning (RL) phase, and SFT is the technique used to achieve this.

[0085] The core of cold start in large model training is to establish a stable starting point for subsequent RL training. This process is accomplished through SFT. If the cold start, which is dominated by the SFT stage, is skipped and RL training is performed directly on the model, problems such as mixed model languages ​​and inconsistent formats will occur.

[0086] The Reinforcement Learning (RL) stage guides the model to optimize its output strategy through a reward mechanism, moving beyond the mechanical imitation of the SFT stage and ensuring that the content generated by the model better aligns with user preferences in terms of security, usefulness, and other aspects. The RL stage is used to train the retrieval capabilities of the reference text processing model.

[0087] Based on this, in a specific embodiment provided in this specification, obtaining the domain sample text corresponding to the target business domain includes: Obtain a preset number of domain retrieval questions and domain knowledge questions corresponding to the target business domain; Search tags are set for the domain search questions, wherein the search tags indicate that the domain search questions are processed by the text search unit; The domain knowledge question-and-answer questions and the domain retrieval question-and-answer questions with search tags are used as domain sample text.

[0088] In the methods provided in the embodiments of this specification, in order to better train the reference text processing model, domain sample text corresponding to the target business domain is obtained. Furthermore, in order to train the answering and retrieval capabilities of the reference text processing model, the domain sample text is divided into domain retrieval question-answering questions and domain knowledge question-answering questions.

[0089] Among them, domain retrieval question-answering questions are used to guide the text retrieval units in the reference text processing model to generate answers, and domain knowledge question-answering questions are used to guide the knowledge units in the reference text processing model to generate answers.

[0090] In practical applications, to better train the reference text processing model's ability to retrieve information from its own memory, a specific structured data format is set for domain retrieval question-answering questions. This structured data format includes a title, content, and answer, thereby generating domain retrieval question-answering questions.

[0091] To ensure the reference text processing model can accurately identify domain retrieval question-and-answer questions and domain knowledge question-and-answer questions, search tags need to be set for these questions. These tags represent how the domain retrieval question-and-answer questions are processed by the text retrieval unit. Furthermore, the ratio of domain retrieval question-and-answer questions to domain knowledge question-and-answer questions should be less than a preset threshold; ideally, the ratio should be one to one.

[0092] At this point, a preset number of domain retrieval questions with search tags and a preset number of domain knowledge questions have been obtained, and the two have been combined as domain sample text.

[0093] Step 606: Input the domain sample question into the reference text processing model to obtain the predicted question response output by the reference text processing model.

[0094] Given domain sample text, which includes domain sample questions and domain sample responses, the domain sample questions are input into a reference text processing model for SFT cold start. The predicted question responses are then obtained from the output of the reference text processing model.

[0095] In one specific embodiment provided in this specification, the domain sample question is input into the reference text processing model to obtain the predicted question response output by the reference text processing model, including: The domain sample question is input into the routing unit of the reference text processing model to obtain the processing unit determined by the routing unit, wherein the processing unit is determined based on whether the domain sample question has a search tag set; When the processing unit is a knowledge unit, the domain sample question is input into the knowledge unit to obtain the predicted question response output by the knowledge unit; When the processing unit is a text retrieval unit, the domain sample question is input into the text retrieval unit to obtain the predicted text information output by the text retrieval unit, and a predicted question response is generated based on the predicted text information.

[0096] In the methods provided in the embodiments of this specification, the domain sample text includes domain retrieval question-and-answer questions and domain knowledge question-and-answer questions. Accordingly, the domain sample question can be either a domain sample question for domain retrieval question-and-answer or a domain sample question for domain knowledge question-and-answer.

[0097] In this embodiment, the domain sample question is input into the routing unit of the reference text processing model. The routing unit can identify whether the domain sample question contains pre-set search tags and learn the question content. During the model training phase, the routing unit determines the subsequent processing unit for the domain sample question based on whether it contains search tags.

[0098] For example, if a domain sample question contains search tags, the routing unit can determine that the domain sample question belongs to a text retrieval task, and thus determine that its corresponding processing unit is the text retrieval unit, and route the domain sample question to the text retrieval unit for processing. If a domain sample question does not contain search tags, the routing unit can determine that the domain sample question does not belong to a text retrieval task, and thus determine that its corresponding processing unit is the knowledge unit, and route the domain sample question to the knowledge unit for processing.

[0099] When the processing unit is determined to be a knowledge unit, the domain sample problem can be input into the knowledge unit for processing to obtain the predicted problem response output by the knowledge unit.

[0100] When the processing unit is determined to be a text retrieval unit, the domain sample question can be input into the text retrieval unit for processing, the predicted text information output by the text retrieval unit can be obtained, and a predicted question response can be generated based on the predicted text information.

[0101] By setting retrieval tags in the domain sample questions using the method provided in this embodiment, it is easier to guide the routing unit to determine the corresponding processing unit during the model training phase. This allows the model to better learn which questions will be routed to the knowledge unit and which questions will be routed to the text retrieval unit, providing a data routing basis for subsequent model applications.

[0102] In yet another specific embodiment provided in this specification, the domain sample problem is provided with search tags; The domain sample question is input into the text retrieval unit to obtain the predicted text information output by the text retrieval unit, and a predicted question response is generated based on the predicted text information, including: The domain sample question is input into the text retrieval unit to obtain the predicted text information output by the text retrieval unit, and a predicted response information is generated based on the predicted text information. A predicted question response is generated based on the predicted text information and the predicted response information.

[0103] In the method provided in the embodiments of this specification, if the domain sample question has search tags, the domain sample question is input into the text retrieval unit. The text retrieval unit predicts the predicted text information related to the domain sample question based on the learned model parameters, and generates predicted response information based on the predicted text information. Finally, for the subsequent reinforcement learning stage, the predicted text information and the predicted response information need to be used together as the predicted question response.

[0104] Step 608: Calculate the model reward value based on the predicted question response and the domain sample response, and adjust the model parameters of the reference text processing model based on the model reward value.

[0105] After the SFT cold start phase, the model enters the RL phase. In the RL phase, to enable the reference text processing model to perform searches better and generate more accurate answers, a corresponding model reward value is designed. Specifically, the model reward value is calculated based on the predicted question response and the domain sample response, and the model parameters of the reference text processing model are adjusted based on this reward value.

[0106] The model reward value is a quantitative score given by the model for the output result. It is a key numerical signal that connects user preferences and model training. Its purpose is to provide a clear direction for model adjustment and ensure that the output of the trained model meets the requirements.

[0107] The model's reward value is not a fixed, manually set document, but a scalar value calculated by a specially trained reward model. The reward value provides a concrete standard for the originally abstract training objective. Higher reward values ​​are assigned to responses that meet the requirements, and lower reward values ​​are assigned to responses that do not meet the requirements, thus instructing the model to output the final answer in a more compliant direction.

[0108] In one specific embodiment provided in this specification, the domain sample response includes sample text information and sample response information; The model reward value is calculated based on the predicted question response and the domain sample response, including: A text reward value is generated based on the predicted text information and the sample text information; A response reward value is generated based on the predicted response information and the sample response information; A model reward value is generated based on the text reward value and the response reward value.

[0109] In the methods provided in the embodiments of this specification, corresponding model reward values ​​are designed for training the reference text processing model. Specifically, those skilled in the art expect the reference text processing model to generate predictive text information that meets the requirements, and to generate predictive response information based on the predictive text information.

[0110] Therefore, the domain sample response specifically includes sample text information and sample response information for answering the domain sample questions. After obtaining the predicted text information and predicted response information from the reference text processing model, they can be compared separately to determine the final model reward value.

[0111] Specifically, the predicted text information can be compared with the sample text information to determine the similarity between the two texts and calculate the text reward value, thereby determining whether the retrieved text meets the requirements. Furthermore, the predicted text also includes a title and content. When calculating the text reward value, it can be calculated more granularly, using the title and content as dimensions. That is, the title reward value is calculated between the predicted title in the predicted text information and the sample title in the sample text information; the content reward value is calculated between the predicted content in the predicted text information and the sample content in the sample text information; and then the text reward value is calculated based on the title reward value and the content reward value.

[0112] Secondly, it is necessary to further compare the predicted response information with the sample response information to determine whether the two are consistent, thereby generating a response reward value to determine whether the answer is right or wrong.

[0113] Finally, the model reward value for this training can be calculated based on the text reward value and the response reward value.

[0114] In another specific embodiment provided in this specification, the corresponding format information can be preset for the output of the model, and it can be determined whether the final output result meets the preset format information, and a format reward value can be given according to the determination result.

[0115] Finally, the model reward value for this training is determined based on the text reward value, response reward value, and format reward value.

[0116] Once the model reward value is obtained, the parameters of the reference text processing model can be adjusted based on the model reward value. More specifically, the parameters of the text retrieval unit in the reference text processing model can be adjusted.

[0117] Step 610: Continue training the reference text processing model based on the domain sample text until the model training stops, and obtain the text processing model.

[0118] The above is an explanation of training the reference text processing model with one or more domain sample texts. The reference text processing model can be continuously trained based on the domain sample texts in the training samples until the model training stops, thereby obtaining the final text processing model.

[0119] In one specific embodiment provided in this specification, the stopping condition for model training may be that the number of training rounds has reached a preset number of training rounds, the training duration has met a preset duration threshold, the model reward value has met a preset reward value threshold, etc. In the method provided in this specification, the stopping condition for model training is not limited, and the actual application shall prevail.

[0120] The model training method for the text processing model provided in this specification involves inputting domain knowledge text from the target business domain into an initial text processing model. This allows the initial text processing model to learn the input domain knowledge text, and the learned content is saved as model parameters in the text retrieval unit. This facilitates subsequent processing without calling an external search engine, providing a data foundation for saving the running time of the text processing model. Furthermore, a bidirectional causal attention mechanism is used during the training of the initial text processing model to learn the domain knowledge text. This allows the initial text processing model to bidirectionally memorize the corpus during the CPT stage while still adhering to the original autoregressive generation paradigm. This allows for bidirectional memorization of documents without affecting the capabilities of the larger model itself, resulting in higher memorization efficiency and reduced sensitivity to the order of the corpus.

[0121] In addition, to ensure that the text processing model can accurately retrieve data from its own memory and optimize the accuracy of answer generation, a reinforcement learning mechanism was designed, and a corresponding model reward value was designed for it. In this way, the relevance of the retrieval results and the quality of the final response are improved by the text retrieval unit during the retrieval process.

[0122] See Figure 7 , Figure 7 This specification illustrates a schematic diagram of the three-stage data flow during model training according to an embodiment, as shown below. Figure 7 As shown, the training of the text processing model is divided into three stages: continuous pre-training, cold start, and reinforcement learning. In the continuous pre-training stage, the input training data is routed to the text retrieval unit via the routing unit to train the text retrieval unit.

[0123] During the cold start phase, the routing unit routes the training data to the knowledge unit or text retrieval unit based on the labels of the training data to train the text processing model.

[0124] During the reinforcement learning phase, the routing unit will still route the training data to the knowledge unit or text retrieval unit based on the labels of the training data, and continue to train the text processing model through the pre-designed model reward value.

[0125] See Figure 8 , Figure 8 This specification illustrates a flowchart of a model training method for a text processing model applied to a cloud-based device, as provided in one embodiment. Figure 8 As shown, the method includes: Step 802: Receive domain knowledge text of the target business domain sent by the receiving end device, and input the domain knowledge text into the initial text processing model to obtain a reference text processing model, wherein the initial text processing model includes knowledge units and text retrieval units.

[0126] Step 804: Obtain the domain sample text corresponding to the target business domain, wherein the domain sample text includes domain sample questions and domain sample answers.

[0127] Step 806: Input the domain sample question into the reference text processing model to obtain the predicted question response output by the reference text processing model.

[0128] Step 808: Calculate the model reward value based on the predicted question response and the domain sample response, and adjust the model parameters of the reference text processing model based on the model reward value.

[0129] Step 810: Continue training the reference text processing model based on the domain sample text until the model training stop condition is met, obtain the text processing model, and send the model parameters of the text processing model to the edge device.

[0130] In practical applications, training a model requires a large amount of data and significant computing resources, which edge devices may lack. Therefore, the model training process can be implemented on cloud devices. After obtaining the model parameters of the text processing model, the cloud device can send these parameters to the edge device. The edge device can then build a text processing model locally based on these parameters and further utilize the model for text data processing.

[0131] The model training method for the text processing model provided in this specification involves inputting domain knowledge text from the target business domain into an initial text processing model. This allows the initial text processing model to learn the input domain knowledge text, and the learned content is saved as model parameters in the text retrieval unit. This facilitates subsequent processing without calling an external search engine, providing a data foundation for saving the running time of the text processing model. Furthermore, a bidirectional causal attention mechanism is used during the training of the initial text processing model to learn the domain knowledge text. This allows the initial text processing model to bidirectionally memorize the corpus during the CPT stage while still adhering to the original autoregressive generation paradigm. This allows for bidirectional memorization of documents without affecting the capabilities of the larger model itself, resulting in higher memorization efficiency and reduced sensitivity to the order of the corpus.

[0132] In addition, to ensure that the text processing model can accurately retrieve data from its own memory and optimize the accuracy of answer generation, a reinforcement learning mechanism was designed, and a corresponding model reward value was designed for it. In this way, the relevance of the retrieval results and the quality of the final response are improved by the text retrieval unit during the retrieval process.

[0133] Corresponding to the above method embodiments, this specification also provides data processing apparatus embodiments. Figure 9 A schematic diagram of the structure of a data processing apparatus provided in one embodiment of this specification is shown. For example... Figure 9 As shown, the device includes: Module 902 is configured to acquire text to be processed. The processing module 904 is configured to input the text to be processed into a text processing model and obtain the response text information output by the text processing model. The text processing model includes a knowledge unit and a text retrieval unit, and the response text information is generated based on the knowledge unit or the text retrieval unit.

[0134] In one specific embodiment provided in this specification, the text processing model further includes a routing unit; The processing module 904 is further configured to: The text to be processed is input into the routing unit to obtain the processing unit determined by the routing unit based on the text to be processed, wherein the processing unit includes the knowledge unit or the text retrieval unit; The text to be processed is input into the knowledge unit or the text retrieval unit to obtain the response text information.

[0135] In one specific embodiment provided in this specification, the processing module 904 is further configured to: The text to be processed is input into the knowledge unit to obtain the response text information output by the knowledge unit; or, The text to be processed is input into the text retrieval unit to obtain the associated text information output by the text retrieval unit, and a response text information is generated based on the associated text information.

[0136] In one specific embodiment provided in this specification, the apparatus further includes a model training module, configured to: The domain knowledge text of the target business domain is obtained and input into the initial text processing model to obtain the reference text processing model. The initial text processing model includes a routing unit, a knowledge unit, and a text retrieval unit. Obtain the domain sample text corresponding to the target business domain, wherein the domain sample text includes domain sample questions and domain sample answers; The domain sample question is input into the reference text processing model to obtain the predicted question response output by the reference text processing model; The model reward value is calculated based on the predicted question response and the domain sample response, and the model parameters of the reference text processing model are adjusted based on the model reward value. The reference text processing model is trained again based on the domain sample text until the model training stops, thus obtaining the text processing model.

[0137] In one specific embodiment provided in this specification, the model training module is further configured as follows: Obtain domain knowledge text and an initial text processing model for the target business domain, wherein the initial text processing model includes knowledge units and text retrieval units; The domain knowledge text is input into the initial text processing model to train the text retrieval unit in the initial text processing model, thereby obtaining a reference text processing model.

[0138] In one specific embodiment provided in this specification, the domain knowledge text includes domain knowledge statements and domain tag statements; The model training module is further configured as follows: The domain knowledge statement and the domain tag statement are input into the text retrieval unit of the initial text processing model to obtain the domain prediction statement generated by the text retrieval unit based on the domain knowledge statement; The retrieval loss value is calculated based on the domain prediction statement and the domain tag statement, and the model parameters of the text retrieval unit are adjusted based on the retrieval loss value.

[0139] In one specific embodiment provided in this specification, the model training module is further configured as follows: The domain sample question is input into the routing unit of the reference text processing model to obtain the processing unit determined by the routing unit, wherein the processing unit is determined based on whether the domain sample question has a search tag set; When the processing unit is a knowledge unit, the domain sample question is input into the knowledge unit to obtain the predicted question response output by the knowledge unit; When the processing unit is a text retrieval unit, the domain sample question is input into the text retrieval unit to obtain the predicted text information output by the text retrieval unit, and a predicted question response is generated based on the predicted text information.

[0140] In one specific embodiment provided in this specification, when the domain sample problem is set with search tags; The model training module is further configured as follows: The domain sample question is input into the text retrieval unit to obtain the predicted text information output by the text retrieval unit, and a predicted response information is generated based on the predicted text information. A predicted question response is generated based on the predicted text information and the predicted response information.

[0141] In one specific embodiment provided in this specification, the domain sample response includes sample text information and sample response information; The model training module is further configured as follows: A text reward value is generated based on the predicted text information and the sample text information; A response reward value is generated based on the predicted response information and the sample response information; A model reward value is generated based on the text reward value and the response reward value.

[0142] In one specific embodiment provided in this specification, the model training module is further configured as follows: Obtain a preset number of domain retrieval questions and domain knowledge questions corresponding to the target business domain; Search tags are set for the domain search questions, wherein the search tags indicate that the domain search questions are processed by the text search unit; The domain knowledge question-and-answer questions and the domain retrieval question-and-answer questions with search tags are used as domain sample text.

[0143] The apparatus provided in the embodiments of this specification offers a novel text processing model, which includes a knowledge unit and a text retrieval unit. The knowledge unit retains the original capabilities of the large language model, storing knowledge already learned by the model. The text retrieval unit stores newly acquired knowledge, enabling the large language model to use the latest information when generating responses. This apparatus allows for dynamic selection based on the input text to be processed, generating the final response text information through either the knowledge unit or the text retrieval unit. It enriches the knowledge update speed and depth of the text processing model through the text retrieval unit, and improves the generation efficiency of the large language model without requiring external search engines. It also shortens the generation time of the first token, enhancing the user experience.

[0144] The above is an illustrative scheme of a data processing apparatus according to this embodiment. It should be noted that the technical solution of this data processing apparatus and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the data processing apparatus, please refer to the description of the technical solution of the data processing method described above.

[0145] Corresponding to the above method embodiments, this specification also provides embodiments of a model training device for a text processing model. Figure 10 This specification illustrates a schematic diagram of a model training apparatus for a text processing model according to an embodiment of this specification. Figure 10 As shown, the device includes: The receiving module 1002 is configured to receive domain knowledge text of the target business domain sent by the end-side device, and input the domain knowledge text into the initial text processing model to obtain a reference text processing model, wherein the initial text processing model includes knowledge units and text retrieval units; The acquisition module 1004 is configured to acquire domain sample text corresponding to the target business domain, wherein the domain sample text includes domain sample questions and domain sample answers; Processing module 1006 is configured to input the domain sample question into the reference text processing model and obtain the predicted question response output by the reference text processing model; The training module 1008 is configured to calculate a model reward value based on the predicted question response and the domain sample response, and to adjust the model parameters of the reference text processing model based on the model reward value. The sending module 1010 is configured to continue training the reference text processing model based on the domain sample text until the model training stop condition is met, obtain the text processing model, and send the model parameters of the text processing model to the end device.

[0146] The model training device for the text processing model provided in this specification's embodiments inputs domain knowledge text from the target business domain into the initial text processing model, enabling the initial text processing model to learn the input domain knowledge text. The learned content is then saved as model parameters in the text retrieval unit, facilitating subsequent processing without requiring external search engines and providing a data foundation for saving the text processing model's runtime. Simultaneously, a bidirectional causal attention mechanism is used during the training of the initial text processing model to learn the domain knowledge text. This allows the initial text processing model to bidirectionally memorize the corpus during the CPT stage while adhering to the original autoregressive generation paradigm. This allows for bidirectional memorization of documents without affecting the capabilities of the larger model itself, resulting in higher memorization efficiency and reduced sensitivity to corpus order.

[0147] In addition, to ensure that the text processing model can accurately retrieve data from its own memory and optimize the accuracy of answer generation, a reinforcement learning mechanism was designed, and a corresponding model reward value was designed for it. In this way, the relevance of the retrieval results and the quality of the final response are improved by the text retrieval unit during the retrieval process.

[0148] The above is an illustrative scheme of a model training device for a text processing model according to this embodiment. It should be noted that the technical solution of this model training device for a text processing model belongs to the same concept as the technical solution of the model training method for a text processing model described above. For details not described in detail in the technical solution of the model training device for a text processing model, please refer to the description of the technical solution of the model training method for a text processing model described above.

[0149] Figure 11 A structural block diagram of a computing device 1100 according to an embodiment of this specification is shown. The components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.

[0150] The computing device 1100 also includes an access device 1140, which enables the computing device 1100 to communicate via one or more networks 1160. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1140 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0151] In one embodiment of this specification, the aforementioned components of the computing device 1100 and Figure 11 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 11 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0152] The computing device 1100 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1100 can also be a mobile or stationary server.

[0153] The processor 1120 is used to execute the following computer program / instruction, which, when executed by the processor, implements the steps of the above-mentioned data processing method and text processing model training method.

[0154] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the technical solutions of the data processing method and the model training method of the text processing model described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the data processing method and the model training method of the text processing model described above.

[0155] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method and text processing model training method.

[0156] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solutions of the data processing method and the model training method of the text processing model described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the data processing method and the model training method of the text processing model described above.

[0157] An embodiment of this specification also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the above-described data processing method and text processing model training method.

[0158] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solutions of the data processing method and the model training method of the text processing model described above. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solutions of the data processing method and the model training method of the text processing model described above.

[0159] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0160] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0161] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this specification is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this specification. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this specification.

[0162] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0163] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. These embodiments have been selected and specifically described in this specification to better explain the principles and practical applications of this specification, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A data processing method, characterized in that, include: Get the text to be processed; The text to be processed is input into a text processing model to obtain the response text information output by the text processing model. The text processing model includes a knowledge unit and a text retrieval unit, and the response text information is generated based on the knowledge unit or the text retrieval unit.

2. The method as described in claim 1, characterized in that, The text processing model also includes a routing unit; The text to be processed is input into the text processing model to obtain the response text information output by the text processing model, including: The text to be processed is input into the routing unit to obtain the processing unit determined by the routing unit based on the text to be processed, wherein the processing unit includes the knowledge unit or the text retrieval unit; The text to be processed is input into the knowledge unit or the text retrieval unit to obtain the response text information.

3. The method as described in claim 2, characterized in that, The text to be processed is input into the knowledge unit or the text retrieval unit to obtain response text information, including: The text to be processed is input into the knowledge unit to obtain the response text information output by the knowledge unit; or, The text to be processed is input into the text retrieval unit to obtain the associated text information output by the text retrieval unit, and a response text information is generated based on the associated text information.

4. The method as described in claim 1, characterized in that, The text processing model is trained through the following steps: The domain knowledge text of the target business domain is obtained and input into the initial text processing model to obtain the reference text processing model. The initial text processing model includes a routing unit, a knowledge unit, and a text retrieval unit. Obtain the domain sample text corresponding to the target business domain, wherein the domain sample text includes domain sample questions and domain sample answers; The domain sample question is input into the reference text processing model to obtain the predicted question response output by the reference text processing model; The model reward value is calculated based on the predicted question response and the domain sample response, and the model parameters of the reference text processing model are adjusted based on the model reward value. The reference text processing model is trained again based on the domain sample text until the model training stops, thus obtaining the text processing model.

5. The method as described in claim 4, characterized in that, Obtain domain knowledge text from the target business domain and input the domain knowledge text into the initial text processing model to obtain a reference text processing model, including: Obtain domain knowledge text and an initial text processing model for the target business domain, wherein the initial text processing model includes knowledge units and text retrieval units; The domain knowledge text is input into the initial text processing model to train the text retrieval unit in the initial text processing model, thereby obtaining a reference text processing model.

6. The method as described in claim 5, characterized in that, The domain knowledge text includes domain knowledge statements and domain tag statements; Inputting the domain knowledge text into the initial text processing model and training the text retrieval unit in the initial text processing model includes: The domain knowledge statement and the domain tag statement are input into the text retrieval unit of the initial text processing model to obtain the domain prediction statement generated by the text retrieval unit based on the domain knowledge statement; The retrieval loss value is calculated based on the domain prediction statement and the domain tag statement, and the model parameters of the text retrieval unit are adjusted based on the retrieval loss value.

7. The method as described in claim 4, characterized in that, The domain sample question is input into the reference text processing model to obtain the predicted question response output by the reference text processing model, including: The domain sample question is input into the routing unit of the reference text processing model to obtain the processing unit determined by the routing unit, wherein the processing unit is determined based on whether the domain sample question has a search tag set; When the processing unit is a knowledge unit, the domain sample question is input into the knowledge unit to obtain the predicted question response output by the knowledge unit; When the processing unit is a text retrieval unit, the domain sample question is input into the text retrieval unit to obtain the predicted text information output by the text retrieval unit, and a predicted question response is generated based on the predicted text information.

8. The method as described in claim 7, characterized in that, When the domain sample problem is tagged with search labels; The domain sample question is input into the text retrieval unit to obtain the predicted text information output by the text retrieval unit, and a predicted question response is generated based on the predicted text information, including: The domain sample question is input into the text retrieval unit to obtain the predicted text information output by the text retrieval unit, and a predicted response information is generated based on the predicted text information. A predicted question response is generated based on the predicted text information and the predicted response information.

9. The method as described in claim 8, characterized in that, The domain sample responses include sample text information and sample response information; The model reward value is calculated based on the predicted question response and the domain sample response, including: A text reward value is generated based on the predicted text information and the sample text information; A response reward value is generated based on the predicted response information and the sample response information; A model reward value is generated based on the text reward value and the response reward value.

10. The method as described in claim 4, characterized in that, Obtaining the domain sample text corresponding to the target business domain, including: Obtain a preset number of domain retrieval questions and domain knowledge questions corresponding to the target business domain; Search tags are set for the domain search questions, wherein the search tags indicate that the domain search questions are processed by the text search unit; The domain knowledge question-and-answer questions and the domain retrieval question-and-answer questions with search tags are used as domain sample text.

11. A method for training a text processing model, characterized in that, Applications in cloud-side devices, including: The receiving end device sends domain knowledge text of the target business domain, and inputs the domain knowledge text into the initial text processing model to obtain a reference text processing model, wherein the initial text processing model includes knowledge units and text retrieval units; Obtain the domain sample text corresponding to the target business domain, wherein the domain sample text includes domain sample questions and domain sample answers; The domain sample question is input into the reference text processing model to obtain the predicted question response output by the reference text processing model; The model reward value is calculated based on the predicted question response and the domain sample response, and the model parameters of the reference text processing model are adjusted based on the model reward value. The reference text processing model is trained again based on the domain sample text until the model training stops, and the text processing model is obtained. The model parameters of the text processing model are then sent to the edge device.

12. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 11.

13. A computer-readable storage medium storing a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 11.

14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 11.