Data processing system based on large model

By generating and expanding information on the original question, the problem of inaccurate search results in existing search enhancement generation techniques is solved, achieving more accurate and comprehensive search result generation and enhancing the answer generation capability of large models.

CN121833869APending Publication Date: 2026-04-10WUHAN TCL CORP RES CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN TCL CORP RES CO LTD
Filing Date
2024-10-08
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing search enhancement generation techniques directly retrieve relevant documents based on the original question, leading to inaccurate search results.

Method used

By acquiring the first problem information to be processed, information generation processing is performed on it to generate the second problem information to be processed, and retrieval is performed based on this information to determine the target processing result information. A large model is used to generate and match text information to improve retrieval accuracy.

Benefits of technology

It expands the scope of retrieval, improves the accuracy and coverage of retrieval results, enhances the reasoning robustness of large models, and ensures the accuracy and comprehensiveness of answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833869A_ABST
    Figure CN121833869A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing system based on a large model, and the system determines target processing result information based on second to-be-processed problem information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically to a data processing system based on a large model. Background Technology

[0002] Retrieval-augmented generation (RAG) generates answers or content by incorporating additional information sources such as external knowledge bases, effectively mitigating the illusion problem caused by existing language models' lack of understanding of new knowledge. However, existing RAG techniques typically perform relevant document searches directly based on the original question, leading to inaccurate search results. Summary of the Invention

[0003] This application provides a data processing system based on a large model.

[0004] In a first aspect, this application provides a method comprising:

[0005] Obtain information on the first pending issue;

[0006] The first problem information to be processed is processed to generate the second problem information to be processed.

[0007] Based on the information of the second problem to be processed, the target processing result information is determined.

[0008] Secondly, this application provides a system comprising:

[0009] The information acquisition module is used to acquire information about the first problem to be processed;

[0010] The information generation module is used to generate information from the first problem information to obtain the second problem information.

[0011] The information determination module is used to determine the target processing result information based on the second problem information to be processed.

[0012] Thirdly, this application also provides a computer device, which includes:

[0013] One or more processors;

[0014] Memory; and

[0015] One or more applications, wherein the applications are stored in memory and configured to be executed by a processor to implement the methods of any one of the first aspects.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps of the method in any of the first aspects. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of a data processing system provided in an embodiment of the present invention;

[0019] Figure 2 This is a flowchart of one embodiment of the data processing method provided by the present invention;

[0020] Figure 3 This is a flowchart illustrating a specific embodiment of the information generation and processing of the first problem information provided by the present invention.

[0021] Figure 4 This is a flowchart illustrating a specific embodiment of the method for determining target processing result information provided in this invention.

[0022] Figure 5 This is a schematic block diagram of the data processing system provided in the embodiments of the present invention;

[0023] Figure 6 This is a schematic diagram of an embodiment of the computer device provided in this invention. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0025] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on this application. Furthermore, the terms "first," "second," and "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, features defined with "first," "second," "third," etc., may explicitly or implicitly include one or more features. In the description of this application, "several" means at least one, unless otherwise explicitly specified.

[0026] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0027] It should be noted that since the method in this application embodiment is executed in a computer device, the processing objects of each computer device exist in the form of data or information, such as time, which is essentially time information. It can be understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data that exist so that the computer device can process them. Specific details will not be elaborated here.

[0028] This application provides a data processing method and system based on a large model, which will be described in detail below.

[0029] Please see Figure 1 , Figure 1 This is a schematic diagram of a data processing system provided in an embodiment of this application. The data processing system may include a computer device 100, such as... Figure 1 Computer equipment in the country.

[0030] In this embodiment, the computer device 100 is mainly used to obtain first problem information to be processed; to perform information generation processing on the first problem information to be processed to obtain second problem information to be processed; and to determine target processing result information based on the second problem information to be processed, which can improve the accuracy and coverage of the search results.

[0031] In this embodiment, the computer device 100 can be a standalone server, a server network, or a server cluster. For example, the computer device 100 described in this embodiment includes, but is not limited to, a computer, a network host, a single network server, a set of multiple network servers, or a cloud server composed of multiple servers. The cloud server is composed of a large number of computers or network servers based on cloud computing.

[0032] It is understood that the computer device 100 used in the embodiments of this application can be a device that includes both receiving and transmitting hardware, that is, a device having receiving and transmitting hardware capable of performing bidirectional communication on a bidirectional communication link. Such a device may include: cellular or other communication devices having a single-line display, a multi-line display, or a cellular or other communication device without a multi-line display. Specifically, the computer device 100 may be a desktop terminal or a mobile terminal, and may also be one of a mobile phone, tablet computer, laptop computer, etc.

[0033] Those skilled in the art will understand that Figure 1 The application environment shown is merely one application scenario of the solution in this application and does not constitute a limitation on the application scenario of the solution in this application. Other application environments may include those that are more specific to this application. Figure 1 The number of computer devices shown is more or less, for example Figure 1 Only one computer device is shown in the diagram. It is understood that the data processing system may also include one or more other services, which are not limited here.

[0034] In addition, such as Figure 1 As shown, the data processing system may also include a memory 200 for storing data, such as association information, such as first association information, second association information, etc., and evaluation information, such as first evaluation information, second evaluation information, etc.

[0035] It should be noted that, Figure 1The schematic diagram of the data processing system shown is merely an example. The data processing system and scenario described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of data processing systems and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0036] like Figure 2 The diagram shown is a flowchart of an embodiment of the data processing method in this application. The data processing method may include the following steps S201 to S203, as detailed below:

[0037] S201. Obtain information on the first pending issue.

[0038] The first problem information to be processed is the information acquired by the computer device that needs to be processed. The first problem information to be processed varies depending on the application scenario of the data processing method. For example, when the data processing method is applied to a question-and-answer system, the first problem information to be processed is the question information input by the user. When the data processing method is applied to document generation, the first problem information to be processed can be the prompt words or prompt statements input by the user. This embodiment does not limit this.

[0039] Furthermore, the computer device can obtain the first problem information to be processed in a variety of ways. For example, the computer device can collect the user's voice information through its own microphone and convert the voice information into text information to obtain the first problem information to be processed. The computer device can also receive the first problem information to be processed input by the user through input devices such as keyboard, mouse, and touch screen. The computer device can also obtain the first problem information to be processed from other devices through networks, Bluetooth, etc. This embodiment does not limit these methods.

[0040] The first question information to be processed obtained by the computer device can be text information in any language such as Chinese, English, Japanese, etc. This embodiment does not limit it. For example, the first question information to be processed can be "What is the capital of France?", or it can be "What is the capital of France?".

[0041] S202. Perform information generation processing on the first problem information to be processed to obtain the second problem information to be processed.

[0042] In this embodiment, information generation processing refers to the process of further expanding the query scope by adding information related to the first problem information to be processed. The second problem information to be processed is the information obtained after information generation processing of the first problem information, and the second problem information to be processed includes several pieces of first problem information. This embodiment obtains the second problem information to be processed by information generation processing of the first problem information, and then performs data processing based on the second problem information. This can expand the scope of the search and avoid the problem of incomplete and inaccurate search results caused by ambiguity, unclear description, or inability to fully reflect the content to be searched in the first problem information. It can improve the accuracy and coverage of the search results.

[0043] In one specific implementation, such as Figure 3 As shown, step S202 above involves generating information from the first problem information to obtain the second problem information. This can include steps S301 to S302, as follows:

[0044] S301. Based on the information of the first problem to be processed, determine the first candidate text information.

[0045] In this embodiment, the first candidate text information is text information related to the content and format of the answer information corresponding to the first question to be processed. The first candidate text information includes multiple question information associated with the first question to be processed and the answer information corresponding to each question information. For example, as shown in Table 1, the first question to be processed is "How is the ohmic contact layer prepared during the fabrication of amorphous silicon TFTs?", and the first candidate text information includes four question information associated with the first question to be processed: "What is an amorphous silicon TFT?", "What is an ohmic contact layer?", "Why is an ohmic contact layer needed when fabricating an amorphous silicon TFT?", and "What are the steps for preparing the ohmic contact layer during the fabrication of amorphous silicon TFTs?", and the answer information corresponding to these four question information.

[0046] Table 1 Information on the First and Second Problems to be Processed

[0047]

[0048]

[0049] In one specific embodiment, the step of determining the first candidate text information based on the first problem information to be processed specifically includes: performing fusion processing on the first problem information to be processed and the prompt information to obtain the first fused information; and using the first information processing model to perform text information generation processing on the first fused information to obtain the first candidate text information.

[0050] The prompt information is used to guide the first information processing model to output the first candidate text information. The prompt information can be pre-set or constructed based on the first question information to be processed; this embodiment does not limit this. For example, the prompt information could be: "Please list the semiconductor display field background knowledge required to answer the following question step by step, and provide the corresponding answers. {question} Your task is not to answer..." <query>Instead of presenting the question, it lists the background knowledge required to answer it in a question-and-answer format. The format is as follows: question1 content\n answer1 content\n...questionk content\n answerk content\n".

[0051] In one specific embodiment, the first information processing model can be constructed by combining a generative pre-trained (GPT) model and a chain-of-thought (CoT) technique, which is a widely used hinting method that can significantly enhance the reasoning ability of large models by adding intermediate reasoning steps.

[0052] S302. Based on the first candidate text information and the first problem information to be processed, determine the second problem information to be processed.

[0053] In one specific embodiment, the step of determining the second problem information to be processed based on the first candidate text information and the first problem information to be processed specifically includes: performing text segmentation processing on the first candidate text information to obtain a plurality of second candidate text information; and performing fusion processing on the plurality of second candidate text information and the first problem information to be processed to obtain the second problem information to be processed.

[0054] Text segmentation refers to the process of dividing the first candidate text information into multiple text fragments according to specified rules or tags. Several second candidate text information fragments are the text fragments segmented from the first candidate text information. As mentioned in the preceding steps, the first candidate text information includes multiple question information associated with the first question information to be processed and the answer information corresponding to each question. When performing text segmentation on the first candidate text information, it can be performed according to the multiple question information and the answer information corresponding to each question. For example, referring to Table 1, the first candidate text information includes 4 question information associated with the first question information to be processed and 4 answer information corresponding to the questions. Performing text segmentation on the first candidate text information can yield 8 second candidate text information fragments (i.e., fragments 1 to 8 in Table 1).

[0055] Furthermore, the fusion processing of several second candidate text information and first problem information can be performed by merging several second candidate text information and first problem information, by splicing several second candidate text information and first problem information, or by performing union operation on several second candidate text information and first problem information. This embodiment does not limit the specific processing.

[0056] In one specific embodiment, fusing several second candidate text information and first problem information to be processed means merging the several second candidate text information and the first problem information to be processed. Therefore, the resulting second problem information to be processed includes several second candidate text information and first problem information to be processed. For example, performing text segmentation processing on the first candidate text information in Table 1 can yield 8 second candidate text information. Fusing the 8 second candidate text information and the first problem information to be processed can yield the second problem information to be processed, as shown in Table 1, which includes 9 pieces of first problem information, from fragment 1 to fragment 9.

[0057] S203. Based on the information of the second problem to be processed, determine the target processing result information.

[0058] In this embodiment, the target processing result information is the search text information corresponding to the first problem information to be processed, which is determined based on the second problem information to be processed. Since the second problem information to be processed is the information obtained by information generation processing of the first problem information to be processed, information retrieval based on the second problem information to be processed can expand the scope of retrieval and avoid the problem that the retrieval results are not comprehensive and accurate due to ambiguity, unclear expression or inability to fully reflect the content to be retrieved in the first problem information to be processed. This can improve the accuracy and coverage of the retrieval results.

[0059] In one specific implementation, such as Figure 4 As shown, the second problem information to be processed includes several first problem information. The determination of the target processing result information based on the second problem information to be processed in step S203 above may include steps S401 to S402, as follows:

[0060] S401. Based on several pieces of information about the first problem, determine the first related information.

[0061] In one specific embodiment, the second problem information to be processed includes several first problem information items. For example, continuing to refer to Table 1, the second problem information to be processed includes nine first problem information items, namely fragment 1 to fragment 9. The first associated information is the search text information obtained by text retrieval based on several first problem information items. This embodiment performs text retrieval based on several first problem information items. Compared with the existing method of directly performing text retrieval based on first problem information to be processed, this can avoid the problem of incomplete and inaccurate search results caused by ambiguity, unclear expression, or inability to fully reflect the content to be searched in the first problem information to be processed. This can improve the accuracy and coverage of the search results.

[0062] In one specific embodiment, the step of determining the first association information based on a plurality of first problem information specifically includes: for any first problem information among the plurality of first problem information, determining the target evaluation information between each of the plurality of second association information and the first problem information; filtering the plurality of second association information based on the target evaluation information to obtain the third association information corresponding to the first problem information; and fusing the third association information corresponding to the plurality of first problem information to obtain the first association information.

[0063] In this embodiment, several second-related information items are text information in a knowledge base, and the third-related information items are second-related information items among the several second-related information items that satisfy the first condition. Satisfying the first condition can be that the evaluation value corresponding to the target evaluation information is greater than the evaluation threshold, or that the evaluation value corresponding to the target evaluation information is the highest among a predetermined number of evaluation values ​​corresponding to several second-related information items, or that the evaluation value corresponding to the target evaluation information is within the evaluation value range. This embodiment does not impose any limitations on this. For example, several second-related information items include text information 1 to text information 100. The target evaluation information between several second-related information items and the first question information A is sorted from high to low as text information 15 > text information 20 > text information 48 > text information 10 > text information 60 > ... The third-related information corresponding to the first question information A is the five second-related information items with the highest evaluation values, sorted from high to low according to the target evaluation information. Therefore, the third-related information corresponding to the first question information A includes text information 15, text information 20, text information 48, text information 10, and text information 60.

[0064] In one specific embodiment, each first problem information has a corresponding third associated information. The third associated information corresponding to different first problem information can be the same or different. The fusion processing of the third associated information corresponding to several first problem information can be a union operation on the third associated information corresponding to several first problem information, or it can be a merging processing of the third associated information corresponding to several first problem information. This embodiment does not limit this.

[0065] In one specific embodiment, fusing the third associated information corresponding to several first problem information refers to performing a union operation on the third associated information corresponding to several first problem information. For example, the third associated information corresponding to first problem information A includes text information 15, text information 20, text information 48, text information 10, and text information 60; the third associated information corresponding to first problem information B includes text information 15, text information 12, text information 48, text information 11, and text information 65; and the third associated information corresponding to first problem information C includes text information 31, text information 12, text information 45, text information 11, and text information 65. Then, by fusing the third associated information corresponding to first problem information A, first problem information B, and first problem information C, the resulting first associated information includes text information 15, text information 20, text information 48, text information 10, text information 60, text information 12, text information 11, text information 65, text information 31, and text information 45.

[0066] In one specific embodiment, the step of determining the target evaluation information between each of the several second related information and the first problem information for any one of the several first problem information specifically includes: for any one of the several first problem information, performing similarity calculation on the first problem information and the several second related information to obtain first evaluation information between each second related information and the first problem information; performing information matching processing on the first problem information and the several second related information to obtain second evaluation information between each second related information and the target problem information; and performing weighted summation processing on the first evaluation information and the second evaluation information to obtain target evaluation information between each second related information and the first problem information.

[0067] This embodiment calculates the similarity between the first problem information and several second related information to obtain the first evaluation information, performs information matching processing on the first problem information and several second related information to obtain the second evaluation information, and performs weighted summation processing on the first evaluation information and the second evaluation information to obtain the target evaluation information. The combination of similarity calculation and information matching can comprehensively evaluate the degree of association between each first problem information and each second related information, thereby improving the accuracy of the obtained first related information.

[0068] In one specific embodiment, there are multiple ways to calculate the similarity between the first question information and several second related information. For example, the similarity between the first question information and several second related information can be calculated using a cosine similarity calculation method, the similarity between the first question information and several second related information can be calculated using a Jaccard similarity calculation method, or the similarity between the first question information and several second related information can be calculated using an Euclidean distance calculation method. This embodiment does not limit the methods.

[0069] In one specific embodiment, the step of performing information matching processing on the first problem information and several second related information to obtain second evaluation information between each second related information and the first problem information specifically includes: extracting features from the first problem information to obtain feature word information of the first problem information; determining third evaluation information for each feature word based on the frequency of each feature word in each second related information and the length information of each second related information; determining first weight information for each feature word based on the number of second related information containing each feature word and the number of second related information; and performing weighted summation processing on the third evaluation information of several feature words based on the first weight information to obtain second evaluation information between each second related information and the first problem information.

[0070] The length information represents the number of characters contained in the second associated information. The number of second associated information containing each feature word refers to the number of second associated information containing each feature word in several second associated information. The total number of second associated information refers to the total number of several second associated information. Feature word information includes several feature words, which are words or phrases in the first question information that can describe the content of the first question information. In order to extract features from the first question information, a feature extraction function can be predefined. For a given first question information, the feature word extraction function can extract keyword information from the first question information by identifying nouns and proper nouns in the first question information. The feature extraction process of the feature extraction function can be expressed as: K(q)={k1,k2,...,kn}. For example, if the first question information is "What is the capital of France?", then the extracted feature word information is ["capital","France"].

[0071] In one specific embodiment, the process of determining the third evaluation information can be represented as follows: R(q i D) represents the feature word q i The third evaluation information in the second related information D, f(q) i D) represents the feature word q i The number of times the second associated information D appears is used, where |D| represents the length of the second associated information D, and avgdl represents the average length of several second associated information items.

[0072] Furthermore, the process of determining the first weight information can be expressed as: The process of determining the second assessment information can be represented as follows: Wherein, Score1 represents the second evaluation information between the second correlation information D and the first problem information, W i Represents the i-th feature word q i The first weight information, R(q) i D) represents the i-th feature word q i The third evaluation information in the second related information D, where N represents the total number of second related information, n(q) i ) indicates that the i-th feature word q is included. i The number of second-related information.

[0073] S402. Based on several first problem information and first related information, determine the target processing result information.

[0074] As can be seen from the aforementioned step S401, the first associated information is related text information obtained by retrieving knowledge base information based on several first question information. However, the relevance between the several first question information and the first question information to be processed varies. Therefore, the degree of association between the associated information in the first associated information and the first question information to be processed also varies. To further improve the accuracy of the obtained target processing result information, this embodiment, after obtaining the first associated information, filters the first associated information based on several first question information to obtain the target processing result information corresponding to the first question information to be processed.

[0075] In one specific embodiment, the first association information includes several association information items. The step of determining the target processing result information based on several first problem information items and the first association information specifically includes: calculating the relevance of several first problem information items and first problem information items to be processed to obtain second weight information corresponding to each first problem information item; performing weighted summation processing on the target evaluation information between each association information item and several first problem information items based on the second weight information to obtain fourth evaluation information corresponding to each association information item; and performing filtering processing on several association information items based on the fourth evaluation information to obtain the target processing result information.

[0076] The process of determining the target evaluation information between several first problem information and each associated information has been discussed in the aforementioned step S401. For details, please refer to the discussion in the aforementioned step S401. To avoid repetition, this embodiment will not repeat it here.

[0077] In one specific embodiment, the step of calculating the relevance of several first problem information and first problem information to obtain the second weight information corresponding to each first problem information specifically includes: determining several model input information based on several first problem information and first problem information to be processed; each model input information includes one first problem information and one first problem information to be processed from the several first problem information; inputting the several model input information into a third information processing model, and outputting the relevance information between each first problem information and the first problem information to be processed through the third information processing model; and determining the relevance information between each first problem information and the first problem information to be processed as the second weight information corresponding to each first problem information.

[0078] The third information processing model can be built based on the BCEranker model, which is a re-ranking model specifically used in the fields of information retrieval and natural language processing. The BCEranker model takes an information pair consisting of a first question information and a first question information to be processed as input and outputs a score, which represents the relevance between the first question information and the first question information to be processed.

[0079] In one specific embodiment, the process of determining the fourth evaluation information can be represented as follows: Score t D represents the t-th related information. t The corresponding fourth assessment information, R i S(K) represents the second weight information corresponding to the i-th first question information. i D t ) represents the i-th first problem information K i With the t-th related information D t The target evaluation information between them, where m represents the amount of information for the first question.

[0080] In one specific embodiment, when filtering several pieces of related information based on the fourth evaluation information, the related information can be sorted from high to low according to the fourth evaluation information, and the related information ranked first can be determined as the target processing result information. Of course, when filtering several pieces of related information based on the fourth evaluation information, the fourth evaluation information corresponding to each piece of related information can also be compared with a preset threshold, and the related information whose fourth evaluation information is greater than the preset threshold can be determined as the target processing result information. This embodiment does not limit this.

[0081] In one specific implementation, after determining the target processing result information based on the second problem information to be processed in step S203 above, the following steps may be included: fusing the first problem information to be processed and the target processing result information to obtain the second fused information; and using the second information processing model to perform text information generation processing on the second fused information to obtain the first target text information.

[0082] In this embodiment, the first target text information is text information obtained by information generation processing based on the first question information to be processed and the target processing result information. For example, the first question information to be processed is the question information input by the user, and the first target text information is the answer information corresponding to the question information generated based on the first question information to be processed and the target processing result information. This embodiment performs information generation processing on the first question information to be processed to obtain the second question information to be processed, determines the target processing result information based on the second question information to be processed, and then generates the answer information corresponding to the question information based on the first question information to be processed and the target processing result information. This can avoid the problem of incomplete and inaccurate search results caused by ambiguity, unclear expression, or inability to fully reflect the content to be searched in the first question information to be processed, and improve the accuracy of the generated first target text information.

[0083] In one specific embodiment, the second information processing model can be built based on a Convolutional Neural Network (CNN) model, a Dynamic Graph Convolutional Neural Network (DGCNN) model, or a Large Language Model (LLM) model. This embodiment of the invention does not limit the specific model.

[0084] In one specific embodiment, after the above-mentioned text information generation processing of the second fused information using the second information processing model to obtain the first target text information, the method further includes: performing matching analysis processing on the first target text information and the benchmark information; if the first target text information matches the benchmark information, performing information retrieval processing based on the first problem information to be processed to obtain the fourth related information; parsing the fourth related information to obtain the fifth related information; and performing text information generation processing on the fifth related information and the first problem information to be processed to obtain the second target text information.

[0085] The benchmark information is pre-set data used to measure whether the second information processing model can infer the answer information corresponding to the first question information based on the first question information and the target processing result information. When the first target text information does not match the benchmark information, it indicates that the second information processing model can infer the answer information corresponding to the first question information based on the first question information and the target processing result information. In this case, the first target text information is the answer information corresponding to the first question information, and the first target text information can be directly output. Conversely, when the first target text information matches the benchmark information, it indicates that the second information processing model cannot infer the answer information corresponding to the first question information based on the first question information and the target processing result information. In this case, information retrieval processing is performed based on the first question information to obtain the fourth related information, parsing processing is performed on the fourth related information to obtain the fifth related information, and text information generation processing is performed on the fifth related information and the first question information to obtain the second target text information. For example, if the preset benchmark information is "cannot obtain results through documents", and the first target text information is "cannot obtain results", then it is determined that the first target text information matches the benchmark information.

[0086] In this embodiment, when the first target text information matches the benchmark information, information retrieval processing is performed based on the first question information to be processed to obtain the fourth related information. The fourth related information is parsed to obtain the fifth related information. Text information generation processing is performed on the fifth related information and the first question information to be processed to obtain the second target text information. If the second information processing model cannot infer the answer information corresponding to the first question information to be processed based on the target processing result information, the fifth related information obtained from the network search can be used as a supplementary document to infer the answer information corresponding to the first question information to be processed, thereby improving the robustness of the large model inference.

[0087] In one specific embodiment, the step of performing information retrieval processing based on the first problem information to obtain the fourth related information specifically includes: determining the representation of information to be retrieved based on the first problem information; and performing information retrieval processing based on the representation of information to be retrieved to obtain the fourth related information. The representation of information to be retrieved consists of search keywords dynamically generated based on the first problem information, and the fourth related information consists of data related to the first problem information crawled from a predetermined data source or website based on the representation of information to be retrieved. The predetermined data source may include one or more of specific news websites, academic paper databases, government databases, social media platforms, etc., and this embodiment does not limit this.

[0088] Furthermore, regular expressions can be used to parse and process the fourth association information, or Python libraries (e.g., BeautifulSoup) can be used to parse and process the fourth association information, or C language-based libraries (e.g., lxml) can be used to parse and process the fourth association information. This embodiment does not limit the specific methods.

[0089] In one specific embodiment, the second target text information is text information obtained by text information generation processing based on the fifth associated information and the first question information to be processed. For example, the first question information to be processed is the question information input by the user, and the second target text information is the answer information corresponding to the question information generated based on the fifth associated information and the first question information to be processed. The step of performing information generation processing on the fifth associated information and the first question information to obtain the second target text information corresponding to the first question information to be processed specifically includes: performing fusion processing on the fifth associated information and the first question information to be processed to obtain third fused information; and using a second information processing model to perform text information generation processing on the third fused information to obtain the second target text information.

[0090] To verify the effectiveness of the data processing method provided in this invention, the inventors selected 178 objective single-choice test questions related to semiconductors. These 178 questions cover various levels from basic knowledge to advanced applications, including LCD, LCD & OLED, materials technology, product development, process technology, module development, question classification, and display technology. The questions were answered directly using models such as GPT3.5turbo, GPT4, Kimi, ChatGLM4, and Qwen-max, as well as using existing RAG technology and the data processing method provided in this invention. The accuracy rates of the answers were statistically analyzed. Experimental analysis showed that using existing RAG technology improved the accuracy rate by approximately 5% compared to answering directly using models, while using the data processing method provided in this invention improved the accuracy rate by approximately 5% compared to using existing RAG technology. This indicates that using the data processing method provided in this invention can improve the accuracy of the obtained answer information.

[0091] In summary, the data processing method provided in this implementation plan, by acquiring first problem information to be processed, performing information generation processing on the first problem information to be processed to obtain second problem information to be processed, and determining the target processing result information based on the second problem information to be processed, can expand the scope of retrieval compared to existing retrieval enhancement generation techniques. It avoids the problem of incomplete and inaccurate retrieval results caused by ambiguity, unclear description, or inability to fully reflect the content to be retrieved in the first problem information to be processed, thus improving the accuracy and coverage of retrieval results. Furthermore, the first problem information to be processed and the target processing result information are fused to obtain second fused information, and a second information processing model is used to generate text information from the second fused information. The first target text information is obtained through processing, which can improve the accuracy of the generated first target text information. Furthermore, if the first target text information matches the benchmark information, information retrieval processing is performed based on the first question information to be processed to obtain the fourth related information. The fourth related information is parsed to obtain the fifth related information. Text information generation processing is performed on the fifth related information and the first question information to be processed to obtain the second target text information. If the second information processing model cannot infer the answer information corresponding to the first question information to be processed based on the target processing result information, the fifth related information obtained from the network search can be used as a supplementary document to infer the answer information corresponding to the first question information to be processed, thereby improving the robustness of the large model's inference.

[0092] To better implement the data processing method in the embodiments of this application, a data processing system is also provided in the embodiments of this application, such as... Figure 5 As shown, the data processing system 600 includes:

[0093] Information acquisition module 610 is used to acquire information about the first problem to be processed;

[0094] The information generation module 620 is used to generate information from the first problem information to obtain the second problem information.

[0095] The information determination module 630 is used to determine the target processing result information based on the second problem information to be processed.

[0096] In this embodiment, the second problem information is obtained by generating information from the first problem information to be processed, and the target processing result information is determined based on the second problem information to be processed. Compared with the existing search enhancement generation technology, this can expand the scope of the search and avoid the problem that the search results are not comprehensive and accurate due to the ambiguity, unclear expression or inability to fully reflect the content to be searched in the first problem information to be processed. This can improve the accuracy and coverage of the search results.

[0097] In some embodiments of this application, the information generation module 620 performs information generation processing on the first problem information to be processed to obtain the second problem information to be processed, including:

[0098] Based on the information of the first problem to be processed, the first candidate text information is determined;

[0099] Based on the first candidate text information and the first problem information to be processed, the second problem information to be processed is determined.

[0100] In some embodiments of this application, the information generation module 620 determines first candidate text information based on the first problem information to be processed, including:

[0101] The first pending problem information and the prompt information are merged to obtain the first merged information;

[0102] The first information processing model is used to process the first fused information into text information to obtain the first candidate text information.

[0103] In some embodiments of this application, the information generation module 620 determines the second problem information to be processed based on the first candidate text information and the first problem information to be processed, including:

[0104] The first candidate text information is segmented to obtain several second candidate text information;

[0105] Several second candidate text information and first problem information to be processed are fused together to obtain second problem information to be processed.

[0106] In some embodiments of this application, the second problem information to be processed includes several first problem information items. The information determination module 630 determines the target processing result information based on the second problem information to be processed, including:

[0107] Based on several pieces of information about the first question, the first related information is determined;

[0108] Based on several primary problem information and primary related information, the target processing result information is determined.

[0109] In some embodiments of this application, the information determination module 630 determines first related information based on several first problem information, including:

[0110] For any one of several first problem information, determine the target evaluation information between each of several second related information and the first problem information;

[0111] Based on the target evaluation information, several second related information are filtered and processed to obtain the third related information corresponding to the first problem information; the third related information is the second related information that satisfies the first condition of the target evaluation information.

[0112] The third related information corresponding to several first question information is fused to obtain the first related information.

[0113] In some embodiments of this application, the information determination module 630, for any one of a plurality of first problem information, determines target evaluation information between each of a plurality of second related information and the first problem information, including:

[0114] For any one of the several first question information, a similarity calculation is performed between the first question information and several second related information to obtain the first evaluation information between each second related information and the first question information;

[0115] The first problem information and several second related information are matched to obtain second evaluation information between each second related information and the first problem information.

[0116] The first evaluation information and the second evaluation information are weighted and summed to obtain the target evaluation information between each second related information and the first problem information.

[0117] In some embodiments of this application, the information determination module 630 performs information matching processing on the first problem information and several second related information to obtain second evaluation information between each second related information and the first problem information, including:

[0118] Feature extraction is performed on the first question information to obtain the feature word information of the first question information; the feature word information includes several feature words;

[0119] Based on the frequency of each feature word in each second association information and the length of each second association information, the third evaluation information of each feature word is determined;

[0120] Based on the number of second association information items for each feature word, determine the first weight information for each feature word;

[0121] Based on the first weight information, the third evaluation information of several feature words is weighted and summed to obtain the second evaluation information between each second association information and the first question information.

[0122] In some embodiments of this application, the first associated information includes several associated information items. The information determination module 630 determines the target processing result information based on several first problem information items and the first associated information items, including:

[0123] The relevance of several first problem information and first problem to be processed information is calculated to obtain the second weight information corresponding to each first problem information;

[0124] Based on the second weight information, the target evaluation information between each related information and several first question information is weighted and summed to obtain the fourth evaluation information corresponding to each related information.

[0125] Based on the fourth evaluation information, several related information are filtered and processed to obtain the target processing result information.

[0126] In some embodiments of this application, after the information determination module 630 determines the target processing result information based on the second problem information, the information determination module 630 is further configured to:

[0127] The information on the first problem to be processed and the information on the target processing result are fused together to obtain the second fused information;

[0128] The second information processing model is used to process the second fused information into text information to obtain the first target text information.

[0129] In some embodiments of this application, after the information determination module 630 uses a second information processing model to perform text information generation processing on the second fused information to obtain the first target text information, the information determination module 630 is further used to:

[0130] The first target text information is matched and analyzed with the benchmark information;

[0131] If the first target text information matches the benchmark information, information retrieval processing is performed based on the first problem information to be processed to obtain the fourth related information;

[0132] The fourth association information is parsed and processed to obtain the fifth association information;

[0133] The fifth related information and the first problem information to be processed are processed to generate text information to obtain the second target text information.

[0134] This application embodiment also provides a computer device, the computer device including:

[0135] One or more processors;

[0136] Memory; and

[0137] One or more applications, wherein the applications are stored in memory and configured to be executed by a processor from the steps of the data processing method in any of the embodiments described above.

[0138] This application also provides a computer device, such as... Figure 6 As shown, it illustrates a structural schematic diagram of the computer device involved in the embodiments of this application, specifically:

[0139] The computer device may include components such as a processor 801 with one or more processing cores, a memory 802 with one or more computer-readable storage media, a power supply 803, and an input unit 804. Those skilled in the art will understand that... Figure 6 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0140] The processor 801 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 802, and by calling data stored in the memory 802, it performs various functions of the computer device and processes data, thereby providing overall monitoring of the computer device. Optionally, the processor 801 may include one or more processing cores; preferably, the processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 801.

[0141] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 802 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.

[0142] The computer device also includes a power supply 803 that supplies power to the various components. Preferably, the power supply 803 can be logically connected to the processor 801 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 803 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0143] The computer device may also include an input unit 804, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0144] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 801 in the computer device loads the executable files corresponding to the processes of one or more application programs into the memory 802 according to the following instructions, and the processor 801 runs the application programs stored in the memory 802 to realize various functions, as follows:

[0145] Obtain information on the first pending issue;

[0146] The first problem information to be processed is processed to generate the second problem information to be processed.

[0147] Based on the information of the second problem to be processed, the target processing result information is determined.

[0148] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0149] Therefore, embodiments of this application provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any of the data processing methods provided in embodiments of this application. For example, the computer program loaded by the processor can execute the following steps:

[0150] Obtain information on the first pending issue;

[0151] The first problem information to be processed is processed to generate the second problem information to be processed.

[0152] Based on the information of the second problem to be processed, the target processing result information is determined.

[0153] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.

[0154] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.

[0155] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0156] The above provides a detailed description of a data processing method and system based on a large model provided by the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.< / query>

Claims

1. A method, characterized in that, include: Obtain information on the first pending issue; The first problem information to be processed is processed to generate information, thereby obtaining the second problem information to be processed. Based on the second problem information to be processed, the target processing result information is determined.

2. The method according to claim 1, characterized in that, The second problem information to be processed includes several first problem information items. The step of determining the target processing result information based on the second problem information to be processed includes: Based on several pieces of information related to the first question, the first related information is determined; Based on several pieces of the first problem information and the first related information, the target processing result information is determined.

3. The method according to claim 2, characterized in that, The determination of first related information based on several pieces of information about the first problem includes: For any one of the first problem information from a plurality of first problem information, determine the target evaluation information between each of the plurality of second related information and the first problem information; Based on the target evaluation information, several pieces of second related information are filtered to obtain the third related information corresponding to the first problem information; the third related information is the second related information in which the target evaluation information satisfies the first condition. The third associated information corresponding to several of the first problem information is fused to obtain the first associated information.

4. The method according to claim 3, characterized in that, The step of determining the target evaluation information between each of the several second related pieces of information and the first question information for any one of the several first question information includes: For any one of the several first question information, a similarity calculation is performed between the first question information and several second related information to obtain first evaluation information between each second related information and the first question information; The first problem information and several second related information are matched to obtain second evaluation information between each second related information and the first problem information. The first evaluation information and the second evaluation information are weighted and summed to obtain the target evaluation information between each second related information and the first problem information.

5. The method according to claim 4, characterized in that, The step of performing information matching processing on the first problem information and several second related information to obtain second evaluation information between each second related information and the first problem information includes: Feature extraction is performed on the first problem information to obtain the feature word information of the first problem information; the feature word information includes several feature words; Based on the frequency of each feature word in each of the second association information and the length of each of the second association information, a third evaluation information for each feature word is determined; Based on the number of second association information containing each of the feature words and the number of second association information, a first weight information for each of the feature words is determined; Based on the first weight information, the third evaluation information of several feature words is weighted and summed to obtain the second evaluation information between each second association information and the first question information.

6. The method according to claim 3, characterized in that, The first association information includes several association information items. The step of determining the target processing result information based on the several first question information items and the first association information includes: The relevance of several pieces of the first problem information and the first problem information to be processed is calculated to obtain the second weight information corresponding to each piece of the first problem information. Based on the second weight information, the target evaluation information between each of the associated information and several of the first question information is weighted and summed to obtain the fourth evaluation information corresponding to each of the associated information; Based on the fourth evaluation information, several related information items are filtered and processed to obtain the target processing result information.

7. The method according to claim 1, characterized in that, The step of generating information from the first problem information to obtain the second problem information includes: Based on the first problem information to be processed, the first candidate text information is determined; Based on the first candidate text information and the first problem information to be processed, the second problem information to be processed is determined.

8. The method according to claim 7, characterized in that, The step of determining the first candidate text information based on the first problem information to be processed includes: The first problem information to be processed and the prompt information are fused together to obtain the first fused information; The first information processing model is used to process the first fused information into text information to obtain the first candidate text information.

9. The method according to claim 7, characterized in that, The step of determining the second problem information to be processed based on the first candidate text information and the first problem information to be processed includes: The first candidate text information is processed by text segmentation to obtain several second candidate text information; The second candidate text information and the first problem information to be processed are fused together to obtain the second problem information to be processed.

10. The method according to any one of claims 1 to 9, characterized in that, After determining the target processing result information based on the second problem information to be processed, the process includes: The first problem information to be processed and the target processing result information are fused together to obtain the second fused information; The second information processing model is used to process the second fused information into text information to obtain the first target text information.

11. The method according to claim 10, characterized in that, After processing the second fused information using the second information processing model to generate text information and obtain the first target text information, the process includes: The first target text information is matched and analyzed with the benchmark information; If the first target text information matches the benchmark information, information retrieval processing is performed based on the first problem information to be processed to obtain the fourth related information; The fourth association information is parsed and processed to obtain the fifth association information; The fifth related information and the first problem information to be processed are processed to generate text information to obtain the second target text information.

12. A system, characterized in that, include: The information acquisition module is used to acquire information about the first problem to be processed; The information generation module is used to generate information from the first problem information to obtain the second problem information. The information determination module is used to determine the target processing result information based on the second problem information to be processed; Optionally, the information generation module performs information generation processing on the first problem information to be processed to obtain the second problem information to be processed, including: Based on the first problem information to be processed, the first candidate text information is determined; Based on the first candidate text information and the first problem information to be processed, the second problem information to be processed is determined; Optionally, the information generation module determines first candidate text information based on the first problem information to be processed, including: The first problem information to be processed and the prompt information are fused together to obtain the first fused information; The first information processing model is used to perform text information generation processing on the first fused information to obtain the first candidate text information; Optionally, the information generation module determines the second problem information to be processed based on the first candidate text information and the first problem information to be processed, including: The first candidate text information is processed by text segmentation to obtain several second candidate text information; The second candidate text information and the first problem information to be processed are fused together to obtain the second problem information to be processed. Optionally, the second problem information to be processed includes several first problem information items, and the information determining module determines the target processing result information based on the second problem information to be processed, including: Based on several pieces of information related to the first question, the first related information is determined; Based on several pieces of the first problem information and the first related information, the target processing result information is determined; Optionally, the information determining module determines first related information based on several pieces of the first question information, including: For any one of the first problem information from a plurality of first problem information, determine the target evaluation information between each of the plurality of second related information and the first problem information; Based on the target evaluation information, several pieces of second related information are filtered to obtain the third related information corresponding to the first problem information; the third related information is the second related information in which the target evaluation information satisfies the first condition. The third associated information corresponding to several first problem information is fused to obtain the first associated information; Optionally, the information determination module, for any one of the plurality of first problem information, determines target evaluation information between each of the plurality of second related information and the first problem information, including: For any one of the several first question information, a similarity calculation is performed between the first question information and several second related information to obtain first evaluation information between each second related information and the first question information; The first problem information and several second related information are matched to obtain second evaluation information between each second related information and the first problem information. The first evaluation information and the second evaluation information are weighted and summed to obtain the target evaluation information between each second related information and the first problem information. Optionally, the information determination module performs information matching processing on the first problem information and several second related information to obtain second evaluation information between each second related information and the first problem information, including: Feature extraction is performed on the first problem information to obtain the feature word information of the first problem information; the feature word information includes several feature words; Based on the frequency of each feature word in each of the second association information and the length of each of the second association information, a third evaluation information for each feature word is determined; Based on the number of second association information containing each of the feature words and the number of second association information, a first weight information for each of the feature words is determined; Based on the first weight information, the third evaluation information of several feature words is weighted and summed to obtain the second evaluation information between each second association information and the first question information; Optionally, the first associated information includes several associated information items, and the information determination module determines the target processing result information based on several first problem information items and the first associated information items, including: The relevance of several pieces of the first problem information and the first problem information to be processed is calculated to obtain the second weight information corresponding to each piece of the first problem information. Based on the second weight information, the target evaluation information between each of the associated information and several of the first question information is weighted and summed to obtain the fourth evaluation information corresponding to each of the associated information; Based on the fourth evaluation information, several related information items are filtered and processed to obtain the target processing result information; Optionally, after determining the target processing result information based on the second problem-to-be-processed information, the information determining module is further configured to: The first problem information to be processed and the target processing result information are fused together to obtain the second fused information; The second information processing model is used to perform text information generation processing on the second fused information to obtain the first target text information; Optionally, after the information determination module uses the second information processing model to perform text information generation processing on the second fused information to obtain the first target text information, the information determination module is further used to: The first target text information is matched and analyzed with the benchmark information; If the first target text information matches the benchmark information, information retrieval processing is performed based on the first problem information to be processed to obtain the fourth related information; The fourth association information is parsed and processed to obtain the fifth association information; The fifth related information and the first problem information to be processed are processed to generate text information to obtain the second target text information.

13. A computer device, characterized in that, The computer device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to perform the steps of the method according to any one of claims 1 to 11.