Data processing method and device, electronic equipment, storage medium and program product

By splitting and concatenating form data to generate feature data and analyzing it using a large language model, the problem of incomplete information when complex form data is stored in the database is solved, thus improving the accuracy of the large model's responses.

CN120910113APending Publication Date: 2025-11-07BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510894047.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Retrieval Augmentation (RAG) based on large models has poor accuracy when processing complex form data. This is mainly because the information in complex form data is incomplete when it is stored in the database, resulting in the selection of information items that do not contain the optimal information, which in turn leads to a low accuracy rate of the large model's response.

Method used

By splitting the merged cells in the form, the field data from the merged cells is filled into the split cells to generate a second form. The field data of each form data is then concatenated to form feature data. A large language model is used to extract and analyze the feature data to match the target retrieval question and output an accurate answer.

Benefits of technology

It improves the accuracy of the target answer, ensures that the feature data contains all the valid information of each form data, can accurately match the optimal candidate feature data, and improves the answer accuracy of the large model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910113A_ABST
    Figure CN120910113A_ABST
Patent Text Reader

Abstract

The invention relates to a data processing method and device, electronic equipment, a storage medium and a program product. The method comprises the steps that a target retrieval problem is received; a first form corresponding to the target retrieval question is obtained, the first form comprises header information and multiple pieces of form data, and the header information comprises multiple field names; the combined cells in the first form are subjected to cell splitting processing, the cells obtained through splitting are filled with field data included in the combined cells, a second form is obtained, and in the second form, each piece of form data comprises multiple pieces of field data corresponding to the multiple field names respectively; splicing a plurality of pieces of field data included in each piece of form data in the second form to obtain a plurality of pieces of feature data; obtaining multiple pieces of candidate feature data matched with the target retrieval problem from the multiple pieces of feature data; and inputting the target retrieval question and the plurality of candidate feature data into the target analysis model, and outputting a target answer.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of data processing, and particularly relates to a data processing method and device, electronic equipment, storage medium and program product. BACKGROUND

[0002] Retrieval Augmented Generation (RAG) based on a large model usually has poor accuracy when processing complex form data. Generally, form data is obtained from an Application Programming Interface (API) of an online form. Based on the form data, fixed-length segmentation (or segmentation according to a specified delimiter) is performed, and the data is stored in an embedding database. When a user raises a question, the complete question is submitted as a search word to the embedding database to match the top k information items. The k information items, the user question and the prompt word are submitted to the large model to obtain an answer.

[0003] However, in actual operation, it is found that in the above scheme, because the complex form data is incomplete when stored in the database, the k information items screened out are not satisfactory, and the optimal information is not included, thereby causing the accuracy of the answer given by the large model to be low. SUMMARY

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a data processing method and device, electronic equipment, storage medium and program product.

[0005] In a first aspect of the embodiments of the present disclosure, a data processing method is provided, which includes: receiving a target retrieval question; obtaining a first form corresponding to the target retrieval question, the first form including table header information and a plurality of form data, the table header information including a plurality of field names, and the first form including a merged cell; performing cell splitting processing on the merged cell in the first form, and filling field data included in the merged cell into each cell obtained by splitting to obtain a second form, in which each form data includes a plurality of field data corresponding to the plurality of field names; splicing the plurality of field data included in each form data in the second form to obtain a plurality of feature data, the plurality of feature data corresponding one-to-one to the plurality of form data; obtaining a plurality of candidate feature data matched with the target retrieval question from the plurality of feature data; inputting the target retrieval question and the plurality of candidate feature data into a target analysis model to output a target answer; and the target analysis model is configured to analyze the target retrieval question and the plurality of candidate feature data to obtain the target answer matched with the target retrieval question.

[0006] In some embodiments of the present disclosure, before the obtaining the plurality of candidate feature data matching the target retrieval question from the plurality of feature data, the method further comprises: performing feature extraction on the target retrieval question based on the plurality of field names to obtain target feature data, the target feature data being obtained by splicing field data in the target retrieval question matching the plurality of field names respectively based on the plurality of field names, the plurality of field names corresponding to the plurality of field data one by one; and the obtaining the plurality of candidate feature data matching the target retrieval question from the plurality of feature data comprises: obtaining the plurality of candidate feature data matching the target feature data from the plurality of feature data.

[0007] In some embodiments of the present disclosure, the performing feature extraction on the target retrieval question based on the plurality of field names to obtain target feature data comprises: inputting the target retrieval question and a first prompt word into a first large language model to output the target feature data; wherein the first prompt word is used to instruct the first large language model to extract and splice field data in the target retrieval question matching the plurality of field names respectively to obtain the target feature data.

[0008] In some embodiments of the present disclosure, each feature data is obtained by cross-splicing the plurality of field data and the plurality of field names according to the order of one field name and field data corresponding to the one field name; and the target feature data is obtained by cross-splicing field data in the target retrieval question matching the plurality of field names and the plurality of field names according to the order of one field name and field data corresponding to the one field name.

[0009] In some embodiments of the present disclosure, the arrangement order of the plurality of field names in the target feature data is the same as the arrangement order of the plurality of field names in each feature data.

[0010] In some embodiments of the present disclosure, the target analysis model is a second large language model; and the inputting the target retrieval question and the plurality of candidate feature data into the target analysis model to output the target answer comprises: inputting the target retrieval question, the plurality of candidate feature data and a second prompt word into the second large language model to output the target answer; wherein the second prompt word is used to instruct the second large language model to analyze the target retrieval question and the plurality of candidate feature data to obtain the target answer.

[0011] In a second aspect, the present disclosure provides a data processing apparatus, comprising: a receiving module configured to receive a target search question; an obtaining module configured to obtain a first form corresponding to the target search question, the first form comprising form header information and a plurality of form data, the form header information comprising a plurality of field names, and the first form comprising a merged cell; a processing module configured to perform cell splitting processing on the merged cell in the first form, fill field data included in the merged cell into each cell obtained by splitting, and obtain a second form, wherein each form data in the second form comprises a plurality of field data corresponding to the plurality of field names respectively; splice the plurality of field data included in each form data in the second form to obtain a plurality of feature data, the plurality of feature data corresponding to the plurality of form data one by one; the obtaining module is further configured to obtain a plurality of candidate feature data matched with the target search question from the plurality of feature data; and an output module configured to input the target search question and the plurality of candidate feature data into a target analysis model, and output a target answer; the target analysis model is configured to analyze the target search question and the plurality of candidate feature data, and obtain the target answer matched with the target search question.

[0012] In some embodiments of the present disclosure, the apparatus further comprises a feature extraction module configured to, before obtaining the plurality of candidate feature data matched with the target search question from the plurality of feature data, perform feature extraction on the target search question based on the plurality of field names, and obtain target feature data, the target feature data being obtained by splicing field data matched with the plurality of field names in the target search question based on the plurality of field names, the plurality of field names corresponding to the plurality of field data one by one; and the obtaining module is specifically configured to obtain the plurality of candidate feature data matched with the target feature data from the plurality of feature data.

[0013] In some embodiments of the present disclosure, the feature extraction module is specifically configured to input the target search question and a first prompt word into a first large language model, and output the target feature data; wherein the first prompt word is used to instruct the first large language model to extract and splice the field data matched with the plurality of field names in the target search question to obtain the target feature data.

[0014] In some embodiments of the present disclosure, each feature data is obtained by cross-splicing in the order of one field name and field data corresponding to the one field name based on the plurality of field data and the plurality of field names; and the target feature data is obtained by cross-splicing in the order of one field name and field data corresponding to the one field name based on the field data matched with the plurality of field names in the target search question and the plurality of field names.

[0015] In some embodiments of the present disclosure, the arrangement order of the plurality of field names in the target feature data is the same as the arrangement order of the plurality of field names in each piece of feature data.

[0016] In some embodiments of the present disclosure, the acquisition module is further configured to acquire the form before acquiring the plurality of pieces of candidate feature data matching the target search question from the plurality of pieces of feature data; and splice the plurality of field data included in each piece of form data in the form and the plurality of field names to obtain corresponding feature data, so as to generate the plurality of pieces of feature data.

[0017] In some embodiments of the present disclosure, the device further comprises a splitting module configured to, before splicing the plurality of field data included in each piece of form data in the form and the plurality of field names to obtain corresponding feature data, perform cell splitting processing on the merged cells in the form in the case that the form includes the merged cells.

[0018] In some embodiments of the present disclosure, the target analysis model is a second large language model; and the output module is specifically configured to input the target search question, the plurality of pieces of candidate feature data and a second prompt word into the second large language model to output the target answer; wherein the second prompt word is used to instruct the second large language model to analyze the target search question and the plurality of pieces of candidate feature data to obtain the target answer.

[0019] In a third aspect, an electronic device is provided. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the data processing method according to the first aspect is implemented.

[0020] In a fourth aspect, a computer readable storage medium is provided. The computer readable storage medium stores a computer program. When the computer program is executed by a processor, the data processing method according to the first aspect is implemented.

[0021] In a fifth aspect, a computer program product is provided. The computer program product includes a computer program. When the computer program product is executed on a processor, the processor executes the computer program to implement the data processing method according to the first aspect.

[0022] In a sixth aspect, a chip is provided. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to execute program instructions to implement the data processing method according to the first aspect.

[0023] Compared with the prior art, the technical scheme provided by the embodiments of the present disclosure has the following advantages: a target retrieval question is received; the target retrieval question is received; a first form corresponding to the target retrieval question is obtained, the first form includes table header information and multiple pieces of form data, the table header information includes multiple field names, and there is a merged cell in the first form; cell splitting processing is performed on the merged cell in the first form, and field data included in the merged cell is filled into each cell obtained by splitting to obtain a second form, in which each piece of form data includes multiple field data corresponding to the multiple field names respectively; the multiple field data included in each piece of form data in the second form is spliced to obtain multiple pieces of feature data, which correspond one-to-one to the multiple pieces of form data; multiple pieces of candidate feature data matched with the target retrieval question are obtained from the multiple pieces of feature data; the target retrieval question and the multiple pieces of candidate feature data are input into a target analysis model, and a target answer is output; the target analysis model is used for analyzing the target retrieval question and the multiple pieces of candidate feature data to obtain the target answer matched with the target retrieval question. In the embodiments of the present disclosure, by splitting the merged cell in the form and splicing all field data included in each piece of form data after splitting to obtain multiple pieces of feature data, all valid information in each piece of form data can be included in the feature data, and then the optimal candidate feature data that can accurately answer the target retrieval question can be matched from the multiple pieces of feature data, so that the target analysis model can be used to analyze the target retrieval question and the multiple pieces of candidate feature data, and the target answer matched with the target retrieval question can be obtained, thereby improving the accuracy of the target answer. BRIEF DESCRIPTION OF DRAWINGS

[0024] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced as follows, and obviously, other drawings can also be obtained by those skilled in the art without creative labor.

[0026] Figure 1 One of the flowcharts of the data processing method provided by the embodiments of the present disclosure;

[0027] Figure 2 The second flowchart of the data processing method provided by the embodiments of the present disclosure;

[0028] Figure 3 The third flowchart of the data processing method provided by the embodiments of the present disclosure;

[0029] Figure 4 A structural block diagram of a data processing apparatus provided by an embodiment of the present disclosure is shown in FIG. 1.

[0030] Figure 5 A structural block diagram of an electronic device provided by an embodiment of the present disclosure is shown in FIG. 1. DETAILED DESCRIPTION

[0031] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0032] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present disclosure, but the present disclosure can also be implemented in other manners different from those described herein; obviously, the embodiments described in the specification are only a part of the embodiments of the present disclosure, and not all the embodiments.

[0033] The terms "first", "second", and the like in the specification and claims of the present disclosure are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of a kind and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally means that the objects before and after are in an "or" relationship.

[0034] The electronic device in the embodiments of the present disclosure can be an electronic device, such as a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc., or a server or a server cluster, etc.; the embodiments of the present disclosure are not specifically limited.

[0035] The execution subject of the data processing method provided by the embodiments of the present disclosure can be the electronic device described above, or a functional module and / or functional entity (such as a question and answer system or platform, etc.) capable of implementing the data processing method in the electronic device; specifically, it can be determined according to actual use requirements, and the embodiments of the present disclosure are not limited.

[0036] The data processing method provided by the embodiments of the present disclosure will be described in detail below in combination with the accompanying drawings, specific embodiments and application scenarios thereof.

[0037] As shown in Figure 1 The embodiments of the present disclosure provide a data processing method, which can include the following steps 101 to 106.

[0038] 101. Receiving a target retrieval question.

[0039] The target retrieval question can be a question related to form data in a form.

[0040] The target retrieval question can be a question in voice form or a question in text form, which is not limited here.

[0041] If the target retrieval question is a question in voice form, voice recognition needs to be performed on the target retrieval question after receiving the target retrieval question to obtain the text content corresponding to the target retrieval question, which is not limited here.

[0042] The target retrieval question can be a question input by a user directly received by an electronic device, or a question received from another electronic device, which is not limited here.

[0043] 102. Obtaining a first form corresponding to the target retrieval question.

[0044] The first form includes header information and multiple pieces of form data, the header information includes multiple field names, and the first form includes a merged cell.

[0045] The retrieval question can correspond to one form, or different types of retrieval questions can correspond to different forms (the form corresponding to the target retrieval question can be determined according to the type to which the target retrieval question belongs), which can be determined according to actual conditions, which is not limited here.

[0046] 103. Performing cell splitting processing on the merged cell in the first form, and filling the field data included in the merged cell into each cell obtained by splitting to obtain a second form.

[0047] In the second form, each piece of form data includes multiple field data corresponding to the multiple field names respectively.

[0048] Complex form data often contains many merged cells, which is a key reason why the data is difficult to fully utilize. Data obtained directly from online documentation APIs often contains merged cells that span multiple cells in rows or columns. Since the field data in merged cells can only be stored as field data for one form data entry, and other form data does not contain that field data, the corresponding field information is lost. This results in inaccurate data features stored in the database, and consequently, inaccurate similarity matching.

[0049] In this embodiment, merged cells in the form are split. For example, if a merged cell spans 3 rows and 6 columns, it can be broken down into 3*6=18 cells, with each cell containing the same data. Through this progressive refinement, m*n merged cells are reduced to m*n individual cells. When storing data in the database, this ensures that the data features stored in the database are accurate, leading to more accurate similarity matching.

[0050] 104. Concatenate the multiple fields of each form data in the second form to obtain multiple feature data.

[0051] Each of these multiple feature data points corresponds one-to-one with each of these multiple form data points.

[0052] These multiple feature data can be stored in a database. The database can be a database that stores form data in text form (natural language form) or a vector database that stores form data in vector form (the vector representation of the form data obtained by extracting features from the form data). There is no limitation here.

[0053] 105. Obtain multiple candidate feature data that match the target retrieval question from these multiple feature data.

[0054] This involves performing similarity matching between multiple feature data and the target retrieval question, thereby identifying multiple candidate feature data that match the target retrieval question from the multiple feature data. Specific similarity matching algorithms may include, but are not limited to, cosine similarity algorithms, Euclidean distance algorithms, etc., or a similarity matching model may be used to identify multiple candidate feature data that match the target retrieval question from multiple feature data; no limitation is imposed here.

[0055] The similarity matching model can be a model specially trained for similarity matching of feature data in the database and the question, or can be a language reasoning model (i.e., a large language model, referred to as a large model) based on a Transformer architecture, capable of generating natural and fluent text, and understanding and processing various natural language tasks, by performing unsupervised learning on a large amount of text data to learn knowledge in aspects such as grammar, semantics, and usage of language.

[0056] In the embodiments of the present disclosure, the electronic device can obtain a plurality of feature data from the database, and then determine a plurality of candidate feature data matched with the target search question from the plurality of feature data; or the electronic device can transmit the target search question to the database, and then the database determines a plurality of candidate feature data matched with the target search question from the plurality of feature data, and then returns the plurality of candidate feature data to the electronic device; which is not limited here.

[0057] Exemplarily, the plurality of feature data, the target search question, and the similarity matching prompt word can be input into the large language model, and a plurality of candidate feature data is output.

[0058] The target search question can include at least one field data and at least one question field, and the matching degree of the plurality of field data included in the candidate feature data and the at least one field data is greater than the matching degree of the plurality of field data included in any non-candidate feature data in the plurality of feature data and the at least one field data. The non-candidate feature data is feature data other than the plurality of candidate feature data in the plurality of feature data.

[0059] 106, input the target search question and the plurality of candidate feature data into a target analysis model, and output a target answer.

[0060] The target analysis model is used to analyze the target search question and the plurality of candidate feature data to obtain the target answer matched with the target search question.

[0061] The target answer is an answer to the at least one question field.

[0062] The target analysis model can be a model specially trained for analyzing the target search question and the plurality of candidate feature data to obtain the target answer matched with the target search question, or can be a language reasoning model (i.e., a large language model, referred to as a large model) based on a Transformer architecture, capable of generating natural and fluent text, and understanding and processing various natural language tasks, by performing unsupervised learning on a large amount of text data to learn knowledge in aspects such as grammar, semantics, and usage of language.

[0063] Exemplarily, the target retrieval question, the plurality of pieces of candidate feature data, and the corresponding prompt words are input into a target analysis model, and a target answer is output.

[0064] In the embodiments of the present disclosure, by splitting the merged cells in the form, and splicing all field data included in each piece of form data after splitting to obtain a plurality of pieces of feature data, all valid information in each piece of form data can be included in the feature data, and further, the optimal candidate feature data that can accurately answer the target retrieval question can be matched from the plurality of pieces of feature data, so that the target analysis model can be used to analyze the target retrieval question and the plurality of pieces of candidate feature data, and obtain the target answer matched with the target retrieval question, and the accuracy of the target answer can be improved.

[0065] In some embodiments of the present disclosure, in combination with Figure 1 As shown in the above step 105, the data processing method provided by the embodiments of the present disclosure can further include the following step 107, and the above step 105 can be implemented by the following step 105a. Figure 2

[0066] 107. Feature extraction is performed on the target retrieval question based on the plurality of field names to obtain target feature data.

[0067] The target feature data is obtained by splicing field data matched with the plurality of field names in the target retrieval question, and the plurality of field names correspond to the plurality of field data one by one.

[0068] The field name can be header information in the form.

[0069] 105a. The plurality of pieces of candidate feature data matched with the target feature data are obtained from the plurality of pieces of feature data.

[0070] The matching degree of the candidate feature data and the target feature data is greater than the matching degree of any feature data other than the plurality of pieces of candidate feature data and the target feature data in the plurality of pieces of feature data.

[0071] In the embodiments of the present disclosure, the target feature data is obtained by extracting and splicing field data corresponding to the plurality of field names in the form in the target retrieval question, so that the key point information in the target retrieval question is retained, the interference information in the target retrieval question is eliminated, the target retrieval question is simplified, the plurality of pieces of candidate feature data matched with the target retrieval question can be accurately matched from the plurality of pieces of feature data, the matching speed is improved, and the accuracy of the answer is further improved.

[0072] In some embodiments of the present disclosure, in combination with Figure 2 As shown in the above step 105, the data processing method provided by the embodiments of the present disclosure can further include the following step 107, and the above step 105 can be implemented by the following step 105a.​Figure 3 As shown, the step 107 can be implemented by the following step 107a.

[0073] 107a, input the target search question and the first prompt word into the first large language model, and output the target feature data.

[0074] The first prompt word is used to instruct the first large language model to extract and splice the field data in the target search question that matches the plurality of field names respectively to obtain the target feature data.

[0075] The first prompt word can be designed according to actual needs, and how the first large language model extracts features from the target search question, such as extracting field data from the plurality of field names, and the data format of the extracted target feature data, is not limited herein.

[0076] In the embodiments of the present disclosure, based on the first prompt word, the first large language model can quickly and accurately extract features from the target search question, and the processing efficiency can be improved.

[0077] In some embodiments of the present disclosure, each piece of feature data is obtained by cross-splicing a field name and field data corresponding to the field name in sequence based on the plurality of field data and the plurality of field names; and the target feature data is obtained by cross-splicing a field name and field data corresponding to the field name in sequence based on the field data in the target search question that matches the plurality of field names respectively and the plurality of field names.

[0078] For example, a piece of feature data corresponds to the content of field name 1, field data 1, field name 3, field data 3, field name 4, field data 4, field name 2, field data 2, and so on. Wherein, field name 1 corresponds to field data 1, field name 2 corresponds to field data 2, field name 3 corresponds to field data 3, and field name 4 corresponds to field data 4, and so on.

[0079] In the embodiments of the present disclosure, in the process of obtaining the corresponding feature data based on the form data, the plurality of field data and the plurality of field names included in each piece of form data are spliced to obtain the corresponding feature data, which can ensure that each piece of feature data stores information containing all key information points in the form data, thereby improving the accuracy and efficiency of similarity matching.

[0080] In the embodiments of the present disclosure, the field data in the target search question that matches the plurality of field names respectively and the plurality of field names are spliced to obtain the target feature data, which can further improve the accuracy and efficiency of similarity matching.

[0081] In some embodiments of the present disclosure, the arrangement order of the plurality of field names can be determined according to actual conditions, which is not limited here.

[0082] In the embodiments of the present disclosure, the corresponding field name is spliced at the position of the corresponding field data in each feature data in the database, and the corresponding field name is spliced at the position of the corresponding field data of the target search question in the target feature data. In this way, by adding the field name, the data type information is added in the feature data, so that the similarity matching process can improve the accuracy and efficiency of the similarity matching.

[0083] In some embodiments of the present disclosure, the arrangement order of the plurality of field names in the target feature data is the same as the arrangement order of the plurality of field names in each feature data. In this way, the accuracy and efficiency of the similarity matching can be further improved.

[0084] In some embodiments of the present disclosure, the target analysis model is a second large language model; and the step 106 can be implemented by the following step 106a.

[0085] 106a, inputting the target search question, the plurality of candidate feature data and a second prompt word into the second large language model, and outputting the target answer.

[0086] The second prompt word is used to instruct the second large language model to analyze the target search question and the plurality of candidate feature data to obtain the target answer.

[0087] The second prompt word can be determined according to actual use requirements, which is not limited here.

[0088] In the embodiments of the present disclosure, the second large language model can improve the efficiency and accuracy of obtaining the target answer.

[0089] For example, 1, part of the data in the original form is shown in Table 1, and the table header is module, meaning, macro name, error code tracking, and analysis description.

[0090] Table 1

[0091]

[0092] It should be noted that in the first merged cell in Table 1, “M” indicates that the “meaning” and “macro name” corresponding to module B1 and module C1 are both M; “D4” in the second merged cell indicates that the “macro name” and “error code tracking macro” corresponding to module D1 are both D4; and “N” in the second merged cell indicates that the “analysis description” corresponding to module C1 to module E1 is N.

[0093] 2. Original question: On January 1, 2025, Xiaoming encountered a system error, he first noticed that the problem occurred in module Al, then he located to belong to M meaning, and then found the error code B4, give analysis.

[0094] 3. After splitting the merged cells in the original form, the form is as shown in Table 2.

[0095] Table 2

[0096] Module Meaning Macro name Error code tracking macro Analysis explanation A1 A2 X A3 A4 B1 M M B4 B5 C1 M M C4 N D1 D2 D4 D4 N E1 E2 E3 E4 N

[0097] 4. Combine the table header (field name) data into the cell data and store the entire row into the embedding database.

[0098] The first feature data corresponds to module Al, meaning A2, macro name X, error code tracking A3, and analysis A4.

[0099] The second feature data corresponds to module Al, meaning M, macro name M, error code tracking B4, and analysis B5.

[0100] The third feature data corresponds to module Al, meaning M, macro name M, error code tracking C4, and analysis N.

[0101] The fourth feature data corresponds to module Al, meaning D2, macro name D4, error code tracking D4, and analysis N.

[0102] The fifth feature data corresponds to module El, meaning E2, macro name E3, error code tracking E4, and analysis N.

[0103] 5. Based on the first language model, the original question is converted into the most suitable search question (target feature data).

[0104] The prompt words are as follows:

[0105]

[0106] In this way, the original question is converted into a dictionary:

[0107]

[0108]

[0109] Based on this dictionary, the original question is converted into target feature data, i.e. module Al, meaning M, macro name, error code tracking B4, and analysis.

[0110] 6. Send the new question above to the embedding database. The top k options include the second feature data (module is A1, meaning is M, macro name is M, error code tracking is B4, analysis description is B5), that is, the second feature data is candidate feature data, and other candidate feature data are also included.

[0111] 7. Using the second feature data (candidate feature data) and other candidate feature data (multiple candidate feature data) as context, and combining them with the original question and prompt words, we feed them into the second language model, outputting the answer: Analysis explanation is B5. The prompt words are used to illustrate how to find an answer that matches the user's expectations from the context that matches the original question; their specific meaning can be determined based on actual usage needs and is not limited here.

[0112] Figure 4 This is a structural block diagram of a data processing apparatus shown in an embodiment of the present disclosure, such as... Figure 3 As shown, it includes:

[0113] The receiving module 401 is used to receive a target retrieval question; the obtaining module 402 is used to obtain a first form corresponding to the target retrieval question, the first form including header information and multiple form data, the header information including multiple field names, and the first form containing merged cells; the processing module 403 is used to split the merged cells in the first form and fill the field data included in the merged cells into the split cells to obtain a second form, in which each form data includes multiple field data corresponding to multiple field names; the multiple field data included in each form data in the second form are concatenated to obtain multiple feature data, which correspond one-to-one with the multiple form data; the obtaining module 402 is also used to obtain multiple candidate feature data matching the target retrieval question from the multiple feature data; the output module 404 is used to input the target retrieval question and the multiple candidate feature data into a target analysis model and output a target answer; the target analysis model is used to analyze the target retrieval question and the multiple candidate feature data to obtain the target answer matching the target retrieval question.

[0114] In some embodiments of the present disclosure, the device further comprises a feature extraction module configured to perform feature extraction on the target search question based on the plurality of field names to obtain target feature data before obtaining a plurality of pieces of candidate feature data matching the target feature data from the plurality of pieces of feature data, wherein the target feature data is obtained by splicing field data in the target search question matching the plurality of field names respectively, and the plurality of field names correspond to the plurality of field data one by one; and the obtaining module 402 is specifically configured to obtain the plurality of pieces of candidate feature data matching the target feature data from the plurality of pieces of feature data.

[0115] In some embodiments of the present disclosure, the feature extraction module is specifically configured to input the target search question and a first prompt word into a first large language model to output the target feature data, wherein the first prompt word is used to instruct the first large language model to extract and splice field data in the target search question matching the plurality of field names respectively to obtain the target feature data.

[0116] In some embodiments of the present disclosure, each piece of feature data is obtained by cross-splicing a field name and field data corresponding to the field name in the order of the field name and the field data corresponding to the field name based on the plurality of field data and the plurality of field names; and the target feature data is obtained by cross-splicing field data in the target search question matching the plurality of field names and the plurality of field names in the order of a field name and field data corresponding to the field name.

[0117] In some embodiments of the present disclosure, the arrangement order of the plurality of field names in the target feature data is the same as the arrangement order of the plurality of field names in each piece of feature data.

[0118] In some embodiments of the present disclosure, the target analysis model is a second large language model; and the output module 404 is specifically configured to input the target search question, the plurality of pieces of candidate feature data and a second prompt word into the second large language model to output the target answer, wherein the second prompt word is used to instruct the second large language model to analyze the target search question and the plurality of pieces of candidate feature data to obtain the target answer.

[0119] In the embodiments of the present disclosure, each module can implement the data processing method provided by the above-mentioned method embodiments, and achieve the same technical effects. To avoid repetition, it will not be described here.

[0120] Figure 5 A structural schematic diagram of an electronic device is provided for the embodiments of the present disclosure to exemplarily illustrate the electronic device implementing any data processing method in the embodiments of the present disclosure, which should not be understood as a specific limitation on the embodiments of the present disclosure.

[0121] As Figure 5As shown, the electronic device 500 can include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 501 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 502 or loaded into a random access memory (RAM) 503 from a storage device 508. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processor 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0122] Generally, the following devices can be connected to the I / O interface 505: input devices 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 508 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 509. The communication devices 509 can allow the electronic device 500 to communicate wirelessly or wired with other devices to exchange data. Although the electronic device 500 is shown with various devices, it should be understood that all of the illustrated devices are not required to implement or possess. More or less devices can alternatively be implemented or possessed.

[0123] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 509, or installed from the storage devices 508, or installed from the ROM 502. When the computer program is executed by the processor 501, the functions defined in any of the data processing methods provided by embodiments of the present disclosure can be performed.

[0124] It should be noted that the computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal that propagates in a baseband or as part of a carrier wave, carrying computer-readable program code. Such a propagated data signal can take on many forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, transmit, or propagate programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted using any suitable medium, including but not limited to a wire, cable, optical fiber, RF (radio frequency), or the like, or any suitable combination thereof.

[0125] In some embodiments, the client, server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications (e.g., a communications network) of any form or medium, such as the Internet. Examples of communications networks include local area networks ("LAN"), wide area networks ("WAN"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.

[0126] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.

[0127] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: receive a target search question; obtain a first form corresponding to the target search question, the first form including form header information and a plurality of form data, the form header information including a plurality of field names, and the first form including a merged cell; perform cell splitting processing on the merged cell in the first form, and fill field data included in the merged cell into each cell obtained by splitting, to obtain a second form, in which each form data includes a plurality of field data corresponding to the plurality of field names respectively; splice the plurality of field data included in each form data in the second form to obtain a plurality of feature data, the plurality of feature data corresponding to the plurality of form data one by one; obtain a plurality of candidate feature data matching the target search question from the plurality of feature data; input the target search question and the plurality of candidate feature data into a target analysis model, and output a target answer; and the target analysis model is configured to analyze the target search question and the plurality of candidate feature data, to obtain the target answer matching the target search question.

[0128] In an embodiment of the present disclosure, the computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as "C" or similar programming languages. The program code can execute entirely on the computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0129] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or computer readable storage device from a computer readable storage medium or computer readable storage device to implement processes described herein. As used herein, the term "computer readable storage medium" or "computer readable storage device" is intended to encompass a non-transitory computer readable medium or device that is accessible by a computer, other programmable data processing apparatus, or computer readable storage device. Computer readable storage media and computer readable storage devices can include, but are not limited to, volatile memory, non-volatile memory, and / or storage devices such as a magnetic disks, optical disks, and / or tape. A data storage medium can be embodied as a set of computer readable instructions encoded on the medium. A computer readable storage medium can also be any medium that is accessible by a computer or a computer readable storage device to retrieve and execute the computer readable instructions. A computer readable storage medium can be a tangible device that can retain, store, or maintain the computer readable instructions for use by a computer or a computer readable storage device. A computer readable storage medium can also be a computer readable signal medium that stores the computer readable instructions in a transitory or non-transitory fashion.

[0130] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0131] The functions described above in the detailed description can be performed in at least part by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0132] In the context of the present disclosure, a computer readable medium can be a tangible medium that can contain or store the program for use by or in connection with an instruction execution system, apparatus, or device. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium can include one or more lines of electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0133] The above description merely illustrates the preferred embodiments of the disclosure and a principle for applying the technologies. It is understood by those skilled in the art that the disclosed scope of the disclosure is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by the combinations of the technical features described above or their equivalent features without departing from the disclosed concept. For example, the technical solutions formed by the mutual replacement of the above-described features and the technical features with similar functions disclosed in the disclosure (but not limited to) can be formed.

[0134] Further, although operations are depicted in a particular, sequential order, this should not be understood as requiring or implying that the operations are performed in the order illustrated or sequentially. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, although specific implementation details are contained in the above discussion, these should not be construed as limiting the scope of the disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0135] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A data processing method, characterized by, The method comprises: receiving a target retrieval question; obtaining a first form corresponding to the target retrieval question, the first form comprising form header information and a plurality of form data, the form header information comprising a plurality of field names, and the first form comprising a merged cell; performing cell splitting processing on the merged cell in the first form, and filling field data included in the merged cell into each cell obtained by splitting to obtain a second form, wherein each form data in the second form comprises a plurality of field data corresponding to the plurality of field names respectively; splicing the plurality of field data included in each form data in the second form to obtain a plurality of feature data, the plurality of feature data corresponding to the plurality of form data one by one; obtaining a plurality of candidate feature data matching the target retrieval question from the plurality of feature data; inputting the target retrieval question and the plurality of candidate feature data into a target analysis model to output a target answer; the target analysis model is used to analyze the target retrieval question and the plurality of candidate feature data to obtain the target answer matching the target retrieval question.

2. The method of claim 1, wherein, Before the plurality of candidate feature data matching the target retrieval question is obtained from the plurality of feature data, the method further comprises: based on the plurality of field names, extracting features of the target retrieval question to obtain target feature data, the target feature data being obtained by splicing field data in the target retrieval question matching the plurality of field names respectively; the plurality of candidate feature data matching the target feature data is obtained from the plurality of feature data. based on the plurality of field names, extracting features of the target retrieval question to obtain target feature data, comprising:

3. The method of claim 2, wherein, inputting the target retrieval question and a first prompt word into a first large language model to output the target feature data; wherein the first prompt word is used to instruct the first large language model to extract and splice field data in the target retrieval question matching the plurality of field names respectively to obtain the target feature data. each feature data is obtained by cross-splicing a field name and field data corresponding to the field name in the order of the field name and the field data corresponding to the field name based on the plurality of field data and the plurality of field names; 4. The method of claim 2, wherein, the target feature data is obtained by cross-splicing a field name and field data corresponding to the field name in the order of the field name and the field data corresponding to the field name based on field data in the target retrieval question matching the plurality of field names respectively and the plurality of field names. the arrangement order of the plurality of field names in the target feature data is the same as the arrangement order of the plurality of field names in each feature data.

5. The method of claim 4, wherein, comprises:

6. A data processing apparatus, characterized by a receiving module configured to receive a target retrieval question; ​ The acquisition module is configured to acquire a first form corresponding to the target search question, the first form including form header information and a plurality of pieces of form data, the form header information including a plurality of field names, and the first form including a merged cell; The processing module is configured to perform cell splitting processing on the merged cell in the first form, fill field data included in the merged cell into each cell obtained by splitting, and obtain a second form, in which each piece of form data includes a plurality of field data corresponding to the plurality of field names respectively; The plurality of pieces of field data included in each piece of form data in the second form are spliced to obtain a plurality of pieces of feature data, the plurality of pieces of feature data corresponding to the plurality of pieces of form data one by one; The acquisition module is further configured to acquire, from the plurality of pieces of feature data, a plurality of pieces of candidate feature data matching the target search question; The output module is configured to input the target search question and the plurality of pieces of candidate feature data into a target analysis model, and output a target answer. The target analysis model is configured to analyze the target search question and the plurality of pieces of candidate feature data, and obtain the target answer matching the target search question.

7. The apparatus of claim 6, wherein, The processing module is further configured to, before acquiring, from the plurality of pieces of feature data, a plurality of pieces of candidate feature data matching the target search question, perform feature extraction on the target search question based on the plurality of field names, and obtain target feature data, the target feature data being obtained by splicing field data in the target search question matching the plurality of field names respectively. The acquisition module is specifically configured to acquire, from the plurality of pieces of feature data, the plurality of pieces of candidate feature data matching the target feature data.

8. An electronic device, comprising: comprising: a memory and a processor, the memory being configured to store a computer program, and the processor being configured to execute the computer program to perform the data processing method in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, a computer program product comprising a computer program, the computer program being configured to perform the data processing method in any one of claims 1 to 6 when executed by a processor.

10. A computer program product, characterised in that, the computer program product comprising a computer program, the computer program being configured to perform the data processing method in any one of claims 1 to 6 when executed by a processor.