Training data processing method and device and electronic equipment
By dividing text blocks and identifying logical relationships of scientific and technological documents, generating training samples with prompt words, and identifying the domain probability of text blocks, the problem of inefficient extraction and evaluation of scientific and technological documents in the prior art is solved, and more efficient and accurate model training is achieved.
Patent Information
- Application Number
- CN202510299167.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to efficiently extract information from scientific and technological documents and conduct multi-dimensional quantitative evaluation, resulting in inefficient information integration and project evaluation.
By dividing the sample documents, identifying the logical relationship between text blocks, building text block combinations, and generating training samples with prompt words, identifying the probability set of each text block belonging to each field, and finally generating a training text set for model training.
It improves the training efficiency and effect of the model, improves the training accuracy of the model, and can make full use of the information in scientific and technological documents, and improves the efficiency of information integration and project evaluation.
Smart Images

Figure CN120146048A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method, apparatus, and electronic device for training data processing. Background Art
[0002] To enhance technological competitiveness, enterprises have been continuously increasing their R & D investment, resulting in a sharp increase in the number of scientific and technological documents such as scientific research project application forms and bidding documents, which has brought huge challenges to information integration, project evaluation, and other tasks. There is an urgent need for an efficient information extraction and multi-dimensional quantitative evaluation method for scientific and technological documents. Summary of the Invention
[0003] This application aims to at least solve one of the technical problems in the related art to some extent.
[0004] To this end, the first objective of this application is to propose a method for training data processing to achieve efficient information extraction from scientific and technological documents.
[0005] The second objective of this application is to propose a training data processing apparatus.
[0006] The third objective of this application is to propose an electronic device.
[0007] The fourth objective of this application is to propose a computer-readable storage medium.
[0008] The fifth objective of this application is to propose a computer program product.
[0009] To achieve the above objectives, an embodiment of the first aspect of this application proposes a method for training data processing, including:
[0010] Dividing each sample document in the sample set to obtain one or more text blocks of the sample document;
[0011] For any one of the sample documents, identifying the logical relationship between the text blocks and constructing one or more text block combinations based on the logical relationship;
[0012] Determining a first training sample based on the text block combination and a first prompt;
[0013] Obtaining a probability set of each text block belonging to each field, and determining a second training sample based on the probability set and a second prompt;
[0014] Obtaining a training text set according to the first training sample and the second training sample for model training.
[0015] To achieve the above objectives, an embodiment of the second aspect of this application proposes a training data processing apparatus, including:
[0016] A first acquisition module, configured to divide each sample document in a sample set to obtain one or more text blocks of the sample document;
[0017] A second acquisition module, configured to, for any one of the sample documents, identify the logical relationships between the text blocks, and construct one or more text block combinations based on the logical relationships;
[0018] A third acquisition module, configured to determine a first training sample based on the text block combination and a first prompt;
[0019] A fourth acquisition module, configured to obtain a probability set of each text block belonging to each field, and determine a second training sample based on the probability set and a second prompt;
[0020] A fifth acquisition module, configured to obtain a training text set according to the first training sample and the second training sample for model training.
[0021] To achieve the above object, an embodiment of the third aspect of the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0022] The memory stores computer-executable instructions;
[0023] The processor executes the computer-executable instructions stored in the memory to implement the method described in the embodiment of the first aspect.
[0024] To achieve the above object, an embodiment of the fourth aspect of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the embodiment of the first aspect.
[0025] To achieve the above object, an embodiment of the fifth aspect of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described in the embodiment of the first aspect.
[0026] The training data processing method, device, and electronic device provided by this application divide sample documents to obtain one or more text blocks, analyze the text blocks of each sample document, mine the logical relationships between the text blocks, construct text block combinations based on the logical relationships, determine the first training sample with the text block combinations and the first prompt, identify the probability set of each text block belonging to each field, determine the second training sample according to the probability set and the second prompt, generate a training sample set based on the first training sample and the second training sample for model training, and use more sufficient text information for model training to improve the training efficiency and training effect of the model and enhance the training accuracy of the model.
[0027] Additional aspects and advantages of this application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of this application. Brief Description of the Drawings
[0028] The above and / or additional aspects and advantages of this application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0029] Figure 1 is a schematic flowchart of a training data processing method provided by an embodiment of this application;
[0030] Figure 2 is a schematic flowchart of another training data processing method provided by an embodiment of this application;
[0031] Figure 3 is a schematic structural diagram of a training data processing device provided by an embodiment of this application. Detailed Description of the Embodiments
[0032] The embodiments of this application will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain this application and should not be construed as limiting this application.
[0033] With the rapid development of natural language processing technology, large models have shown significant advantages in processing complex text tasks and have high application value for tasks such as text duplicate checking and innovation evaluation in scientific and technological document evaluation. Since enterprise scientific and technological documents are often closely related to the enterprise's professional fields and have extremely high professional barriers, fine-tuning technology is required to improve the text understanding ability of general large models in vertical professional fields.
[0034] The fine-tuning of large models requires a large amount of high-quality datasets as support. Currently, the construction of large model fine-tuning datasets mainly relies on manual annotation or screening data from public resources. Manual annotation is costly, and the data quality and pertinence of public resources often fail to meet the requirements of specific tasks.
[0035] Existing technical solutions for constructing fine-tuning datasets generally rely on vertical domain knowledge bases as data sources. However, these methods often overlook the unique structural features of scientific and technological documents themselves. In scientific and technological documents, each paragraph undertakes a specific functional role, such as elaborating on the research background, describing technical solutions, analyzing expected results, etc. For large model fine-tuning datasets used to support tasks such as knowledge extraction and multi-dimensional evaluation, simply relying on vertical domain knowledge information is insufficient. It is also necessary to fully reflect the functional content of document paragraphs to facilitate the large model's deeper understanding and grasp of the structure and logic of scientific and technological documents.
[0036] The following describes the training data processing method, device, and electronic device according to the embodiments of the present application with reference to the accompanying drawings.
[0037] Figure 1 It is a schematic flowchart of a training data processing method provided by an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0038] S101, divide each sample document in the sample set to obtain one or more text blocks of the sample document.
[0039] In some implementations, the sample documents in the sample set can be scientific and technological documents obtained from public or private enterprise channels; the sample documents at least include content such as main titles, sub-titles, and the main text.
[0040] Optionally, the main text can be divided, for example, based on the number of paragraphs in the main text to obtain one or more text blocks; it can also be divided based on sub-titles, and all the main text content within one sub-title is divided into one text block, thereby obtaining one or more text blocks.
[0041] S102, for any sample document, identify the logical relationships between text blocks and construct one or more text block combinations based on the logical relationships.
[0042] It is understandable that scientific and technological documents are generally documents that record scientific and technological facts, precipitate scientific and technological experience, publish scientific and technological achievements, and transmit scientific and technological information in an objective and rigorous language. They can be used in the fields of product R & D and manufacturing, and include various aspects such as research methods, research materials, results, and discussions. Therefore, do scientific and technological documents include multiple components? There is a certain correlation between the descriptive texts corresponding to different text blocks. For example, the content under two subheadings is about the acquisition of two materials, which may be parallel or progressive. That is, after obtaining one material, another material is obtained based on this material.
[0043] Optionally, professionals can identify and analyze the logical relationships between different text blocks in the sample document. The logical relationships can include but are not limited to causal relationships, progressive relationships, parallel relationships, and prerequisite relationships, etc.
[0044] Furthermore, at least two text blocks corresponding to the logical relationship can be obtained and combined, and a text block combination is obtained by combining the corresponding logical relationship. That is to say, the text block combination includes at least the logical relationship and a group of text blocks corresponding to this logical relationship.
[0045] S103. Determine the first training sample based on the text block combination and the first prompt.
[0046] Optionally, the first prompt can be adaptively determined based on the logical relationship corresponding to the text block combination. The first training sample can be a question-and-answer pair generated based on the text block combination and the first prompt. For example, the first training sample is obtained by processing the text block combination and the first prompt through a pre-trained model, where the text blocks in the text block combination are respectively used to form the question and answer of the first training sample.
[0047] S104. Obtain the probability set of the text block belonging to each field, and determine the second training sample based on the probability set and the second prompt.
[0048] Optionally, the field refers to the subject field or engineering application field to which the document belongs, such as subject fields like electrical automation, electronic information, water conservancy engineering, environmental engineering, or engineering application fields like thermal power, hydropower, wind power, and photovoltaic, etc., to achieve more accurate classification of scientific and technological documents.
[0049] Optionally, a pre-trained model can be used to identify the main content of the text block to determine the probability of the text block belonging to each field, thereby obtaining the probability set; in some implementations, the number of probabilities in the probability set is the same as the number of fields. When the text block completely does not belong to a certain field, the corresponding probability is 0.
[0050] Further, a second training sample is generated based on the combination of the probability set and the second prompt. In some implementations, the text block, the probability set corresponding to the text block, and the second prompt can also be input into a pre-trained large model to generate the second training sample.
[0051] S105. Obtain a training text set according to the first training sample and the second training sample for model training.
[0052] It can be understood that the first training sample includes the logical relationships between text blocks, and the second training sample includes the probabilities of each text block belonging to each field, fully mining the relationships between text blocks in the document. Using the first training sample and the second training sample to form a training text set for model training can fully mine the information of each text block in the document, improving the efficiency and accuracy of model training.
[0053] In this embodiment, by partitioning the sample document, one or more text blocks in the sample document are obtained, the text blocks of each sample document are analyzed to mine the logical relationships between the text blocks, and text block combinations are constructed based on the logical relationships. The first training sample is determined based on the text block combination and the first prompt, the probability set of each text block belonging to each field is identified, and the second training sample is determined according to the probability set and the second prompt. A training sample set is generated based on the first training sample and the second training sample for model training, and the model is trained with more comprehensive text information, improving the training efficiency and training effect of the model and enhancing the model training accuracy.
[0054] Figure 2 It is a schematic flowchart of another training data processing method provided by an embodiment of the present application. As Figure 2 shown, the method includes the following steps:
[0055] S201. Obtain one or more titles in each sample document.
[0056] Optionally, a title matching rule can be obtained. The title matching rule includes the character rule and regular expression of the title; based on the title matching rule, identification and matching are performed in the sample document to determine the text segments that conform to the character rule or regular expression, and one or more titles are obtained.
[0057] It can be understood that titles usually have specific formats, such as starting with specific symbols, fonts, or numbers. For example, the character rule defines that the title is a line starting with a title symbol, or the line starting with a title symbol is the title through the regular expression rule.
[0058] Optionally, Python character matching and regular expressions can be used to identify text segments in the sample document that conform to the rules, obtaining one or more titles. In this embodiment, the titles are the various subheadings in the sample document.
[0059] In some implementations, before obtaining one or more titles in the sample document, the original document can also be obtained; the original document is subjected to format conversion and data cleaning to obtain the sample document; where the original document may be a PDF, image, or other non-pure text format document, and the original document is format-converted through a format conversion tool, such as Optical Character Recognition (OCR), to obtain a pure text format.
[0060] Furthermore, the pure text format document is subjected to data cleaning to remove abnormal characters, page numbers, and other irrelevant data in the pure text format, remove irrelevant format information such as extra spaces and tab characters, and remove irrelevant text content such as footnotes and chart captions, thereby obtaining the sample document after data cleaning.
[0061] S202, Divide the sample document into one or more text blocks based on the titles.
[0062] Optionally, the position of each title can be used as the start of a text segment until the line where the next title is located, obtaining a text block, that is, the text content between every two titles is used as a text block, thereby dividing the sample document into one or more text blocks.
[0063] S203, For any sample document, identify the logical relationships between the text blocks.
[0064] Optionally, the text blocks can be sent to the review department, and the review department analyzes the text content of the text blocks to determine the logical relationships between the text blocks, where the logical relationships include at least one of: causal relationship, parallel relationship, progressive relationship, contrast relationship, and conditional relationship; it can be understood that the review department can include one or more professionals to distinguish and identify the logical relationships between the text contents of the text blocks, thereby extracting the logical relationships between the text blocks.
[0065] S204, Extract at least two text blocks associated with each logical relationship.
[0066] It can be understood that at least two text blocks associated with an existing logical relationship are extracted and marked.
[0067] S205, Based on at least two text blocks and the corresponding logical relationships, form text block combinations.
[0068] Exemplary illustration, assume that there are two text blocks with any logical relationship, then the text block combination can be expressed as (B q , B a ), where B q and B a are the corresponding text blocks respectively.
[0069] S206. Based on the text block combination and the first prompt, determine the first training sample.
[0070] Optionally, the text block combination and the first prompt can be input into a pre-trained large language model, and the large language model processes the text block combination and the first prompt to generate the corresponding first training sample.
[0071] Exemplary illustration, the input content can be expressed as [(B q , B a ), Prompt1], Prompt1 is the first prompt, (B q , B a ) is the text block combination; the first training sample output by the large language model can be expressed as (Q, A), and the first training sample exists in the form of a question-answer pair, that is, the text blocks in the text block combination are used as the question part and the answer part respectively. For example, using B q as the question part in the question-answer pair and B a as the answer part in the question-answer pair, and a complete question-answer pair is generated through the processing of the large language model to obtain the first training sample.
[0072] S207. Obtain the probability set of the text block belonging to each field, and determine the second training sample based on the probability set and the second prompt.
[0073] In some implementations, the text block can be input into a pre-trained large language model to obtain the probability set of the text block belonging to each field. The fields include subject fields and engineering application fields; the subject fields can specifically include water conservancy engineering, environmental engineering, energy and power engineering, and electronic information, etc., and the engineering application fields can specifically include thermal power, hydropower, wind power, and photovoltaic, etc.
[0074] In some implementations, the second training sample can also be obtained according to the text block, the probability set, and the second prompt; exemplary illustration, the text block, the probability set, and the second prompt can be input into a pre-trained large language model, and the corresponding second training sample is output. The second training sample can also exist in the form of a question-answer pair, that is, (Q, A) = LLM(B i , P c (B i ), Prompt 2), where LLM is the large language model; Bi is a text block; P c (B i ) is the probability set corresponding to the text block; Prompt 2 is the second prompt 。
[0075] S208. Obtain a training text set from the first training sample and the second training sample for model training.
[0076] In some implementations, the training samples in the training text set can also be evaluated for quality, such as evaluations of integrity, accuracy, and diversity; by way of example, all the titles in the training text set can be checked to ensure that they have been correctly identified and segmented without omission; the integrity of the question-and-answer pairs can be evaluated to ensure that each question has a corresponding answer; by manually reviewing a certain proportion of the question-and-answer pairs, the logical relationship between the question and the answer can be evaluated for correctness and whether the answer accurately reflects the original content; for classification tasks, the matching degree between the class probabilities predicted by the model and the actual classes can be evaluated; for classification tasks, ensure that the data set covers all elements of the class set and that the sample sizes of the elements are relatively balanced; domain experts or actual users can also be invited to evaluate the training text set, collect feedback, and obtain a training text set of higher quality.
[0077] Furthermore, model training can be performed based on the training text set to obtain a pre-trained recognition model. In this embodiment, a pre-trained large language model is used as the teacher model, and the recognition model to be trained is used as the student model. The student model is trained with the training text set composed of the first training sample and the second training sample output by the teacher model to obtain a trained student model, which is the pre-trained recognition model. Document processing and text understanding are performed based on the pre-trained recognition model to improve the model processing effect.
[0078] In this embodiment, the title of the sample document is recognized through character rules or regular expressions, and the sample document is segmented based on the title to obtain one or more text blocks. The logical relationship between the text blocks is further recognized, and a text block combination is constructed based on the logical relationship. The pre-trained large language model is used to process the text block combination to generate the first training sample of the question-and-answer pair; correspondingly, the probability set of each text block is obtained, and the second training sample of the question-and-answer pair is generated based on the pre-trained large language model and the probability set, etc. The first training sample and the second training sample are used to generate a training text set for model training. Through automated and refined data preprocessing, the quality and pertinence of the training text set are improved. By constructing question-and-answer pairs with rich semantic information, the depth of understanding of the content of scientific and technological documents by the recognition model of the smaller model is enhanced, and the training effect of the smaller model is improved, making the performance of the small-parameter model in document evaluation tasks such as text duplication checking and innovation evaluation significantly improved.
[0079] To implement the above embodiments, the present application further provides a training data processing device.
[0080] Figure 3 The following is a schematic structural diagram of a training data processing device provided by an embodiment of the present application. As Figure 3 shown, the training data processing device 300 includes:
[0081] A first acquisition module 301, configured to divide each sample document in the sample set to obtain one or more text blocks of the sample document;
[0082] A second acquisition module 302, configured to, for any sample document, identify the logical relationship between text blocks, and construct one or more text block combinations based on the logical relationship;
[0083] A third acquisition module 303, configured to determine a first training sample based on the text block combination and the first prompt;
[0084] A fourth acquisition module 304, configured to obtain a probability set of each text block belonging to each field, and determine a second training sample based on the probability set and the second prompt;
[0085] A fifth acquisition module 305, configured to obtain a training text set according to the first training sample and the second training sample for model training.
[0086] Further, in a possible implementation manner of the embodiment of the present application, the second acquisition module 302 includes:
[0087] Extract at least two text blocks associated with each logical relationship;
[0088] Form a text block combination according to the at least two text blocks and the corresponding logical relationship.
[0089] Further, in a possible implementation manner of the embodiment of the present application, the fourth acquisition module 304 includes:
[0090] Input the text block into a pre-trained large language model to obtain a probability set of each text block belonging to each field, where the fields include subject fields and engineering application fields;
[0091] Obtain a second training sample according to the text block, the probability set, and the second prompt.
[0092] Further, in a possible implementation manner of the embodiment of the present application, the first acquisition module 301 includes:
[0093] Obtain one or more titles in each sample document;
[0094] Divide the sample document into one or more text blocks based on the title.
[0095] Further, in a possible implementation manner of the embodiment of the present application, the first acquisition module 301 includes:
[0096] Obtain a title matching rule, where the title matching rule includes a character rule and a regular expression of the title;
[0097] Based on the title matching rule, perform identification and matching in the sample document to determine text segments that conform to the character rule or the regular expression, and obtain one or more titles.
[0098] Further, in a possible implementation manner of the embodiment of the present application, the first acquisition module 301 further includes:
[0099] Obtain the original document;
[0100] Perform format conversion and data cleaning on the original document to obtain the sample document.
[0101] Further, in a possible implementation manner of the embodiment of the present application, the second acquisition module 302 includes:
[0102] Send the text block to the review department, and the review department analyzes the text content of the text block to determine the logical relationship between the text blocks, where the logical relationship includes at least one of: causal relationship, parallel relationship, progressive relationship, contrast relationship, and conditional relationship.
[0103] Further, in a possible implementation manner of the embodiment of the present application, the fifth acquisition module 305 further includes:
[0104] Perform model training based on the training text set to obtain a pre-trained recognition model.
[0105] It should be noted that the foregoing explanation of the embodiment of the training data processing method also applies to the training data processing device of this embodiment, and will not be repeated here.
[0106] In the embodiment of the present application, by dividing the sample document, one or more text blocks in the sample document are obtained, the text blocks of each sample document are analyzed, the logical relationship between the text blocks is mined, and a text block combination is constructed based on the logical relationship. The first training sample is determined with the text block combination and the first prompt, the probability set of each text block belonging to each field is identified, and the second training sample is determined according to the probability set and the second prompt. A training sample set is generated based on the first training sample and the second training sample for model training, so as to train the model with more sufficient text information, improve the training efficiency and training effect of the model, and improve the training accuracy of the model.
[0107] To implement the above embodiments, the present application also provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0108] To implement the above embodiments, the present application also provides a computer-readable storage medium storing computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method provided in the foregoing embodiments.
[0109] To implement the above embodiments, the present application also provides a computer program product including a computer program, and when the computer program is executed by a processor, it implements the method provided in the foregoing embodiments.
[0110] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in the present application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0111] It should be noted that personal information from users should be collected for legal and reasonable purposes and not shared or sold outside of these legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and signing an agreement / authorization including authorizing relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others with access to the personal information data comply with their privacy policies and procedures.
[0112] The present application is expected to provide an implementation for users to selectively block the use or access of personal information data. That is, the present disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of users.
[0113] In the descriptions of the foregoing embodiments, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0114] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0115] Any process or method description shown in the flowchart or described in other ways herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in an opposite order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0116] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0117] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGA), field-programmable gate arrays (FPGA), etc.
[0118] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0119] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0120] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.
Claims
1. A training data processing method, characterized in that: The method comprises: Dividing each sample document in the sample set to obtain one or more text blocks of the sample document; For any of the sample documents, identifying the logical relationship between the text blocks, and constructing one or more text block combinations based on the logical relationship; Determine a first training sample based on the text block combination and a first prompt word prompt; Obtain a probability set of the text block belonging to each field, and determine a second training sample based on the probability set and a second prompt; A training text set is obtained according to the first training sample and the second training sample to perform model training.
2. The method according to claim 1, characterized in that The step of constructing one or more text block combinations based on the logical relationship includes: Extract at least two text blocks associated with each logical relationship; The text block combination is formed according to the at least two text blocks and the corresponding logical relationship.
3. The method according to claim 1, characterized in that The step of obtaining a probability set that the text block belongs to each domain and determining a second training sample based on the probability set and the second prompt includes: Inputting the text block into a pre-trained large language model to obtain a probability set of the text block belonging to each field, wherein the fields include subject fields and engineering application fields; The second training sample is obtained according to the text block, the probability set and the second prompt.
4. The method according to any one of claims 1 to 3, characterized in that: The step of dividing each sample document in the sample set to obtain one or more text blocks of the sample document includes: Get one or more titles in each sample document; The sample document is divided into one or more text blocks based on the title.
5. The method according to claim 4, characterized in that The step of obtaining one or more titles in the sample document includes: Obtaining a title matching rule, wherein the title matching rule includes a character rule and a regular expression of the title; Based on the title matching rule, identification and matching are performed in the sample document to determine the text segment that meets the character rule or the regular expression, and obtain one or more titles.
6. The method according to claim 4, characterized in that Before obtaining one or more titles in the sample document, the method further includes: Get the original document; The original document is format converted and data cleaned to obtain the sample document.
7. The method according to claim 1, characterized in that The identifying the logical relationship between the text blocks includes: The text blocks are sent to the review department, which analyzes the text contents of the text blocks to determine the logical relationship between the text blocks, wherein the logical relationship includes at least one of: a causal relationship, a parallel relationship, a progressive relationship, a contrast relationship and a conditional relationship.
8. The method according to claim 1, characterized in that After obtaining the training text set according to the first training sample and the second training sample, the method further includes: Model training is performed based on the training text set to obtain a pre-trained recognition model.
9. A training data processing device, characterized in that: include: A first acquisition module, used for dividing each sample document in the sample set to obtain one or more text blocks of the sample document; A second acquisition module is used to identify the logical relationship between the text blocks for any of the sample documents, and construct one or more text block combinations based on the logical relationship; A third acquisition module is used to determine a first training sample based on the text block combination and a first prompt word prompt; A fourth acquisition module, used to obtain a probability set that the text block belongs to each field, and determine a second training sample based on the probability set and the second prompt; The fifth acquisition module is used to obtain a training text set according to the first training sample and the second training sample to perform model training.
10. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.