Document processing model training method and device and document processing method and device
By splitting and serially marking documents, combining large models and lightweight fine-tuning, a document processing model is generated, which solves the accuracy and completeness problems in document statement alignment, and achieves efficient and accurate document processing results.
Patent Information
- Application Number
- CN202510534312.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art lacks guidance on text integrity and accuracy in document statement alignment processing, resulting in poor accuracy of regularization results.
By splitting and sequentially marking documents, using large models for statement regular processing, and combining lightweight fine-tuning strategies, a document processing model is generated to ensure the chronological order and logical consistency of document fragments, using complete content and restricting illusions to optimize prompt words, and generating document processing results.
It improves the readability and time alignment accuracy of document processing results, reduces resource consumption, enhances the adaptability and accuracy of the model on specific tasks, and ensures the quality and efficiency of statement regular processing.
Smart Images

Figure CN120409428A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, specifically to the fields of natural language processing, large models, and deep learning technologies, and particularly to a training method and device for a document processing model, a document processing method, and a device therefor. Background Art
[0002] In related technologies, the sentence regularization processing of documents is often based on a pre-trained language model, and the model output is guided by designing a prompt, and a certain post-processing strategy is added to optimize the result of the regularization processing. However, a relatively single prompt is often used during model inference, lacking guiding information in various aspects such as text integrity and accuracy, and unable to guarantee the quality and logical integrity of the regularized text, resulting in poor accuracy of the sentence regularization processing result. Summary of the Invention
[0003] The present disclosure provides a training method for a document processing model, a document processing method, a training device for a document processing model, a document processing device, an electronic device, a storage medium, and a computer program product.
[0004] According to a first aspect of the present disclosure, there is provided a training method for a document processing model, including: obtaining a first sample document, and splitting the first sample document to obtain sample document fragments; determining the serial numbers of the sample document fragments, and in the first sample document, marking the sample document fragments based on the serial numbers to obtain a second sample document; performing sentence regularization processing on the second sample document through a first large model to obtain a document processing result corresponding to the first sample document; using the first sample document and the document processing result as training samples, and fine-tuning a second large model based on the training samples to obtain a document processing model.
[0005] According to a second aspect of the present disclosure, there is provided a document processing method, including: obtaining a first document, and splitting the first document to obtain a plurality of document fragments; determining the serial numbers of the document fragments, and in the first document, marking the document fragments based on the serial numbers to obtain a second document; performing sentence regularization processing on the second document through a document processing model to obtain a document regularization result of the first document; wherein the document processing model is a model trained by using the training method described in the first aspect.
[0006] According to a third aspect of the present disclosure, there is provided a training device for a document processing model, including: a first acquisition module, configured to acquire a first sample document and split the first sample document to obtain sample document segments; a second acquisition module, configured to determine the serial numbers of the sample document segments and mark the sample document segments in the first sample document based on the serial numbers to obtain a second sample document; a processing module, configured to perform sentence regularization processing on the second sample document through a first large model to obtain a document processing result corresponding to the first sample document; and a fine-tuning module, configured to use the first sample document and the document processing result as training samples and fine-tune a second large model based on the training samples to obtain a document processing model.
[0007] According to a fourth aspect of the present disclosure, there is provided a document processing device, including: an acquisition module, configured to acquire a first document and split the first document to obtain a plurality of document segments; a marking module, configured to determine the serial numbers of the document segments and mark the document segments in the first document based on the serial numbers to obtain a second document; and a document processing module, configured to perform sentence regularization processing on the second document through a document processing model to obtain a document regularization result of the first document; wherein the document processing model is a model trained by using the training method described in the first aspect.
[0008] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the training method of the document processing model described in the first aspect of the present disclosure or the document processing method described in the second aspect.
[0009] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the training method of the document processing model described in the first aspect of the present disclosure or the document processing method described in the second aspect.
[0010] According to a seventh aspect of the present disclosure, there is provided a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, it implements the training method of the document processing model described in the first aspect of the present disclosure or the document processing method described in the second aspect.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0012] The accompanying drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them:
[0013] Figure 1 It is a schematic flowchart of a method for training a document processing model according to an embodiment of the present disclosure;
[0014] Figure 2 It is a schematic flowchart of a method for training a document processing model according to an embodiment of the present disclosure;
[0015] Figure 3 It is a schematic flowchart of a method for obtaining the document processing result of a first sample document according to an embodiment of the present disclosure;
[0016] Figure 4 It is a schematic flowchart of a document processing method according to an embodiment of the present disclosure;
[0017] Figure 5 It is a schematic structural diagram of a device for training a document processing model according to an embodiment of the present disclosure;
[0018] Figure 6 It is a schematic structural diagram of a document processing device according to an embodiment of the present disclosure;
[0019] Figure 7 It is a schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0020] The following describes exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0021] The following briefly describes the technical fields related to the solution of the present disclosure:
[0022] Artificial Intelligence (AI) is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include several aspects such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0023] Natural Language Processing (NLP) is an important research direction in the field of artificial intelligence. It integrates knowledge from multiple disciplinary fields such as linguistics, computer science, machine learning, mathematics, and cognitive psychology. It is an interdisciplinary subject that combines computer science, artificial intelligence, and linguistics. It includes two main aspects: natural language understanding and natural language generation. The research content includes various levels such as characters, words, phrases, sentences, paragraphs, and chapters. It is a bridge for communication between machine language and human language. Its aim is to enable machines to understand, interpret, and generate human language, achieve effective communication between humans and machines, and enable computers to perform tasks such as language translation, sentiment analysis, and text summarization.
[0024] A large model refers to a machine learning model with a huge parameter scale and complexity. It requires a large amount of computing resources and storage space for training and storage, and often requires distributed computing and special hardware acceleration technologies. Large models have stronger generalization ability and expressive ability.
[0025] Deep Learning (DL) is a new research direction in the field of Machine Learning (ML). It is introduced into machine learning to make it closer to the original goal - artificial intelligence. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analysis and learning like humans, and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed previous related technologies.
[0026] The following describes a training method and a document processing method of a document processing model according to an embodiment of the present disclosure with reference to the accompanying drawings.
[0027] Figure 1 It is a schematic flowchart of a training method of a document processing model according to an embodiment of the present disclosure. It should be noted that the execution subject of the training method of the document processing model in this embodiment is a training device of the document processing model. The training device of the document processing model can specifically be a hardware device or software in a hardware device. Among them, the hardware device is, for example, a terminal device, a server, etc.
[0028] As Figure 1 shown, the training method of the document processing model proposed in this embodiment includes the following steps:
[0029] S101. Obtain a first sample document, and split the first sample document to obtain sample document segments.
[0030] It should be noted that the present disclosure does not limit the specific manner of obtaining the first sample document, and it can be selected according to the actual situation.
[0031] Optionally, a sample audio file can be obtained, and through Automatic Speech Recognition (ASR) technology, the sample audio file can be converted into text form to obtain the first sample document.
[0032] It should be noted that the present disclosure does not limit the type of the sample audio file.
[0033] For example: it can be an audio file in professional scenarios such as finance, law, and finance; for another example: it can be an audio file in a conversation scenario or a question-and-answer scenario; for still another example: it can be an audio file in a meeting record scenario or a news interview scenario.
[0034] Optionally, a large model can be used to understand the first sample document, and according to the understanding result, the first sample document can be split to obtain sample document fragments.
[0035] It should be noted that if the first sample document is transcribed from a sample audio file, by splitting the first sample document, each sample document fragment can more accurately correspond to the time information of the sample audio file.
[0036] S102. Determine the serial numbers of the sample document fragments, and in the first sample document, mark the sample document fragments based on the serial numbers to obtain a second sample document.
[0037] It should be noted that in the ASR scenario, the first sample document is essentially a transcription of the sample audio file. However, due to various factors such as the randomness, pauses, and repetitions of spoken language, traditional methods often have difficulty maintaining the original time order of the text in the first sample document. By splitting the first sample document and marking the sample document fragments based on serial numbers, it can be ensured that the output order of the processing results of the document processing model is consistent with the transcription of the sample audio file, avoiding information confusion, improving readability and time alignment accuracy, and avoiding various problems. For example: (1) Incorrect merging: Two sentences that are originally separated in time may be incorrectly merged due to logical relevance during sentence regularization processing, resulting in information confusion; (2) Incorrect splitting: ASR may recognize a whole sentence as multiple fragments, and split the text at places where it should not be split during sentence regularization processing, affecting readability; (3) Dislocation: If there is no constraint on the time order, the sentence order may be adjusted, resulting in the reversal of the chronological order of event descriptions and affecting the accuracy of information expression.
[0038] It should be noted that in the related art, the timestamp is often directly input into the sample document segment, enabling the large model to refer to the timestamp for sentence regularization during inference. However, the existing large models have limited ability to parse timestamps, mainly based on text reasoning rather than time series data, and it is difficult to directly use timestamps as constraint conditions. Directly inputting timestamps may increase the burden on the model. The timestamps output by ASR are usually based on the entire audio segment rather than corresponding precisely line by line. Multiple short sentences may appear within similar timestamps, making it difficult for the model to accurately distinguish the order of the text. By marking the sample document segments with serial numbers, it can be more intuitive and more in line with the understanding method of the large model. The large model is more sensitive to the displayed sequence annotations (such as: 1, 2, 3... etc.).
[0039] In the embodiment of the present disclosure, the time information of the sample document segment can be obtained, and according to the time information of the sample document segment, the sample document segments are sorted from earliest to latest. Based on the sorting result, the serial numbers of the sample document segments are determined, and the sample document segments are marked based on the serial numbers to obtain the second sample document.
[0040] S103. Perform sentence regularization processing on the second sample document through the first large model to obtain the document processing result corresponding to the first sample document.
[0041] In the embodiment of the present disclosure, the system prompt of the first large model can be determined, and the preset set of constraint information for sentence regularization is determined. The set of constraint information includes at least constraint information related to serial numbers, constraint information for content integrity, and constraint information for limiting hallucinations. Based on the system prompt and the set of constraint information, a prompt word for the first large model is generated. The first large model performs sentence regularization processing on the second document according to the prompt word to obtain the document processing result corresponding to the first sample document.
[0042] Optionally, based on the system prompt and the set of constraint information, a prompt word "prompt" for the first large model is generated. The first large model performs sentence regularization processing on the second document according to the prompt word "prompt" to obtain the document processing result "response" corresponding to the first sample document, where the prompt word "prompt" and the document processing result "response" are in JSON format.
[0043] S104. Use the first sample document and the document processing result as training samples to fine-tune the second large model based on the training samples to obtain a document processing model.
[0044] In the embodiments of the present disclosure, after obtaining the document processing result, using the first sample document and the document processing result as training samples, adopting a lightweight supervised fine-tuning (SFT) strategy, and fine-tuning the second large model based on the training samples, so that the fine-tuned document processing model has stronger instruction-following ability and adaptability in the sentence regularization processing task.
[0045] It should be noted that compared with the un-fine-tuned general large model, the SFT strategy improves the performance of the document processing model, that is, the document processing model performs better in the sentence regularization processing task, can process complex texts more accurately, reduce redundant and error information, enhances the instruction-following ability of the document processing model, that is, follows the set regularization rules more strictly, ensures that the output meets the expectations, reduces situations such as irrelevant rewriting and information loss, and through lightweight optimization, reduces resource consumption. That is, compared with training the large model comprehensively, the SFT strategy only needs to perform targeted adjustments on specific tasks, reduces the computational cost, improves the deployability of the model in the production environment, and based on the SFT strategy, the document processing model can be adapted to documents in different scenarios, that is, can quickly adapt to the sentence regularization requirements in different fields (such as finance, law, academic, etc.), and improves the wide applicability of the document processing model.
[0046] It should be noted that after obtaining the training samples, the training samples can be cleaned and merged. For example, meaningless lines (annotation lines), marked bracket content "(Note: xxx)", full-text English regularization and other invalid data in the training samples are deleted, and the data is merged according to specific conditions to improve the long-text processing ability of the document processing model, obtain the final training samples, and significantly improve the quality of the training samples and the stability of subsequent inference output.
[0047] In the embodiments of the present disclosure, based on the training samples and through the SFT strategy, the second large model is fine-tuned. By observing the curve change trend of the loss function during the fine-tuning process of the second large model, and verifying the line alignment degree, information integrity, and hallucination content control of the output of the second large model, it is judged whether the second large model meets the fine-tuning end condition. In response to the large model meeting the fine-tuning end condition, a document processing model is obtained, where the document processing model is used to perform sentence regularization processing on the document.
[0048] A training method for a document processing model according to an embodiment of the present disclosure. By obtaining a first sample document, splitting the first sample document to obtain sample document fragments, determining the serial numbers of the sample document fragments, and marking the sample document fragments based on the serial numbers in the first sample document to obtain a second sample document, performing sentence regularization processing on the second sample document through a first large model to obtain a document processing result corresponding to the first sample document, using the first sample document and the document processing result as training samples, and fine-tuning a second large model based on the training samples to obtain a document processing model. Thus, the present disclosure marks the sample document fragments based on serial numbers, avoiding the problem of information disorder in the document processing result, improving the readability and time alignment accuracy of the document processing result, and fine-tuning the second large model based on the training samples, reducing the resources for obtaining the document processing model, improving the accuracy and efficiency in the fine-tuning process of the second large model, enhancing the performance of the document processing model in sentence regularization, and ensuring the accuracy of the sentence regularization processing result.
[0049] Figure 2 It is a flowchart showing the training method of the document processing model according to an embodiment of the present disclosure.
[0050] As Figure 2 shown, the training method of the document processing model proposed in this embodiment includes the following steps:
[0051] S201. Obtain a first sample document and split the first sample document to obtain sample document fragments.
[0052] For the relevant content of step S201, reference can be made to the above embodiment, and details will not be elaborated here.
[0053] S102 "Determine the serial numbers of the sample document fragments" in the above embodiment specifically includes S202 and S204.
[0054] S202. Obtain the time information of the sample document fragments.
[0055] S203. Sort the sample document fragments from earliest to latest according to the time information of the sample document fragments.
[0056] In the embodiment of the present disclosure, after obtaining the time information of the sample document fragments, the sample document fragments can be partially ordered from earliest to latest to obtain the partial order result of the sample document fragments.
[0057] For example, if the time information of the sample document segment a "Today's meeting mainly discussed the market strategy" is 00:01:23, the time information of the sample document segment b "We focused on analyzing the user growth problem" is 00:02:05, and the time information of the sample document segment c "In addition, the team also proposed a new product optimization plan" is 00:02:45, by sorting the sample documents from the earliest to the latest, the sorting result is sample document segment a, sample document segment b, and sample document segment c.
[0058] S204. Determine the serial numbers of the sample document segments according to the sorting result.
[0059] In the embodiment of the present disclosure, after obtaining the sorting result, unique serial numbers can be assigned to each sample document segment according to the sorting result.
[0060] For example, for the sample document segment a "Today's meeting mainly discussed the market strategy", the serial number of the sample document segment a is 1, for the sample document segment b "We focused on analyzing the user growth problem", the serial number of the sample document segment b is 2, and for the sample document segment c "In addition, the team also proposed a new product optimization plan", the serial number of the sample document segment c is 3.
[0061] In the embodiment of the present disclosure, after determining the serial numbers of the sample document segments according to the sorting result, a mapping relationship between the time information of the sample document segments and the serial numbers of the sample document segments is constructed.
[0062] For example, for the sample document segment a, the time information is 00:01:23 and the serial number is 1; for the sample document segment b, the time information is 00:02:05 and the serial number is 2; for the sample document segment c, the time information is 00:02:45 and the serial number is 3.
[0063] S205. Mark the sample document segments based on the serial numbers to obtain a second sample document.
[0064] For example, if the serial number of the sample document segment a is 1, mark it with the sample document segment a "Today's meeting mainly discussed the market strategy" to obtain "1 Today's meeting mainly discussed the market strategy". Repeat the above steps until each sample document segment is marked based on the serial numbers to obtain a second sample document.
[0065] S103 in the above embodiment, "Perform sentence regularization processing on the second sample document through the first large model to obtain the document processing result corresponding to the first sample document" specifically includes S206 and S209.
[0066] S206. Determine the system prompt of the first large model.
[0067] Among them, the system prompts of the first model can be understood as strong constraint information of the first model, so as to enhance the effect of the first model on sentence regularization.
[0068] S207: Determine a preset constraint information set for regularizing sentences, where the constraint information set includes at least constraint information related to sequence numbers, constraint information for complete content, and constraint information for limiting illusions.
[0069] It should be noted that in the design process of the prompt words of the first model, in order to ensure that the first model is output in serial number order and that the number of lines and order are not modified, constraint information related to the serial number can be obtained. This is for the "redundant supplement" problem that may exist in the sentence regularization processing of the first model, for example: arbitrarily adding content that does not exist in the first document during the sentence regularization processing, including background information, subjective speculation, extended explanation, etc., resulting in lengthy and inaccurate text, and constraint information that limits illusions can be obtained. This is for the "information loss" problem that may exist in the sentence regularization processing of the first model, and the omission of key information in the first document during the sentence regularization processing, such as key information such as numbers, proper nouns and time, and complete constraint information can be obtained.
[0070] S208: Generate prompt words of the first large model based on the system prompt and constraint information set.
[0071] In an embodiment of the present disclosure, after generating the prompt words of the first large model based on the system prompts and constraint information set, at least one of the guiding examples of the set scenario of document processing and the Chain of Thought (COT) reasoning logic of document processing can be determined. Based on at least one of the guiding examples and the COT reasoning logic, the prompt words of the first large model are optimized to obtain the final prompt words of the first large model.
[0072] It should be noted that a guiding example of a set scenario for document processing, i.e., a few-shot sample, can be determined and provided to the first largest model to optimize the prompt words of the first largest model and obtain the final prompt words of the first largest model.
[0073] For example, in the scenario where the first model arbitrarily completes the sentence and adds speculative content, if the text in the first document is "Meeting tomorrow morning at nine o'clock, remember not to be late", the corresponding document processing result is "A routine department meeting will be held at nine o'clock tomorrow morning. Please arrive at the meeting room on time so that the discussion can proceed smoothly." Among them, "routine department meeting" is inferred by the first model. The original text does not specify the type of meeting. "So that the discussion can proceed smoothly" is a supplement to the first model, adding content not mentioned in the original text.
[0074] For example, in the set scenario where the first large model misses key information, such as key information like numbers and time, if the text in the first document is "The registration fee for this event is 199 yuan, and the deadline is March 15th", the corresponding document processing result is "The registration for this event has started. Please participate as soon as possible". Among them, the key information "199 yuan" and "March 15th" are lost, resulting in incomplete information.
[0075] For example, in the set scenario where the first large model does not retain complete proper nouns, if the text in the first document is "The cloud storage service we use is Amazon S3, and its performance is very good", the corresponding document processing result is "We use a cloud storage service with very good performance". Among them, "Amazon S3" as a proper noun is omitted, resulting in incomplete information.
[0076] It should be noted that the chain of thought COT reasoning logic for document processing can be determined. According to the COT reasoning logic, the prompt words of the first large model are optimized to obtain the final prompt words of the first large model. By introducing the COT reasoning logic, the first large model can be significantly improved in aspects such as paraphrasing, error correction, information retention, and line alignment, and can achieve more accurate sentence regularization processing under multi-constrained information, improving the fluency and readability of the document processing results.
[0077] For example, add instructions to the prompt words of the first large model, requiring the first large model to follow the reasoning process of full-text understanding, line-by-line processing, and line-by-line inspection, and emphasize in the prompt words of the first large model "Do not add information that does not exist in the original text" to avoid generating hallucinated content, and emphasize that "the number of lines in the regularized document is strictly the same as the number of lines in the original document" to avoid merging or splitting multiple lines.
[0078] S209. Use the first large model to perform sentence regularization processing on the second document according to the prompt words to obtain the document processing result corresponding to the first sample document.
[0079] In the embodiments of the present disclosure, under the guidance of the prompt words, the first large model uses the COT reasoning logic to perform sentence regularization processing on the second sample document to obtain the document processing result corresponding to the first sample document.
[0080] Optionally, as Figure 3 shown, "Under the guidance of the prompt words, the first large model uses the COT reasoning logic to perform sentence regularization processing on the second sample document to obtain the document processing result corresponding to the first sample document" in the above embodiments may specifically include S301 and S303.
[0081] S301. Perform full-text understanding on the second sample document.
[0082] It should be noted that by comprehensively understanding the second sample document, the original text content of the second sample document is understood to ensure a clear overall context, understand the original logic and information points, and avoid information fragmentation caused by subsequent line-by-line processing of the sample document.
[0083] S302. Perform line-by-line processing on the second sample document based on the prompt words to obtain a regular document fragment corresponding to the sample document fragment in the second sample document.
[0084] In the embodiments of the present disclosure, based on the prompt words, identify the core information and redundant parts of the sample document fragment in the second sample document, retain the core information, and delete the redundant parts to obtain a candidate document fragment. Correct errors in the candidate document fragment while retaining the original language style to obtain a regular document fragment.
[0085] In the embodiments of the present disclosure, retaining the core information includes: identifying the expression mode of the core information, and in response to the expression mode being the target expression mode, optimizing the expression mode of the core information, where the target expression mode is colloquial expression.
[0086] For example, if any sample document fragment is "Third, increase the promotion efforts, adopt various methods, such as promoting handicraft products, for example, we can use the Internet, build the first experience hall, or enter the service area, and then, okay? Okay.", it is identified that the core information of the sample document fragment is the promotion method, and it is identified that "such as promoting" is a colloquial expression, and the expression mode is optimized, that is, optimized to a formal expression. "And then, okay? Okay." is a redundant colloquial repetition without actual information, so by retaining the core information and deleting the redundant parts, the candidate document fragment is obtained as "Third, increase the promotion efforts, adopt various methods to promote handicraft products, such as promoting through the Internet, building the first experience hall, and entering the service area, etc.".
[0087] It should be noted that after obtaining the candidate document fragment, correct errors in the candidate document fragment while retaining the original language style, that is, correct obvious errors, such as spelling mistakes, grammar mistakes, and logical mistakes, etc., to obtain a regular document fragment.
[0088] For example, the candidate document fragment is "Then for the second point, we must support well or secondarily ensure well, and we can introduce perfect policies to assist, the development of handicraft products." It is identified that the expression "support well or secondarily ensure well" is unclear and may be a typo of "ensure well", which is corrected to "perfect policies ensure". It is identified that "handicraft products" should be "handicraft products", and by reorganizing the sentence structure to make it clearer and more fluent, the regular document fragment is obtained as "Second, perfect policies ensure, introduce preferential policies such as financial support and tax reduction and exemption to assist the development of handicraft products."
[0089] It should be noted that by preserving the original language style, that is, keeping it consistent with the original language style and not making subjective rewrites, the accuracy of the regularized document fragments is ensured.
[0090] For example, the candidate document fragment is "Right? Then, teacher, I have even higher ideas. Is that okay? Yes, no problem at all. You can elaborate on all your higher ideas. Absolutely no problem, right?" It is recognized that "higher ideas" should be adjusted to "more in-depth ideas" to avoid ambiguity, maintaining the original interactive tone, but removing the repeated "right?" to enhance readability and make the overall expression clearer and more formal, yet still in line with the original language style, resulting in the regularized document fragment "Teacher, can I add more in-depth ideas? Of course, all innovative ideas can be fully elaborated."
[0091] In the embodiment of the present disclosure, in response to the sample document fragment being content not understood by the large model, the sample document fragment is output as the corresponding regularized document fragment.
[0092] For example, if the sample document fragment is "Just two places." which is content not understood by the large model, since this sample document fragment is too short to infer the specific meaning, no modification or supplementation is made, and "Just two places." is directly output as the corresponding regularized document fragment.
[0093] S303. Output the regularized document fragment corresponding to each sample document fragment in sequence according to the serial number of the sample document fragment to obtain the document processing result corresponding to the first sample document.
[0094] In the embodiment of the present disclosure, determine the input-output format of the second sample document, and output the regularized document fragment corresponding to each sample document fragment in sequence under the constraint of the input-output format according to the serial number of the sample document fragment to obtain the document processing result corresponding to the first sample document.
[0095] It should be noted that the input-output format of the second sample document should be kept consistent. Under the constraint of the input-output format, output the regularized document fragment corresponding to each sample document fragment in sequence, that is, keep the serial numbers consistent, do not merge or split sentences, and ensure the format is standardized to obtain the document processing result corresponding to the first sample document.
[0096] For example, if the second sample document includes two sample document fragments:
[0097] 1 Okay, you can say so. Third, increase the promotion efforts and adopt various methods to promote handicraft products. For example, you can utilize the Internet, build the first experience hall, or enter service areas. Third, promote and publicize in various ways to increase the popularity of handicraft products. For example, you can open online stores, shoot short videos, and conduct online sales.
[0098] 2 Right? Teacher, I have even better ideas. Is that okay? Yes, no problem at all. You can elaborate on all your better ideas. There is definitely no problem, right? I think there are better ways to do it. I think there are more professional ways to do it. Is that okay? Yes, no problem.
[0099] Under the guidance of the prompt words by the first large model, using the COT reasoning logic, the sentence regularization process is carried out on the above-mentioned second sample document, and the document processing result corresponding to the first sample document is obtained as follows:
[0100] 1 Okay, you can say so. Third, increase the promotion efforts and adopt various methods. For example, you can use online promotion, build the first experience hall, enter service areas, etc. The specific description of the third point is: increase the popularity of handicraft products through various methods such as opening online stores, shooting short videos, online sales, entering service areas, and building the first offline experience hall.
[0101] 2 Teacher, can I add even more in-depth ideas? Of course you can. All innovative ideas can be fully elaborated. I think there are more professional implementation methods. Is it appropriate to add like this? There is absolutely no problem at all. Any reasonable suggestions can be put forward.
[0102] In the embodiment of the present disclosure, determine the serial number corresponding to the regularized document segment, query the mapping relationship between the pre-constructed time information and the serial number, obtain the time information associated with the serial number corresponding to the regularized document segment, and replace the serial number corresponding to the regularized document segment with the associated time information to obtain the document processing result.
[0103] It should be noted that after sequentially outputting each regularized document segment corresponding to the sample document segment according to the serial number of the sample document segment to obtain the document processing result corresponding to the first sample document, the document processing result can be checked for abnormalities according to the second sample document. In response to the first regularized document segment with abnormalities in the document processing result, determine the first sample document segment corresponding to the first regularized document segment from the second sample document, and perform abnormality correction on the first regularized document segment based on the first sample document segment.
[0104] Optionally, abnormal checks such as "missing key information" and "uncorrected errors in the document processing result" can be performed on the document processing result according to the second sample document. In response to the first regularized document fragment with abnormalities in the document processing result, the first sample document fragment corresponding to the first regularized document fragment is determined from the second sample document, and the first regularized document fragment is corrected for abnormalities based on the first sample document fragment, so as to ensure that the final document processing result has no omissions and no incorrect modifications and meets all requirements.
[0105] It should be noted that by using the COT inference logic under the guidance of the prompt words by the first large model to perform sentence regularization processing on the second sample document, it can be understood that: first, the second sample document is comprehensively understood to understand the content of the second sample document as a whole, ensuring clear context and avoiding misunderstandings caused by isolated processing of single sentences. Then, the second sample document is processed line by line based on the prompt words to ensure complete information, retain all key information (i.e., core information) in each line, and delete redundant parts to obtain candidate document fragments. On this basis, the candidate document fragments are corrected for errors while retaining the original language style. For grammar errors, spelling errors, and logical errors, etc., direct corrections are made without intuitive rewriting and without changing the expression method, only optimizing the smoothness and maintaining the original language style. If the sample document fragment is content that the large model does not understand, the ununderstood content is left unchanged to avoid subjective speculation, and the sample document fragment is used as the corresponding regularized document fragment for output. Determine the input-output format of the second sample document, and the input-output format is strictly matched, keeping the line numbers and the number of lines consistent, without merging or splitting sentences. According to the serial number of the sample document fragment, under the constraint of the input-output format, the corresponding regularized document fragments of each sample document fragment are sequentially output to obtain the document processing result corresponding to the first sample document. And according to the second sample document, an abnormal check is performed on the document processing result to ensure no omissions and no incorrect modifications and meet all modifications. In response to the first regularized document fragment with abnormalities in the document processing result, the first sample document fragment corresponding to the first regularized document fragment is determined from the second sample document, and the first regularized document fragment is corrected for abnormalities based on the first sample document fragment. By using the COT inference logic, it ensures line-by-line repetition, which can not only optimize the readability of the document processing result but also completely retain the original information without introducing any errors or biases.
[0106] S2010: Use the first sample document and the document processing result as training samples to fine-tune the second large model to obtain a document processing model.
[0107] A training method for a document processing model according to an embodiment of the present disclosure includes obtaining a first sample document, splitting the first sample document to obtain sample document fragments, obtaining the time information of the sample document fragments, sorting the sample document fragments from earliest to latest according to the time information of the sample document fragments, determining the serial numbers of the sample document fragments according to the sorting result, marking the sample document fragments based on the serial numbers to obtain a second sample document, determining the system prompt of the first large model, determining a preset set of constraint information for sentence regularization, the set of constraint information at least including constraint information related to serial numbers, constraint information for content integrity, and constraint information for limiting hallucinations, generating a prompt word for the first large model based on the system prompt and the set of constraint information, performing sentence regularization processing on the second document by the first large model according to the prompt word to obtain a document processing result corresponding to the first sample document, using the first sample document and the document processing result as training samples, and fine-tuning a second large model based on the training samples to obtain a document processing model. Thus, the present disclosure generates a prompt word for the first large model based on the system prompt and the set of constraint information, and adopts the COT reasoning logic to ensure the text quality and logical integrity of the sentence regularization processing result. Moreover, it fine-tunes the second large model based on the training samples to obtain a document processing model, reducing the resources for obtaining the document processing model, improving the accuracy and efficiency in the fine-tuning process of the second large model, being able to perform customized optimization on the second large model in combination with industry requirements, having strong scalability and adaptability, enhancing the performance of the document processing model for sentence regularization, and outputting the regularized document fragments corresponding to each sample document fragment in sequence according to the serial numbers of the sample document fragments, ensuring the accuracy of the sentence regularization processing result.
[0108] Figure 4 FIG. is a schematic flowchart of a document processing method according to an embodiment of the present disclosure. It should be noted that the execution subject of the document processing method in this embodiment is a document processing device, and the document processing device may specifically be a hardware device or software in a hardware device. Among them, the hardware device may be, for example, a terminal device, a server, etc.
[0109] As Figure 4 shown, the document processing method proposed in this embodiment includes the following steps:
[0110] S401. Obtain a first document and split the first document to obtain a plurality of document fragments.
[0111] Among them, the first document is any document that needs to be subjected to sentence regularization processing.
[0112] It should be noted that the present disclosure does not limit the specific manner of obtaining the first document, which can be selected according to actual situations.
[0113] Optionally, an audio file can be obtained and converted into text form through ASR technology to obtain a first document.
[0114] It should be noted that the present disclosure does not limit the type of the audio file. For example, it can be an audio file in professional scenarios such as finance, law, and finance; for another example, it can be an audio file in a conversation scenario or a question-and-answer scenario; for still another example, it can be an audio file in a meeting record scenario or a news interview scenario.
[0115] In the embodiment of the present disclosure, after the first document is obtained, the first document can be split to obtain a plurality of document fragments.
[0116] S402. Determine the serial numbers of the document fragments, and in the first document, mark the document fragments based on the serial numbers to obtain a second document.
[0117] In the embodiment of the present disclosure, the time information of the document fragments can be obtained, and according to the time information of the document fragments, the document fragments are sorted from earliest to latest, and the serial numbers of the document fragments are determined according to the sorting result.
[0118] In the embodiment of the present disclosure, after the serial numbers of the document fragments are obtained, the document fragments are marked based on the serial numbers to obtain a second document.
[0119] Optionally, serial number marks can be added before each document fragment. For example: 1\tDocument fragment m, 2\tDocument fragment n, etc.
[0120] S403. Perform sentence regularization processing on the second document through a document processing model to obtain a document regularization result of the first document.
[0121] It should be noted that the goal of sentence regularization is to improve the normativity, logic, and readability of the text. Sentence regularization is mainly carried out for the following problems: (1) Colloquial expressions in the first document. If the text in the first document is derived from daily conversations or informal written expressions, it often contains a lot of colloquial elements. For example, expressions like "and then", "that is to say", "this" affect the professionalism and fluency of the text in the first document. In addition, some sentence patterns may be too casual, lacking the refinement and accuracy of formal written language. Therefore, during the sentence regularization process, these colloquial expressions need to be removed or replaced to make the text more formal and standardized; (2) In the first document, there are problems of loose logic and poor structure. If the text in the first document has problems such as unclear logic and chaotic information organization in its expression, for example, narrative redundancy, that is, the same or similar information is repeatedly expressed in a sentence or paragraph, making the content lengthy and affecting the reading experience, unclear logical chain, that is, there is a lack of reasonable logical connection between sentences, resulting in jumping content and making it difficult to understand the cause and effect of information, loose structure, that is, the text lacks clear hierarchical division and information is piled together, making the overall coherence poor. During the sentence regularization process, it is necessary to straighten out the information logic, remove redundant content, and ensure that the content is expressed concisely, smoothly, and logically clearly; (3) In the first document, there are slips of the tongue and typos. If the text in the first document is derived from oral dictation, automatic transcription, or rapid input, it is inevitable that there will be problems such as slips of the tongue, typos, and grammar errors. For example, a slip of the tongue may lead to ambiguous expressions, and a typo may affect the accuracy of information. These are corrected during the sentence regularization process to ensure that the text is accurate; (4) In the first document, there is a problem of inaccurate time positioning. In related technologies, the entire paragraph is often directly regularized, and the model may adjust the sentence order during the processing, making the regularized text unable to be strictly arranged in the time order of the original audio. During the process of converting speech to text, the original ASR output may contain multiple speech segments. If only processed at the paragraph level, it is difficult for the model to distinguish the time sequence of each sentence, which may lead to information misalignment. During the sentence regularization process, sentences that are separated in time may be mistakenly merged, resulting in the mixing of information at different time points. The model may also mistakenly split complete sentences, making the information expression incomplete, affecting readability and semantic coherence. In addition, due to the lack of constraints on fine-grained time information, the sentence order may be adjusted during the regularization process, resulting in the reversal of the event description order and affecting semantic understanding.
[0122] In the embodiments of the present disclosure, through the document processing model under the guidance of prompt words, using the COT reasoning logic, the sentence regularization of the second document is carried out to obtain the document processing result corresponding to the first document.
[0123] Among them, the document processing model is the model obtained by using the training method of the first aspect.
[0124] For example, for the first document "Good evening, all capital investors. First of all, in our last monthly meeting, we reported our views to you. That is, the industry focus is still on consumer growth stocks", through the document processing model under the guidance of the prompt words, using the COT reasoning logic, the second document is processed for sentence regularization, and the document processing result corresponding to the first document is obtained: "Good evening, all investors. First of all, in our last month's meeting, we reported our views to you, that is, the industry focus is on consumer growth stocks". In the related technology, through the sentence regularization processing of the first document, the document processing result corresponding to the first document is obtained: "Good evening, all fund managers. First of all, let's review our suggestions on the industry in the last service, that is, focus on investing in consumer growth stocks". Analyzing the document processing result in the related technology, it can be seen that there is an error, that is, "fund manager" is incorrect.
[0125] According to the document processing method of the present disclosure embodiment, by obtaining the first document, splitting the first document to obtain multiple document fragments, determining the serial numbers of the document fragments, and marking the document fragments based on the serial numbers in the first document to obtain the second document, and then using the document processing model to perform sentence regularization processing on the second document to obtain the document regularization result of the first document. Thus, the present disclosure marks the document fragments based on the serial numbers and, under the guidance of the prompt words, uses the COT reasoning logic through the document processing model to perform sentence regularization processing on the second document, which can make the document processing result more in line with human thinking habits, more natural and professional, and can make the document processing model better understand the text structure, reduce the situations of information omission, logical confusion and hallucination content, optimize the sentence connection, ensure the clear logic of the document processing result, reduce the redundant expressions in the document processing result, improve the readability of the document processing result, ensure the quality of the document processing result, and while reducing the cost of sentence regularization processing, improve the efficiency in the sentence regularization processing process. It can be applied to the regularization of long documents and multi-speech content documents, and has outstanding effects especially in data-intensive scenarios such as finance, law, and academia.
[0126] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0127] Corresponding to the training methods of the document processing model provided by the above several embodiments, an embodiment of the present disclosure further provides a training device for the document processing model. Since the training device for the document processing model provided by the embodiment of the present disclosure corresponds to the training methods of the document processing model provided by the above several embodiments, the implementation manners of the training methods of the document processing model are also applicable to the training device for the document processing model provided by this embodiment, and will not be described in detail in this embodiment.
[0128] Figure 5 It is a schematic structural diagram of a training device for a document processing model according to an embodiment of the present disclosure.
[0129] As Figure 5 shown, the training device 500 for the document processing model includes: a first acquisition module 510, a second acquisition module 520, a processing module 530, and a fine-tuning module 540.
[0130] The first acquisition module 510 is configured to acquire a first sample document and split the first sample document to obtain sample document fragments.
[0131] The second acquisition module 520 is configured to determine the serial numbers of the sample document fragments, and mark the sample document fragments in the first sample document based on the serial numbers to obtain a second sample document.
[0132] The processing module 530 is configured to perform sentence regularization processing on the second sample document through a first large model to obtain a document processing result corresponding to the first sample document.
[0133] The fine-tuning module 540 is configured to use the first sample document and the document processing result as training samples, and fine-tune a second large model based on the training samples to obtain a document processing model.
[0134] Wherein, the second acquisition module 520 is further configured to: acquire time information of the sample document fragments; sort the sample document fragments from earliest to latest according to the time information of the sample document fragments; and determine the serial numbers of the sample document fragments according to the sorting result.
[0135] Wherein, the processing module 530 is further configured to: determine a system prompt of the first large model; determine a preset set of constraint information for sentence regularization, where the set of constraint information includes at least constraint information related to serial numbers, constraint information for content integrity, and constraint information for limiting hallucinations; generate a prompt word for the first large model based on the system prompt and the set of constraint information; and perform sentence regularization processing on the second document through the first large model according to the prompt word to obtain a document processing result corresponding to the first sample document.
[0136] Wherein, after generating the prompt word for the first large model based on the system prompt and the set of constraint information, the device 500 is further configured to: determine at least one of a guiding example of a set scenario for document processing and a chain of thought COT inference logic for document processing; and optimize the prompt word for the first large model according to at least one of the guiding example and the COT inference logic to obtain a final prompt word for the first large model.
[0137] Among them, the processing module 530 is further configured to: under the guidance of the prompt word by the first large model, adopt the COT reasoning logic to perform sentence regularization processing on the second sample document to obtain a document processing result corresponding to the first sample document.
[0138] Among them, the processing module 530 is further configured to: perform full-text understanding on the second sample document; perform line-by-line processing on the second sample document based on the prompt word to obtain a regularized document fragment corresponding to the sample document fragment in the second sample document; output the regularized document fragments corresponding to each sample document fragment in sequence according to the serial number of the sample document fragment to obtain a document processing result corresponding to the first sample document.
[0139] Among them, the processing module 530 is further configured to: based on the prompt word, identify the core information and redundant parts of the sample document fragment in the second sample document; retain the core information and delete the redundant parts to obtain a candidate document fragment; correct the candidate document fragment on the basis of retaining the original language style to obtain the regularized document fragment.
[0140] Among them, the processing module 530 is further configured to: identify the expression mode of the core information, and in response to the expression mode being the target expression mode, optimize the expression mode of the core information.
[0141] Among them, the processing module 530 is further configured to: determine the input and output format of the second sample document; output the regularized document fragments corresponding to each sample document fragment in sequence under the constraint of the input and output format according to the serial number of the sample document fragment to obtain a document processing result corresponding to the first sample document.
[0142] Among them, after outputting the regularized document fragments corresponding to each sample document fragment in sequence according to the serial number of the sample document fragment to obtain a document processing result corresponding to the first sample document, the device 500 is further configured to: perform anomaly inspection on the document processing result according to the second sample document; in response to the first regularized document fragment with anomalies in the document processing result, determine the first sample document fragment corresponding to the first regularized document fragment from the second sample document; perform anomaly correction on the first regularized document fragment based on the first sample document fragment.
[0143] Among them, the device 500 is further configured to: in response to the sample document fragment being content not understood by the large model, output the sample document fragment as the corresponding regularized document fragment.
[0144] Among them, the processing module 530 is further configured to: determine the serial number corresponding to the regularized document fragment; query the mapping relationship between the time information and the serial number pre-constructed to obtain the time information associated with the serial number corresponding to the regularized document fragment; replace the serial number corresponding to the regularized document fragment with the time information associated with the serial number to obtain the document processing result.
[0145] Among them, after determining the serial number of the sample document fragment according to the sorting result, the apparatus 500 is further configured to: construct a mapping relationship between the time information of the sample document fragment and the serial number of the sample document fragment.
[0146] According to the training apparatus of the document processing model of the embodiments of the present disclosure, by obtaining a first sample document, splitting the first sample document to obtain sample document fragments, determining the serial numbers of the sample document fragments, and marking the sample document fragments based on the serial numbers in the first sample document to obtain a second sample document, performing sentence regularization processing on the second sample document through a first large model to obtain a document processing result corresponding to the first sample document, using the first sample document and the document processing result as training samples, and fine-tuning a second large model based on the training samples to obtain a document processing model. Thus, the present disclosure marks the sample document fragments based on serial numbers, avoiding the problem of information disorder in the document processing result, improving the readability and time alignment accuracy of the document processing result, and fine-tuning the second large model based on the training samples, reducing the resources for obtaining the document processing model, improving the accuracy and efficiency in the fine-tuning process of the second large model, enhancing the performance of the document processing model in sentence regularization, and ensuring the accuracy of the sentence regularization processing result.
[0147] Corresponding to the document processing methods provided in the above several embodiments, an embodiment of the present disclosure further provides a document processing apparatus. Since the document processing apparatus provided in the embodiments of the present disclosure corresponds to the document processing methods provided in the above several embodiments, the implementation manners of the document processing methods are also applicable to the document processing apparatus provided in this embodiment and will not be described in detail in this embodiment.
[0148] Figure 6 is a schematic structural diagram of a document processing apparatus according to an embodiment of the present disclosure.
[0149] As Figure 6 shown, the document processing apparatus 600 includes: an acquisition module 610, a marking module 620, and a document processing module 630. Among them:
[0150] The acquisition module 610 is configured to acquire a first document and split the first document to obtain a plurality of document fragments;
[0151] A tagging module 620, configured to determine the serial numbers of document fragments, and in the first document, tag the document fragments based on the serial numbers to obtain a second document;
[0152] A document processing module 630, configured to perform sentence regularization processing on the second document through a document processing model to obtain a document regularization result of the first document; wherein the document processing model is a model trained by using the method described in the first aspect.
[0153] Wherein, the tagging module 620 is further configured to: obtain the time information of the document fragment; sort the document fragments from early to late according to the time information of the document fragment; and determine the serial numbers of the document fragments according to the sorting result.
[0154] Wherein, the document processing module 630 is further configured to: perform sentence regularization processing on the second document through the document processing model under the guidance of a prompt word, and adopt a COT reasoning logic to obtain a document processing result corresponding to the first document.
[0155] According to the document processing device of the embodiments of the present disclosure, by obtaining a first document, splitting the first document to obtain a plurality of document fragments, determining the serial numbers of the document fragments, and in the first document, tagging the document fragments based on the serial numbers to obtain a second document, and performing sentence regularization processing on the second document through a document processing model to obtain a document regularization result of the first document. Thus, the present disclosure tags the document fragments based on the serial numbers, and through the document processing model under the guidance of a prompt word, adopts a COT reasoning logic to perform sentence regularization processing on the second document, which can make the document processing result more in line with human thinking habits, more natural and professional, and can enable the document processing model to better understand the text structure, reduce the situations of information omission, logical confusion and hallucination content, optimize the sentence connection, ensure the logical clarity of the document processing result, reduce the redundant expressions in the document processing result, improve the readability of the document processing result, ensure the quality of the document processing result, improve the efficiency in the sentence regularization processing while reducing the cost of the sentence regularization processing, and can be applied to the regularization of long documents and multi-speech content documents, and has outstanding effects particularly in data-intensive scenarios such as finance, law, and academia.
[0156] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0157] Figure 7FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0158] As Figure 7 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0159] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0160] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the training method or the document processing method of the document processing model. For example, in some embodiments, the training of the document processing model or the document processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the training of the document processing model or the document processing method described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the training method of the document processing model or the document processing method by any other suitable means (e.g., by means of firmware).
[0161] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0162] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0163] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0164] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0165] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0166] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0167] The present disclosure also provides a computer program product, including a computer program which, when executed by a processor, implements the training method or the document processing method of the document processing model as described above.
[0168] It should be understood that various forms of the processes shown above may be used, with steps reordered, added, or deleted. For example, the steps recited in the present disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitations are imposed herein.
[0169] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for training a document processing model, wherein, The method includes: Obtain a first sample document, and split the first sample document to obtain sample document fragments; Determine the serial numbers of the sample document fragments, and in the first sample document, mark the sample document fragments based on the serial numbers to obtain a second sample document; Perform sentence regularization processing on the second sample document through a first large model to obtain a document processing result corresponding to the first sample document; Use the first sample document and the document processing result as training samples, and fine-tune a second large model based on the training samples to obtain a document processing model.
2. The method according to claim 1, wherein The determining the serial numbers of the sample document fragments includes: Obtain the time information of the sample document fragments; Sort the sample document fragments from earliest to latest according to the time information of the sample document fragments; Determine the serial numbers of the sample document fragments according to the sorting result.
3. The method according to claim 1 or 2, wherein The performing sentence regularization processing on the second sample document through a first large model to obtain a document processing result corresponding to the first sample document includes: Determine the system prompt of the first large model; Determine a preset set of constraint information for sentence regularization, where the set of constraint information includes at least constraint information related to serial numbers, constraint information for content integrity, and constraint information for hallucination limitation; Generate a prompt word for the first large model based on the system prompt and the set of constraint information; Perform sentence regularization processing on the second document through the first large model according to the prompt word to obtain a document processing result corresponding to the first sample document.
4. The method according to claim 3, wherein, After generating the prompt word for the first large model based on the system prompt and the set of constraint information, it further includes: Determine at least one of a guiding example of a set scenario for document processing and a chain of thought (COT) reasoning logic for document processing; Optimize the prompt word of the first large model according to at least one of the guiding example and the COT reasoning logic to obtain the final prompt word of the first large model.
5. The method according to claim 4, wherein The performing sentence regularization processing on the second document through the first large model according to the prompt word to obtain a document processing result corresponding to the first sample document includes: Perform sentence regularization processing on the second sample document through the first large model under the guidance of the prompt word using the COT reasoning logic to obtain a document processing result corresponding to the first sample document.
6. The method according to claim 5, wherein, The performing sentence regularization processing on the second sample document through the first large model under the guidance of the prompt word using the COT reasoning logic to obtain a document processing result corresponding to the first sample document includes: Perform full-text understanding on the second sample document; Perform line-by-line processing on the second sample document based on the prompt word to obtain regularized document fragments corresponding to the sample document fragments in the second sample document; Output the regularized document fragments corresponding to each sample document fragment in sequence according to the serial numbers of the sample document fragments to obtain a document processing result corresponding to the first sample document.
7. The method according to claim 6, wherein Performing line-by-line processing on the second sample document based on the prompt words to obtain a regular document fragment corresponding to the sample document fragment in the second sample document, including: Based on the prompt words, identifying the core information and redundant parts of the sample document fragment in the second sample document; Retaining the core information and deleting the redundant parts to obtain a candidate document fragment; Correcting the candidate document fragment while preserving the original language style to obtain the regular document fragment.
8. The method according to claim 7, wherein The retaining of the core information includes: Identifying the expression mode of the core information, and in response to the expression mode being the target expression mode, optimizing the expression mode of the core information.
9. The method according to claim 6, wherein Sequentially outputting the regular document fragment corresponding to each sample document fragment according to the serial number of the sample document fragment to obtain the document processing result corresponding to the first sample document, including: Determining the input-output format of the second sample document; Sequentially outputting the regular document fragment corresponding to each sample document fragment under the constraint of the input-output format according to the serial number of the sample document fragment to obtain the document processing result corresponding to the first sample document.
10. The method according to claim 6, wherein After sequentially outputting the regular document fragment corresponding to each sample document fragment according to the serial number of the sample document fragment to obtain the document processing result corresponding to the first sample document, further including: Performing anomaly inspection on the document processing result according to the second sample document; In response to the first regular document fragment with anomalies in the document processing result, determining the first sample document fragment corresponding to the first regular document fragment from the second sample document; Performing anomaly correction on the first regular document fragment based on the first sample document fragment.
11. The method according to claim 6, wherein, The method further includes: In response to the sample document fragment being content not understood by the large model, outputting the sample document fragment as the corresponding regular document fragment.
12. The method according to any one of claims 6-11, wherein, Sequentially outputting the regular document fragment corresponding to each sample document fragment according to the serial number of the sample document fragment to obtain the document processing result corresponding to the first sample document, including: Determining the serial number corresponding to the regular document fragment; Querying the mapping relationship between the time information and the serial number pre-constructed to obtain the time information associated with the serial number corresponding to the regular document fragment; Replacing the serial number corresponding to the regular document fragment with the time information associated with the serial number to obtain the document processing result.
13. The method according to claim 12, wherein, After determining the serial number of the sample document fragment according to the sorting result, further including: Constructing the mapping relationship between the time information of the sample document fragment and the serial number of the sample document fragment.
14. A document processing method, wherein, The method includes: Obtaining a first document and splitting the first document to obtain multiple document fragments; Determining the serial number of the document fragment and marking the document fragment based on the serial number in the first document to obtain a second document; Performing sentence regularization processing on the second document through a document processing model to obtain the document regularization result of the first document; Wherein the document processing model is a model trained by using the method described in any one of claims 1-13.
15. The method according to claim 14, wherein, Determining the serial number of the document fragment includes: Obtaining the time information of the document fragment; Sorting the document fragments from earliest to latest according to the time information of the document fragments; Determining the serial number of the document fragment according to the sorting result.
16. The method according to claim 14, wherein, The sentence regularization processing of the second document by the document processing model to obtain the document regularization result of the first document includes: Under the guidance of the prompt words, using the COT reasoning logic, the document processing model performs sentence regularization processing on the second document to obtain the document processing result corresponding to the first document.
17. A training device for a document processing model, wherein, The device includes: A first acquisition module, configured to acquire a first sample document and split the first sample document to obtain sample document fragments; A second acquisition module, configured to determine the serial number of the sample document fragment, and mark the sample document fragment in the first sample document based on the serial number to obtain a second sample document; A processing module, configured to perform sentence regularization processing on the second sample document through a first large model to obtain a document processing result corresponding to the first sample document; A fine-tuning module, configured to use the first sample document and the document processing result as training samples, and fine-tune a second large model based on the training samples to obtain a document processing model.
18. A document processing apparatus, wherein, The device includes: An acquisition module, configured to acquire a first document and split the first document to obtain a plurality of document fragments; A marking module, configured to determine the serial number of the document fragment, and mark the document fragment in the first document based on the serial number to obtain a second document; A document processing module, configured to perform sentence regularization processing on the second document through a document processing model to obtain a document regularization result of the first document; Wherein the document processing model is a model trained by using the method described in any one of claims 1-13.
19. An electronic device, characterized in that, Including a processor and a memory; Wherein, the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to implement the method described in any one of claims 1-13 or claims 14-16.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1-13 or claims 14-16.
21. A computer program product, including a computer program, where the computer program implements the method described in any one of claims 1-13 or claims 14-16 when executed by a processor.
Citation Information
Cited By
Text processing method and device, equipment, storage medium and program product
CN121435924A