Unsupervised record extraction system and method under limited data and resource scene
Through the unsupervised transcript extraction system, paper court transcripts are converted into text, and the teacher model is used to extract structured data and enhance the data, which is then distilled to the student model. This solves the problems of limited court transcript data and insufficient computing power, and realizes efficient structured information extraction in fields such as labor dispute arbitration.
Patent Information
- Application Number
- CN202510750127.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies are limited in court transcript data and lack computing power in areas such as labor dispute arbitration, making it difficult for existing general extraction models to effectively identify and extract structured information.
An unsupervised transcript extraction system is adopted, which includes data conversion, data extraction, data enhancement and model distillation modules. Paper transcripts are converted into text through a large visual model, structured factual data is extracted using a teacher model, and the extraction capability of the student model is improved through data enhancement and model distillation.
It provides an efficient and economical structured extraction method for court transcripts in resource-constrained scenarios, improving the adaptability and practicality of the model in judicial scenarios.
Smart Images

Figure CN120670577A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an unsupervised transcript extraction system and method in a scenario with limited data and resources. Background Art
[0002] With the development of artificial intelligence, particularly natural language processing (NLP), model-based automatic information extraction methods have been widely applied in fields such as legal text processing and judicial analysis. During court hearings, court transcripts capture key content such as the parties' statements, the arbitrators' questions, and the exchange of evidence, serving as a crucial basis for fact determination and legal application. Structured extraction of factual elements from court transcripts can assist in case trials, improve arbitration efficiency, and provide a data foundation for subsequent legal services.
[0003] However, existing methods often rely on large-scale annotated corpora and high-performance computing resources, such as pre-trained language models with enormous training parameters. In real-world scenarios, particularly in areas like labor dispute arbitration, court trial data is highly sensitive and public data is scarce. Model training and deployment face the dual challenges of limited data and computing power. Furthermore, court trial texts are characterized by colloquial language, strong logical leaps, and frequent character transitions. This makes existing general-purpose extraction models perform poorly on such texts, making it difficult to effectively identify and extract structured information.
[0004] Therefore, there is an urgent need to develop a court transcript information extraction method that is suitable for low-resource environments, operates with a small number of samples and limited computing resources, and improves its adaptability and practicality in judicial scenarios. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide an unsupervised transcript extraction system and method under limited data and resource scenarios.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] An unsupervised transcript extraction system for limited data and resources, featuring: a data conversion module, a data extraction module, a data augmentation module, and a model distillation module;
[0008] The data conversion module converts the paper trial transcript file into a readable text document through the visual large model OCR method;
[0009] The data extraction module, based on the set extraction topics and extraction prompt words, uses the teacher model to extract structured factual data from the court transcript documents;
[0010] The data enhancement module decomposes the structured fact data into corresponding prompt words and topics based on the structured fact data, and reconstructs the data based on the decomposed fact data and corresponding topics to form new enhanced data;
[0011] The model distillation module uses the structured fact data extracted by the data extraction module and the enhanced data generated by the data enhancement module as distilled data, and fine-tunes the student model to improve its extraction and instruction-following capabilities. The fine-tuning data of the student model comes from the generation of the teacher model.
[0012] Furthermore, in the above-mentioned unsupervised transcript extraction system under the limited data and resource scenario, the data extraction module extracts the trial transcript data obtained in the data conversion module, and the data extraction module includes a fact label setting module and an extraction prompt word construction module. The fact label setting module summarizes the high-frequency facts existing in labor arbitration cases to obtain multiple fact labels that are commonly present and focused in the trial transcripts, and uses the labels as one of the components of the extraction prompt word construction module. The extraction prompt word construction module constructs prompt word instructions based on the set data extraction structure, and combines the fact labels set by the fact label setting module as part of the instruction prompt words to guide the teacher model to perform structured data extraction based on the readable trial transcript text obtained in the data conversion module, and obtains structured fact data extracted based on the trial transcript and the teacher model.
[0013] Furthermore, in the above-mentioned unsupervised transcript extraction system under the restricted data and resource scenario, the data enhancement module performs data enhancement on the structured fact data obtained by the data extraction module; the data enhancement module includes a data disassembly module and a data reconstruction module, the data disassembly module disassembles the structured court transcript fact data extracted by the data extraction module, and separately classifies the factual subjects and corresponding facts; the data reconstruction module performs instruction splicing and reorganization based on the disassembled structured fact data to form new structured enhanced data.
[0014] Furthermore, in the above-mentioned unsupervised transcript extraction system under the limited data and resource scenario, the model distillation module combines the original structured data obtained by the data extraction module and the enhanced structured data obtained by the data enhancement module to fine-tune the instructions of the student model, and improves the student model's instruction-following and extraction capabilities in the court transcript extraction scenario through the structured data and enhanced data generated by the teacher model.
[0015] Furthermore, in the above-mentioned unsupervised transcript extraction system under the limited data and resource scenario, the data extraction module generates labels for the data given to the model distillation module for training by extracting the teacher model, without the need for manual labeling, and generates the distillation data required by the student model in an unsupervised manner.
[0016] The unsupervised transcript extraction method under limited data and resource scenarios of the present invention includes the following steps:
[0017] First, the court transcript files are processed and OCR recognition is performed on them using a large visual model to obtain a readable text document of the transcripts.
[0018] Then, based on the types of facts that are of interest in the labor arbitration process, fact labels are set and structured extraction prompts are designed. The zero-shot capability of the teacher model is used to extract structured facts from the trial transcripts, obtaining the structured fact data corresponding to the trial transcripts.
[0019] For structured factual data, new enhanced data is generated through data decomposition and reconstruction;
[0020] Afterwards, the structured data and augmented data are integrated as fine-tuning data for supervised fine-tuning of the student model, distilling the extraction and instruction-following capabilities of the teacher model into the student model;
[0021] Finally, the fine-tuned student model is obtained and deployed as a service on a local server, thus obtaining an extraction model that can be run in a resource-constrained scenario.
[0022] Furthermore, in the above-mentioned unsupervised transcript extraction method under the limited data and resource scenario, the data conversion module converts the labor arbitration paper files to obtain a readable text data format; the data extraction module extracts structured fact data from the readable court transcript data; the data enhancement module decomposes and reconstructs the structured fact data extracted by the data extraction module to obtain enhanced data; the model distillation module constructs a supervised fine-tuning dataset based on the structured fact data obtained by the data extraction module and the data enhancement module, and fine-tunes the student model to distill the extraction and instruction-following capabilities of the teacher model into the student model, thereby obtaining an extraction model that can be deployed under the limited resource scenario.
[0023] Furthermore, in the above-mentioned unsupervised transcript extraction method under the limited data and resource scenario, the data conversion module converts the paper trial transcript file into a recognizable text document after performing OCR recognition through the visual big model;
[0024] The data extraction module includes a fact label setting module and an extraction prompt word construction module. The fact label setting module sets multiple fact theme labels designed as follows: degree of work injury compensation, degree of work injury recognition, economic compensation conditions, economic compensation time, economic compensation standards, labor remuneration calculation standards, labor remuneration payment status, labor remuneration agreed period, compensation conditions, compensation standards, compensation payment period, labor contract legality, and labor contract recognition;
[0025] The textual form of the data extraction module is expressed by the following formula. The symbol P is set as the overall input module of the teacher model, where L is the instruction prompt word component in the teacher model input module. In the L component, t n represents n topic label texts, which are presented in the brackets [] of the prompt word presentation part in the prompt word extraction construction module. C is the readable court transcript text processed in the data conversion module. F is the structured fact data set extracted based on the zero-shot capability of the teacher model. m represents the mth topic label, f m represents the fact f extracted based on the topic m; where m≤n, because not all factual topics can be extracted from the trial transcript, the formula is as follows:
[0026] D={P{L(t n ),C}→F(t m ,f m )} (1)
[0027] The data enhancement module (3) performs non-empty true subset decomposition on the extracted structured fact data. The formula is as follows:
[0028]
[0029] P* refers to a non-empty proper subset;
[0030] In the structured fact subset, the extracted prompt word L remains unchanged and the topic label text t is disassembled n The corresponding fact f n The corresponding structure in the data D, an F sub It manifests itself in the following form:
[0031] F sub ={F(t1,f1),F(t2,f2),F(t3,None)} (3)
[0032] For the disassembled structured subset, reorganize the extraction instructions required for the corresponding subset data; let t k is the selected label, t αFor the fact labels that are not covered in the fact data of this subset, new enhanced data is constructed by instruction reorganization. The enhanced data contains the enhanced prompt word instruction method and part of the fact data. The dimension of the enhanced data is smaller than the seed data. The reorganized enhanced data D aug The formula is as follows, where represents the input module of the new teacher model, under the structured fact data extracted by the teacher model, t k Indicates the facts involved in the trial transcript, t ∝ Indicates facts not mentioned in the trial transcript; k Represents the corresponding fact, the formula is as follows:
[0033] D aug ={P new {L(t k ,t α ),C}→F(t k , f k )} (4)
[0034] The model distillation module combines the structured fact data obtained by the data extraction module with the enhanced data obtained by the data enhancement module to perform supervised instruction fine-tuning on the student model to obtain a student model that can be deployed in resource-constrained scenarios.
[0035] Furthermore, in the above-mentioned unsupervised transcript extraction method under the limited data and resource scenario, the extraction prompt word construction module and the instruction prompt word extraction are designed as follows:
[0036] You are a fact extraction expert, and you extract facts entered into the trial transcript;
[0037] The data format of a fact finding contains four keys: subject, fact finding, claimant's claim, and respondent's claim;
[0038] The theme is the main summary of the content covered by the extracted facts
[0039] Fact finding is a summary of facts
[0040] The applicant's claim is the applicant's point of view
[0041] The respondent's claim is the respondent's opinion. If the respondent does not appear in court or is absent, the respondent's claim output is None;
[0042] Select topics from [Work Injury Compensation Extent, Work Injury Recognition Extent, Economic Compensation Conditions, Economic Compensation Time, Economic Compensation Standards, Labor Remuneration Calculation Standards, Labor Remuneration Payment Situation, Labor Remuneration Agreement Period, Compensation Conditions, Compensation Standards, Compensation Payment Period, Labor Contract Recognition, Labor Contract Legality]. Do not create new topics out of thin air or create illusions.
[0043] The output data format is as follows: \n[{\"Subject\":\"\",\"Facts\":\"\",\"Claim\":\"\",\"Claim\":\"\"},{\"Subject\":\"\",\"Facts\":\"\",\"Claim\":\"\",\"Claim\":\"\",\"Claim\":\"None\"}]
[0044] Output according to the given data format without any extra characters.
[0045] Compared with the prior art, the present invention has significant advantages and beneficial effects, which are specifically reflected in the following aspects:
[0046] ① The present invention converts paper transcript file data into readable court transcript data, and uses a large visual model to perform OCR automatic extraction and processing of paper text data;
[0047] ② The transcript data is summarized into fact themes, and the teacher model is used to extract the facts corresponding to the themes based on the instruction prompt words. Based on the extracted data structure, a new form of data enhancement is constructed on the basis of the original structured extracted fact data, expanding the data quantity and feature dimensions under limited data resources;
[0048] ③ We propose a new unsupervised distillation method. By constructing structured extraction data and data augmentation methods, we splice the extraction instructions corresponding to structured facts and use instruction fine-tuning to distill the extraction capabilities of the teacher model to the small model. This method can meet the needs of model deployment under limited computing resources.
[0049] ④ It can provide an efficient, economical and deployable structured extraction method for court transcripts in real judicial scenarios that lack large-scale annotated data and high-performance computing power support, which has strong practicality and application prospects.
[0050] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the specific embodiments of the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0052] Figure 1 : Schematic diagram of the module architecture of the system of the present invention;
[0053] Figure 2 : Schematic diagram of the process of the present invention;
[0054] Figure 3 : Schematic diagram of the data conversion module;
[0055] Figure 4 : Schematic diagram of the principle of data extraction module;
[0056] Figure 5 : Schematic diagram of the principle of data enhancement module;
[0057] Figure 6 : Schematic diagram of the model distillation module;
[0058] Figure 7 : Schematic diagram of enhanced data construction. DETAILED DESCRIPTION
[0059] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0060] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, directional terms and order terms are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0061] like Figure 1 、 Figure 2 The unsupervised transcript extraction system under limited data and resource scenarios includes a data conversion module 1, a data extraction module 2, a data enhancement module 3, and a model distillation module 4;
[0062] Data conversion module 1 converts paper trial transcripts into readable text documents using large-scale visual OCR. Leveraging large-scale OCR recognition, paragraph structure regularization, and character separation preprocessing technology, this module improves the conversion of paper trial transcripts into readable data formats, significantly reducing the labor cost of converting paper transcripts into readable documents.
[0063] Data extraction module 2 sets labels and extraction prompts corresponding to the trial facts. Based on the set extraction topics and extraction prompts, combined with the fact labels to be extracted and structured extraction instructions, it uses the zero-shot extraction capability of the large encoder-only model to extract structured factual data from the trial transcript documents. This generates data for supervised instruction fine-tuning of the student model in an unsupervised manner.
[0064] Data augmentation module 3, based on the structured factual data, decomposes the corresponding prompt words and themes into new augmented data. Based on the decomposed factual data and the corresponding themes, data is reconstructed to form new augmented data. Since the amount of court transcript data is very scarce, the training data directly used as small models is too scarce. Therefore, data augmentation is used to greatly increase the number and feature dimensions of the training data set to compensate for the shortage of data.
[0065] The model distillation module 4 uses the structured fact data extracted by the data extraction module 2 and the enhanced data generated by the data enhancement module 3 as distilled data, and fine-tunes the student model to improve its extraction and instruction-following capabilities. The fine-tuning data of the student model are all generated by the teacher model. The zero-shot capability and instruction-following capability of the teacher model are distilled into the model, and the zero-shot capability of the teacher model is used to construct labeled data to realize unsupervised student model training without additional data labeling costs. It realizes the improvement of the small model extraction capability in limited data and resource scenarios in an unsupervised manner, increases the data quantity and feature dimension when data is limited, and satisfies the model deployment requirements under limited computing power.
[0066] The student model and the teacher model are general models. The teacher model is a large-parameter model that requires high resources to run and usually has a larger number of parameters. The present invention selects a teacher model with a parameter amount of 72 billion, which requires more GPU resources to deploy. The advantages of this type of model are that it is more intelligent, more compliant with user instructions, and more accurate; the student model usually refers to a small-parameter model. This type of model can run with lower hardware resource consumption and requires far less computing power than the teacher model. The present invention selects a student model with a parameter amount of 1.5 billion, which consumes less resources and has insufficient ability to comply with user instructions.
[0067] Usually, the student model will choose a small model that is dozens of times smaller than the teacher model. The low parameter count of this type of model can meet the needs of model deployment in resource-constrained scenarios, and even meet the needs of model deployment on mobile devices. Through model fine-tuning, the extraction capability of the teacher model is distilled to the small model, thereby realizing model deployment in scenarios with limited computing resources.
[0068] Among them, the data extraction module 2 extracts data from the trial transcript data obtained by the data conversion module 1. The data extraction module 2 includes a fact label setting module 201 and an extraction prompt word construction module 202. The fact label setting module 201 summarizes the high-frequency facts in the labor arbitration case and obtains multiple fact labels that are commonly present and focused in the trial transcript. The labels are used as one of the components of the extraction prompt word construction module 202. The labels are in natural language form, and the instructions of the teacher model are also in natural language form. These constructed labels are used as part of the teacher model prompt words, which can specify the teacher model to extract facts for the corresponding topic. Multiple labels meet the arbitration fact extraction topic. The extraction prompt word construction module 202 constructs prompt word instructions based on the set data extraction structure, combines the fact labels set by the fact label setting module 201 as part of the instruction prompt words, and guides the teacher model to extract structured data based on the readable trial transcript text obtained in the data conversion module 1, and obtains structured fact data extracted based on the trial transcript and the teacher model.
[0069] The data enhancement module 3 enhances the structured fact data obtained by the data extraction module 2; the data enhancement module 3 includes a data disassembly module 301 and a data reconstruction module 302. The data disassembly module 301 disassembles the structured court transcript fact data extracted by the data extraction module 2 and separately classifies the factual subject and the corresponding facts; the data reconstruction module 302 performs instruction splicing and reorganization based on the disassembled structured fact data to form new structured enhanced data.
[0070] Model distillation module 4 combines the original structured data obtained by data extraction module 2 and the enhanced structured data obtained by data enhancement module 3 to fine-tune the instructions of the student model. The structured data and enhanced data generated by the teacher model improve the student model's instruction-following and extraction capabilities in the court transcript extraction scenario.
[0071] The data extraction module 2 generates the data for training the model distillation module 4. The labels are generated by the teacher model without manual labeling, and the distillation data required by the student model is generated in an unsupervised manner.
[0072] The unsupervised transcript extraction method under limited data and resource scenarios includes the following steps:
[0073] First, the court transcript files are processed and OCR recognition is performed on them using a large visual model to obtain a readable text document of the transcripts.
[0074] Then, based on the types of facts that are of interest in the labor arbitration process, fact labels are set and structured extraction prompts are designed. The zero-shot capability of the teacher model is used to extract structured facts from the trial transcripts, obtaining the structured fact data corresponding to the trial transcripts.
[0075] For structured factual data, new enhanced data is generated through data decomposition and reconstruction;
[0076] Afterwards, the structured data and augmented data are integrated as fine-tuning data for supervised fine-tuning of the student model, distilling the extraction and instruction-following capabilities of the teacher model into the student model;
[0077] Finally, the fine-tuned student model is obtained and deployed as a service on a local server, thus obtaining an extraction model that can be run in a resource-constrained scenario.
[0078] The data conversion module 1 converts the labor arbitration paper files into a readable text data format; Figure 3 The data conversion module 1 starts processing from the paper file of the trial transcript, takes the image scan of the trial transcript as the input of the visual model, performs OCR recognition, and then performs automated script splicing on the recognized text files to obtain readable trial transcript text data.
[0079] The data extraction module 2 extracts structured fact data from the readable court transcript data; Figure 4 The teacher model receives three types of combined input: ① readable court transcript text; ② multiple factual topic tags set by fact labels; ③ extraction prompts that require the teacher model to extract and output structured facts; the above three types of data are used together as input to the teacher model to guide the teacher model to extract structured factual information from the court transcript; the teacher model generates corresponding structured fact data based on the input court transcript and extraction instructions; the structured fact data covers the key factual elements in the trial and forms a standardized data representation.
[0080] The data enhancement module 3 decomposes and reconstructs the structured fact data extracted by the data extraction module 2 to obtain enhanced data; Figure 5On the basis of obtaining structured extracted data, data enhancement processing is performed on the structured fact data; by disassembling and reorganizing the fact subject labels and the corresponding facts, enhanced data different from the original structured fact data is formed; the structured fact data (i.e., seed data) is merged with the enhanced fact data to generate a "seed + enhancement" data set, which is used for the subsequent distillation process of the student model.
[0081] The model distillation module 4 constructs a supervised fine-tuning dataset based on the structured fact data obtained by the data extraction module 2 and the data enhancement module 3, and fine-tunes the student model to distill the extraction and instruction-following capabilities of the teacher model to the student model, thereby obtaining an extraction model that can be deployed in resource-constrained scenarios. Figure 6 , receiving the generated structured fact data and its corresponding augmented fact data, and using them together to train the student model. The student model is a small language model to be compressed and optimized, and its initial parameters can come from a pre-trained model or a lightweight model structure; the structured fact data and augmented data are input into the student model together for instruction fine-tuning training; during the fine-tuning process, an instruction template is introduced, combined with the factual content in the data, to guide the model to learn output behaviors guided by user needs; after fine-tuning is completed, the optimized student model is further deployed as a "resource-constrained deployment model", which can run efficiently on edge devices or local environments with limited computing power; this model retains the key reasoning and fact extraction capabilities of the teacher model, and has a lower parameter scale and reasoning latency.
[0082] Taking a court trial transcript as an example, a paper court trial transcript file is converted into a readable document through the data conversion module (1). The document-based court trial transcript is output by the model extraction module (2) as follows: [
[0084] {
[0085] "Topic":"Work Injury Compensation Level",
[0086] "Facts found": "The applicant requested the respondent to pay a one-time disability allowance of RMB 47,600, a one-time disability employment allowance of RMB 30,000, a one-time disability employment allowance of RMB 15,000, and transportation expenses of RMB 60, totaling RMB 92,660."
[0087] "Applicant's claim": "The applicant requests the respondent to pay a one-time disability allowance of RMB 47,600, a one-time disability employment allowance of RMB 30,000, a one-time disability employment allowance of RMB 15,000, and transportation expenses of RMB 60, totaling RMB 92,660."
[0088] The respondent claimed: "The respondent agreed to pay the applicant a one-time disability allowance, a one-time work-related injury medical allowance, a one-time disability employment allowance, and transportation expenses totaling 69,000 yuan. 34,000 yuan has been paid, and the remaining 35,000 yuan is due before March 31, 2021. If payment is not made within the deadline, the applicant may apply for compulsory execution based on 90,500 yuan deducted from the amount already paid."
[0089] },
[0090] {
[0091] "Subject":"Compensation Payment Deadline",
[0092] "Facts": "The two parties reached an agreement that the respondent would pay 34,000 yuan on January 20, 2021, and the remaining 35,000 yuan would be paid before March 31, 2021."
[0093] "Applicant's claim": "The respondent paid 34,000 yuan on January 20, 2021, and the remaining 35,000 yuan was due before March 31, 2021. If payment is overdue, the applicant may apply for compulsory execution based on 90,500 yuan, minus the amount already paid."
[0094] The respondent claimed: "The respondent agreed to pay 34,000 yuan on January 20, 2021, and the remaining 35,000 yuan to be paid before March 31, 2021. If payment is not made within the deadline, the applicant may apply for compulsory execution based on 90,500 yuan, minus the amount already paid."
[0095] },
[0096] {
[0097] "Topic":"Labor Contract Determination",
[0098] "Facts found": "The labor relationship between the two parties was terminated on August 4, 2020."
[0099] "Applicant's claim": "The labor relationship between the two parties was terminated on August 4, 2020."
[0100] The respondent claimed that the labor relationship between the two parties was terminated on August 4, 2020.
[0101] } ]
[0103] The above structured fact data is the output of the teacher model after receiving the instruction prompt word and the textual trial transcript; the structured output is used as the input of the data enhancement module (3), and three factual themes appearing in the trial transcript are obtained: the degree of work-related injury compensation, the compensation payment period, and the labor contract recognition; the subset data generated based on the above data are listed as follows: [degree of work-related injury compensation, the compensation payment period] [compensation payment period, the labor contract recognition] [degree of work-related injury compensation, the labor contract recognition] [degree of work-related injury compensation] [compensation payment period] [labor contract recognition]; based on these subsets, the enhanced data is spliced. Here, the subset [degree of work-related injury compensation, the labor contract recognition] is taken as an example to demonstrate how to construct the enhanced data: splice [degree of work-related injury compensation, the labor contract recognition] into the extraction prompt word construction module (202) in the data extraction module (2), replace the multiple themes in the original [], and at the same time, based on the three themes contained in the transcript, randomly add some themes that do not exist in the transcript to construct a new prompt word module. The facts of the enhanced data are the facts contained in [degree of work-related injury compensation, the labor contract recognition]: [
[0105] {
[0106] "Topic":"Work Injury Compensation Level",
[0107] "Facts found": "The applicant requested the respondent to pay a one-time disability allowance of RMB 47,600, a one-time disability employment allowance of RMB 30,000, a one-time disability employment allowance of RMB 15,000, and transportation expenses of RMB 60, totaling RMB 92,660."
[0108] "Applicant's claim": "The applicant requests the respondent to pay a one-time disability allowance of RMB 47,600, a one-time disability employment allowance of RMB 30,000, a one-time disability employment allowance of RMB 15,000, and transportation expenses of RMB 60, totaling RMB 92,660."
[0109] The respondent claimed: "The respondent agreed to pay the applicant a one-time disability allowance, a one-time work-related injury medical allowance, a one-time disability employment allowance, and transportation expenses totaling 69,000 yuan. 34,000 yuan has been paid, and the remaining 35,000 yuan is due before March 31, 2021. If payment is not made within the deadline, the applicant may apply for compulsory execution based on 90,500 yuan deducted from the amount already paid."
[0110] },
[0111] "Topic":"Labor Contract Determination",
[0112] "Facts found": "The labor relationship between the two parties was terminated on August 4, 2020."
[0113] "Applicant's claim": "The labor relationship between the two parties was terminated on August 4, 2020."
[0114] The respondent claimed that the labor relationship between the two parties was terminated on August 4, 2020.
[0115] } ]
[0117] The enhanced data processed by the data enhancement module (3) contains the following structure: a new prompt word module, text court transcripts, and new structured factual data. The enhanced data and structured data are used as training data for the student model in the model distillation module (4), where the structured prompt words serve as prompt words during model distillation, the text court transcripts serve as model input, and the structured fact book serves as model output. After passing through the model distillation module, the student model can have the extraction and instruction-following capabilities of the teacher model.
[0118] like Figure 7 As shown, the data conversion module 1 converts the paper trial transcript file into a recognizable text document after performing OCR recognition through the visual large model;
[0119] The data extraction module 2 includes a fact label setting module 201 and an extraction prompt word construction module 202. The fact label setting module 201 sets multiple fact theme labels designed as follows: degree of work injury compensation, degree of work injury recognition, economic compensation conditions, economic compensation time, economic compensation standards, labor remuneration calculation standards, labor remuneration payment status, labor remuneration agreed period, compensation conditions, compensation standards, compensation payment period, labor contract legality, and labor contract recognition;
[0120] Extract prompt word construction module 202, extract instruction prompt word design as follows:
[0121] You are a fact extraction expert, and you extract facts entered into the trial transcript;
[0122] The data format of a fact finding contains four keys: subject, fact finding, claimant's claim, and respondent's claim;
[0123] The theme is the main summary of the content covered by the extracted facts
[0124] Fact finding is a summary of facts
[0125] The applicant's claim is the applicant's point of view
[0126] The respondent's claim is the respondent's opinion. If the respondent does not appear in court or is absent, the respondent's claim output is None;
[0127] The topics were selected from [work injury compensation level, work injury recognition level, economic compensation conditions, economic compensation time, economic compensation standards, labor remuneration calculation standards, labor remuneration payment situation, labor remuneration agreement period, compensation conditions, compensation standards, compensation payment period, labor contract recognition, labor contract legality]. No new topics were generated out of thin air, and no illusions were created.
[0128] The output data format is as follows: \n[{\"Subject\":\"\",\"Facts\":\"\",\"Claim\":\"\",\"Claim\":\"\"},{\"Subject\":\"\",\"Facts\":\"\",\"Claim\":\"\",\"Claim\":\"\",\"Claim\":\"None\"}]
[0129] Output according to the given data format without any extra characters;
[0130] Through instruction extraction, all factual topics to be extracted are spliced together using a set prompt word, which serves as the extraction instruction for the teacher model. The court transcript is used as user input, and the extraction instruction is input as the prompt word. The court transcript fact text corresponding to the fact label is extracted. This instruction extraction method allows the teacher model to extract structured factual data. By leveraging the teacher model's zero-shot capability, label data is obtained in this way, saving the cost of manual labeling.
[0131] The textual expression of the data extraction module 2 is expressed by the following formula, where symbol P is the overall input module of the teacher model, where L is the instruction prompt word component in the teacher model input module, and in the L component, t n represents n topic label texts, which are presented in the brackets [] of the prompt word presentation part in the prompt word extraction and construction module 202. C is the readable court transcript text processed in the data conversion module 1; F is the structured fact data set extracted based on the zero-shot capability of the teacher model, t m represents the mth topic label, f m represents the fact f extracted based on the topic m; where m≤n, because not all factual topics can be extracted from the trial transcript text, the formula is as follows:
[0132] D={P{L(t n ),C}→F(t m ,f m )} (1)
[0133] The data enhancement module (3) performs non-empty true subset decomposition on the extracted structured fact data. The formula is as follows:
[0134]
[0135] P* refers to a non-empty proper subset;
[0136] In the structured fact subset, the extracted prompt word L remains unchanged and the topic label text t is disassembled n The corresponding fact f n The corresponding structure in the data D, an F sub It manifests itself in the following form:
[0137] F sub ={F(t1,f1),F(t2,f2),F(t3,None)} (3)
[0138] For the disassembled structured subset, reorganize the extraction instructions required for the corresponding subset data; let t k is the selected label, t α For the fact labels that are not covered in the fact data of this subset, new enhanced data is constructed by instruction reorganization. The enhanced data contains the enhanced prompt word instruction method and part of the fact data. The dimension of the enhanced data is smaller than the seed data. The reorganized enhanced data D aug The formula is as follows, where represents the input module of the new teacher model, under the structured fact data extracted by the teacher model, t k Indicates the facts involved in the trial transcript, t ∝ Indicates facts not mentioned in the trial transcript; k To express the corresponding facts, the formula is as follows;
[0139] D aug ={P new {L(t k , t α ),C}→F(t k ,f k )} (4)
[0140] Model distillation module 4, based on the structured fact data obtained by data extraction module 2 and the enhanced data obtained by data enhancement module 3, performs supervised instruction fine-tuning on the student model to obtain a student model that can be deployed in resource-constrained scenarios; the supervised fine-tuning data used by the student model comes from the teacher model generation and the original text of the court transcript, without the need for additional manual labeling, thus realizing an unsupervised model training method.
[0141] In summary, the present invention converts paper trial transcript file data into readable form, and uses a large visual model to perform OCR automatic extraction and processing of paper text data.
[0142] The court transcript data is summarized into factual themes, and the teacher model is used to extract the facts corresponding to the themes based on the instruction prompt words. Based on the extracted data structure, a new form of data enhancement is constructed on the basis of the original structured extracted fact data, expanding the data quantity and feature dimensions under limited data resources;
[0143] A new unsupervised distillation method is proposed. By constructing structured extraction data and data augmentation methods, splicing the extraction instructions corresponding to structured facts, and using instruction fine-tuning, the extraction capability of the teacher model is distilled to the small model. In this way, model deployment under limited computing resources is achieved.
[0144] It can provide an efficient, economical and deployable structured extraction method for court transcripts in real judicial scenarios that lack large-scale annotated data and high-performance computing power support, and has strong practicality and application prospects.
[0145] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Various modifications and variations are readily apparent to those skilled in the art. Any modifications, equivalent substitutions, improvements, and the like made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention. It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it need not be further defined or explained in subsequent figures.
[0146] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
[0147] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
Claims
1. Unsupervised transcript extraction system for limited data and resources, characterized by: It includes a data conversion module (1), a data extraction module (2), a data enhancement module (3) and a model distillation module (4); The data conversion module (1) converts the paper trial transcript file into a readable text document through the visual large model OCR method; The data extraction module (2) extracts structured factual data from the court trial transcript using a teacher model based on the set extraction topics and extraction prompt words; The data enhancement module (3) decomposes the structured fact data into corresponding prompt words and topics based on the structured fact data, and reconstructs the data based on the decomposed fact data and the corresponding topics to form new enhanced data; The model distillation module (4) uses the structured fact data extracted by the data extraction module (2) and the enhanced data generated by the data enhancement module (3) as distillation data to fine-tune the student model to improve the student model's extraction and instruction-following capabilities. The fine-tuning data of the student model are all generated by the teacher model.
2. The unsupervised transcript extraction system in a limited data and resource scenario according to claim 1 is characterized by: The data extraction module (2) extracts data from the court trial transcript data obtained in the data conversion module (1). The data extraction module (2) includes a fact label setting module (201) and an extraction prompt word construction module (202). The fact label setting module (201) summarizes the high-frequency facts existing in the labor arbitration case to obtain multiple fact labels that are commonly present and focused in the court trial transcript, and uses the labels as one of the components of the extraction prompt word construction module (202). The extraction prompt word construction module (202) constructs a prompt word instruction based on the set data extraction structure, combines the fact label set by the fact label setting module (201) as a part of the instruction prompt word, and guides the teacher model to extract structured data based on the readable court trial transcript text obtained in the data conversion module (1), and obtains structured fact data extracted based on the court trial transcript and the teacher model.
3. The unsupervised transcript extraction system in a limited data and resource scenario according to claim 1 is characterized by: The data enhancement module (3) performs data enhancement on the structured fact data obtained by the data extraction module (2); the data enhancement module (3) includes a data decomposition module (301) and a data reconstruction module (302); the data decomposition module (301) decomposes the structured court trial transcript fact data extracted by the data extraction module (2) and separately classifies the fact subject and the corresponding facts; the data reconstruction module (302) performs instruction splicing and reorganization based on the decomposed structured fact data to form new structured enhanced data.
4. The unsupervised transcript extraction system in a limited data and resource scenario according to claim 1 is characterized by: The model distillation module (4) combines the original structured data obtained by the data extraction module (2) and the enhanced structured data obtained by the data enhancement module (3) to fine-tune the instructions of the student model, and improves the instruction-following and extraction capabilities of the student model in the trial transcript extraction scenario through the structured data and enhanced data generated by the teacher model.
5. The unsupervised transcript extraction system in a limited data and resource scenario according to claim 1 is characterized by: The data extraction module (2) generates labels for the data given to the model distillation module (4) for training by extracting the data from the teacher model, and generates the distillation data required by the student model in an unsupervised manner.
6. Unsupervised transcript extraction method in a limited data and resource scenario, characterized by: The following steps are involved: First, the court transcript files are processed and OCR recognition is performed on them using a large visual model to obtain a readable text document of the transcripts. Then, based on the types of facts that are of interest in the labor arbitration process, fact labels are set and structured extraction prompts are designed. The zero-shot capability of the teacher model is used to extract structured facts from the trial transcripts, obtaining the structured fact data corresponding to the trial transcripts. For structured factual data, new enhanced data is generated through data decomposition and reconstruction; Afterwards, the structured data and augmented data are integrated as fine-tuning data for supervised fine-tuning of the student model, distilling the extraction and instruction-following capabilities of the teacher model into the student model; Finally, the fine-tuned student model is obtained and deployed as a service on a local server, thus obtaining an extraction model that can be run in a resource-constrained scenario.
7. The unsupervised transcript extraction method in a limited data and resource scenario according to claim 6 is characterized by: The data conversion module (1) converts the labor arbitration paper files to obtain readable text data format; the data extraction module (2) extracts structured fact data from the readable court transcript data; the data enhancement module (3) decomposes and reconstructs the structured fact data extracted by the data extraction module (2) to obtain enhanced data; the model distillation module (4) constructs a supervised fine-tuning dataset based on the structured fact data obtained by the data extraction module (2) and the data enhancement module (3), and fine-tunes the student model to distill the extraction and instruction-following capabilities of the teacher model to the student model, thereby obtaining an extraction model that can be deployed in resource-constrained scenarios.
8. The unsupervised transcript extraction method in a limited data and resource scenario according to claim 6, characterized in that: The data conversion module (1) converts the paper trial transcript file into a recognizable text document after OCR recognition through the visual big model; The data extraction module (2) includes a fact label setting module (201) and an extraction prompt word construction module (202). The fact label setting module (201) sets a plurality of fact subject labels designed as follows: degree of work injury compensation, degree of work injury recognition, economic compensation conditions, economic compensation time, economic compensation standards, labor remuneration calculation standards, labor remuneration payment situation, labor remuneration agreed period, compensation conditions, compensation standards, compensation payment period, legality of labor contract, and labor contract recognition; The textual expression of the data extraction module (2) is expressed by the following formula, where symbol P is the overall input module of the teacher model, where L is the instruction prompt word component in the teacher model input module, and in the L component, t n represents n topic label texts, which are presented in the brackets [ ] of the prompt word presentation part in the prompt word extraction and construction module (202), and C is the readable court transcript text processed in the data conversion module (1); F is a set of structured fact data extracted based on the zero-shot capability of the teacher model, t m represents the mth topic label, f m represents the fact f extracted based on the topic m; where m≤n, because not all factual topics can be extracted from the trial transcript, the formula is as follows: D={P{L(t n ),C}→F(t m ,f m )} (1) The data enhancement module (3) performs non-empty true subset decomposition on the extracted structured fact data. The formula is as follows: P* refers to a non-empty proper subset; In the structured fact subset, the extracted prompt word L remains unchanged and the topic label text t is disassembled n The corresponding fact f n The corresponding structure in the data D, an F sub It manifests itself in the following form: F sub ={F(t1,f1),F(t2,f2),F(t3,None)} (3) For the disassembled structured subset, reorganize the extraction instructions required for the corresponding subset data; let t k is the selected label, t α For the fact labels that are not covered in the fact data of this subset, new enhanced data is constructed by instruction reorganization. The enhanced data contains the enhanced prompt word instruction method and part of the fact data. The dimension of the enhanced data is smaller than the seed data. The reorganized enhanced data D aug The formula is as follows, where represents the input module of the new teacher model, under the structured fact data extracted by the teacher model, t k Indicates the facts involved in the trial transcript, t ∝ Indicates facts not mentioned in the trial transcript; k Represents the corresponding fact, the formula is as follows: D aug ={P new {L(t k ,t α ),C}→F(t k ,f k )} (4) The model distillation module (4) combines the structured fact data obtained by the data extraction module (2) and the enhanced data obtained by the data enhancement module (3) to fine-tune the student model with supervised instructions to obtain a student model that can be deployed in resource-constrained scenarios.
9. The unsupervised transcript extraction method in a limited data and resource scenario according to claim 8, characterized in that: Extract prompt word construction module (202), extract instruction prompt word design is as follows: You are a fact extraction expert, and you extract facts entered into the trial transcript; The data format of a fact finding contains four keys: subject, fact finding, claimant's claim, and respondent's claim; The theme is the main summary of the content covered by the extracted facts Fact finding is a summary of facts The applicant's claim is the applicant's point of view The respondent's claim is the respondent's opinion. If the respondent does not appear in court or is absent, the respondent's claim output is None; Select topics from [Work Injury Compensation Extent, Work Injury Recognition Extent, Economic Compensation Conditions, Economic Compensation Time, Economic Compensation Standards, Labor Remuneration Calculation Standards, Labor Remuneration Payment Situation, Labor Remuneration Agreement Period, Compensation Conditions, Compensation Standards, Compensation Payment Period, Labor Contract Recognition, Labor Contract Legality]. Do not create new topics out of thin air or create illusions. The output data format is as follows: \n[{\"Subject\":\"\",\"Facts\":\"\",\"Claim\":\"\",\"Claim\":\"\"},{\"Subject\":\"\",\"Facts\":\"\",\"Claim\":\"\",\"Claim\":\"\",\"Claim\":\"None\"}] Output according to the given data format without any extra characters.