Large language model disability identification method and device and electronic equipment

Through the methods of feature convolution classification, vector similarity matching and knowledge base search matching, the disability description text is classified, which solves the problem of insufficient accuracy of vector similarity method in the prior art and improves the accuracy of disability identification.

CN120216693APending Publication Date: 2025-06-27ZHONGNAN UNIVERSITY OF ECONOMICS AND LAW +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510219316.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, the vector similarity method is limited by the embedded model performance, making it difficult to obtain accurate prompt words, resulting in a low accuracy rate of disability identification.

Method used

Classify the identification disability description text through three methods: feature convolution classification, vector similarity matching and knowledge base search matching, and design the prompt word using the classification results as reference information.

Benefits of technology

It effectively improves the accuracy of prompt words, and thus improves the accuracy of disability identification of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216693A_ABST
    Figure CN120216693A_ABST
Patent Text Reader

Abstract

The invention provides a large language model disability identification method and device and electronic equipment, and belongs to the technical field of natural language processing, and the method comprises the steps: obtaining a to-be-identified disability description text, carrying out text vectorization to obtain a disability text vector, carrying out feature convolution classification on the disability text vector to obtain a first classification result, and obtaining a second classification result; performing similarity matching with a preset vectorization knowledge base to obtain a second classification result, and performing retrieval matching with a preset mapping knowledge base to obtain a third classification result; and according to the first classification result, the second classification result and the third classification result, constructing a disability identification cue word, and inputting a large language model to carry out question and answer mode disability identification to obtain a disability identification result. According to the method, the to-be-identified disability description text is classified in three modes of feature convolution point class, vector similarity matching and knowledge base retrieval matching, and cue word design is performed by taking the classified disability description text as reference information, so that the cue word accuracy can be effectively improved, and the accuracy of large language model disability identification is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a method, device and electronic device for disability appraisal using a large language model. Background Art

[0002] With the development of large language models, various generative large language models emerge in an endless stream and show satisfactory performance in multiple general fields. Using a large language model for disability appraisal means inputting the disability scenario text into the large language model and asking questions about the large language model to obtain the judgment result of the large language model on the disability. However, since the accuracy of the large language model depends on the training corpus, for disability appraisal, because the training corpus often does not involve professional knowledge, the performance of the large language model in this field is generally average. Therefore, the existing technology usually retrieves and enhances the text input into the large language model, that is, retrieves relevant content in a pre-set knowledge base for the input text, and then constructs a prompt to ask questions to the large language model.

[0003] The existing technology usually vectorizes the knowledge base, and then calculates the semantic similarity between the embedded vectors to return the top N most similar texts, so as to retrieve and enhance the input text. However, for the disability appraisal task, due to the large difference in the length and complexity of the content between the texts in the knowledge base and the small difference in the word usage style, limited by the performance of the embedding model for text vectorization, there may be a situation where the vector calculation does not match the actual text, and it is difficult to obtain accurate prompts through the vector similarity method. Therefore, the accuracy of disability appraisal is relatively low.

[0004] Therefore, the existing technology has the technical problem that the vector similarity method is limited by the performance of the embedding model, it is difficult to obtain accurate prompts, resulting in a relatively low accuracy of disability appraisal. Summary of the Invention

[0005] In view of this, it is necessary to provide a method, device and electronic device for disability appraisal using a large language model to solve the technical problem in the existing technology that the vector similarity method is limited by the performance of the embedding model, it is difficult to obtain accurate prompts, resulting in a relatively low accuracy of disability appraisal.

[0006] To solve the above technical problem, on the one hand, the present invention provides a method for disability appraisal using a large language model, including: Obtain the disability description text to be appraised, perform text vectorization on the disability description text to be appraised to obtain a disability text vector, perform feature convolution classification on the disability text vector to obtain a first classification result, perform similarity matching on the disability text vector and a pre-set vectorized knowledge base to obtain a second classification result, and perform retrieval matching on the disability description text to be appraised and a pre-set graphical knowledge base to obtain a third classification result; Construct a disability appraisal prompt based on the first classification result, the second classification result, and the third classification result, and input the disability appraisal prompt into a large language model for disability appraisal in a question-and-answer mode to obtain a disability appraisal result.

[0007] In a possible implementation, the text to be appraised for disability is vectorized to obtain a disability text vector, including: Perform word decomposition embedding, position embedding, and separation embedding on the text to be appraised for disability based on the Chinese word embedding method to obtain word decomposition embedding features, position embedding features, and separation embedding features; Combine the word decomposition embedding features, position embedding features, and separation embedding features to obtain a disability text vector.

[0008] In a possible implementation, the first classification result includes an injury site classification result and an injury degree classification result. Perform feature convolution classification on the disability text vector to obtain the first classification result, including: Perform injury site classification on the disability text vector after sequentially performing feature splitting, first feature encoding, and first convolution pooling operations to obtain an injury site classification result; Perform injury degree classification on the disability text vector after sequentially performing feature splitting, second feature encoding, and second convolution pooling operations to obtain an injury degree classification result.

[0009] In a possible implementation, a preset vectorized knowledge base is obtained by vectorizing the injury appraisal standard text. Vectorize the injury appraisal standard text, including: Construct an injury category index for each injury appraisal standard text according to the injury site category and the injury degree category; Vectorize each injury appraisal standard text to obtain an appraisal standard vector, and construct a preset vectorized knowledge base according to the appraisal standard vector and the corresponding injury category index.

[0010] In a possible implementation, perform similarity matching between the disability text vector and the preset vectorized knowledge base to obtain the second classification result, including: Determine the Euclidean distance between the disability text vector and each appraisal standard vector in the preset disability appraisal knowledge base based on the Euclidean distance method; Sort the appraisal standard vectors according to the Euclidean distance and determine the second classification result according to the similarity sorting result.

[0011] In a possible implementation, a preset graph-based knowledge base is obtained by constructing an appraisal standard triple for the injury appraisal standard text. Construct an appraisal standard triple for the injury appraisal standard text, including: Construct an injury category index for each injury appraisal standard text according to the injury site category and the injury degree category; Extract the disability circumstances text based on the injury location and injury manifestations in the injury appraisal standard text, form an appraisal standard triple with the disability circumstances text and the corresponding injury category index, and construct a preset graph-based knowledge base according to the appraisal standard triple.

[0012] In a possible implementation, retrieve and match the disability description text to be appraised with the preset graph-based knowledge base to obtain a third classification result, including: Extract the disability circumstances to be appraised based on the injury location and injury manifestations in the disability description text to be appraised; Based on the TOP-N retrieval method, retrieve and match the disability circumstances to be appraised and the appraisal standard triples in the preset graph-based knowledge base to obtain a third classification result.

[0013] In a possible implementation, construct a disability appraisal prompt word according to the first classification result, the second classification result, and the third classification result, including: Use the average method for the first classification result, the second classification result, and the third classification result to determine the injury category judgment result and the injury degree judgment result; Construct a disability appraisal prompt word according to the disability description text to be appraised, the injury category judgment result, the injury degree judgment result, and the preset prompt word template.

[0014] On the other hand, the present invention also provides a large language model disability appraisal device, including: A retrieval enhancement unit, configured to obtain the disability description text to be appraised, perform text vectorization on the disability description text to be appraised to obtain a disability text vector, perform feature convolution classification on the disability text vector to obtain a first classification result, perform similarity matching on the disability text vector and a preset vectorized knowledge base to obtain a second classification result, and retrieve and match the disability description text to be appraised with the preset graph-based knowledge base to obtain a third classification result; A question-and-answer appraisal unit, constructs a disability appraisal prompt word according to the first classification result, the second classification result, and the third classification result, and inputs the disability appraisal prompt word into the large language model for question-and-answer mode disability appraisal to obtain a disability appraisal result.

[0015] On the other hand, the present invention also provides an electronic device, including a memory and a processor, wherein, The memory is used to store a computer program; The processor is coupled to the memory and is configured to execute the computer program to implement the steps in the above-mentioned large language model disability appraisal method.

[0016] The beneficial effects of the above implementation method are as follows: in the large language model disability appraisal method provided by the present invention, the disability description text to be appraised is classified through three methods: feature convolution classification, vector similarity matching, and knowledge base retrieval matching, and the classification results are used as reference information for prompt word design, which can effectively improve the accuracy of prompt words, and further improve the accuracy of large language model disability appraisal. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 It is a flowchart of an embodiment of the large language model disability appraisal method provided by the present invention; Figure 2 It is a flowchart of the text vectorization in the embodiment of the present invention; Figure 3 It is a flowchart of the feature convolution classification in the embodiment of the present invention; Figure 4 It is a flowchart of constructing a preset vectorized knowledge base in the embodiment of the present invention; Figure 5 It is a flowchart of the vectorized knowledge base similarity matching in the embodiment of the present invention; Figure 6 It is a flowchart of constructing a preset graph knowledge base in the embodiment of the present invention; Figure 7 It is a flowchart of the graph knowledge base retrieval matching in the embodiment of the present invention; Figure 8 It is a flowchart of constructing a disability appraisal prompt word in the embodiment of the present invention; Figure 9 It is a structural diagram of an embodiment of the large language model disability appraisal device provided by the present invention; Figure 10 It is a structural diagram of an embodiment of the electronic device provided by the present invention. Detailed Embodiments

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0020] It should be understood that the drawings of the schematic diagrams are not drawn to the actual scale. The flowcharts used in the present invention illustrate the operations implemented according to some embodiments of the present invention. It should be understood that the operations of the flowchart may not be implemented in sequence, and the steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present invention.

[0021] Some of the block diagrams shown in the drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor systems and / or microcontroller systems.

[0022] Referring to "embodiments" herein means that the specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of the present invention. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0023] Figure 1 A flowchart of an embodiment of the large language model disability appraisal method provided by the present invention is as Figure 1 shown. The large language model disability appraisal method includes: S101. Obtain the disability description text to be appraised, perform text vectorization on the disability description text to be appraised to obtain a disability text vector, perform feature convolution classification on the disability text vector to obtain a first classification result, perform similarity matching between the disability text vector and a preset vectorized knowledge base to obtain a second classification result, and perform retrieval matching between the disability description text to be appraised and a preset graphical knowledge base to obtain a third classification result; S102. Construct a disability appraisal prompt word according to the first classification result, the second classification result and the third classification result, input the disability appraisal prompt word into the large language model for question-and-answer mode disability appraisal, and obtain a disability appraisal result.

[0024] Compared with the prior art, the present invention classifies the disability description text to be appraised through three methods of feature convolution classification, vector similarity matching and knowledge base retrieval matching, and uses the classification results as reference information for prompt word design, which can effectively improve the accuracy of the prompt word, and further improve the accuracy of the large language model disability appraisal.

[0025] In some embodiments of the present invention, Figure 2The flowchart of text vectorization according to an embodiment of the present invention is shown as Figure 2 follows. The disability description text to be identified is subjected to text vectorization to obtain a disability text vector, including: S201. Respectively perform word and character disassembling embedding, position embedding, and separation embedding on the disability description text to be identified based on the Chinese word embedding method to obtain a word and character disassembling embedding feature, a position embedding feature, and a separation embedding feature; S202. Combine the word and character disassembling embedding feature, the position embedding feature, and the separation embedding feature to obtain a disability text vector.

[0026] Specifically, the embodiment performs vectorization processing on the disability description text to be identified through the Chinese word embedding method, processes the text through three methods of word and character disassembling embedding, position embedding, and separation embedding, and combines the embedding features obtained by the three methods to obtain a disability text vector.

[0027] Among them, WordPiece (word and character disassembling embedding) can disassemble and encode words and divide words into multiple common sub-word units to achieve a more balanced compromise between the effectiveness of words and the flexibility of characters. Position embedding refers to encoding the position information of words into feature vectors. Since there is no fixed iteration direction in the Transformer model, it is necessary to record the position information of the input data and introduce the position relationship of words into the model. Separation embedding is used to distinguish two sentences. For example, B is the following text of A. For a sentence pair, the values corresponding to the words in the first sentence are 0, and the values corresponding to the second sentence are 1. The embodiment combines the vectors obtained by the three embedding methods to obtain a vectorized disability text vector.

[0028] In some embodiments of the present invention, the first classification result includes an injury site classification result and an injury degree classification result. Figure 3 The flowchart of feature convolution classification according to an embodiment of the present invention is shown as Figure 3 follows. The disability text vector is subjected to feature convolution classification to obtain a first classification result, including: S301. Perform injury site classification after sequentially performing feature splitting, first feature encoding, and first convolution pooling operations on the disability text vector to obtain an injury site classification result; S302. Perform injury degree classification after sequentially performing feature splitting, second feature encoding, and second convolution pooling operations on the disability text vector to obtain an injury degree classification result.

[0029] Specifically, considering that although deep learning models have poor interpretability, they often have excellent results in classification problems. Therefore, in this embodiment, deep learning is used as the first retrieval enhancement generation method. By establishing a RoBERTa-CNN model to perform feature convolution classification on the disability text vector, a first classification result is obtained as a reference for subsequent generation of prompt words. Among them, the RoBERTa model is trained through the whole-word masking strategy, which is mainly used to solve the problem that during the pre-training process of the model, due to the design defect of the tokenizer, only a part of a word is masked, resulting in limited model performance, and can effectively improve the model performance. In the embodiment, the input disability text vector is first split by a tokenizer and then input into the RoBERTa model. In the RoBERTa model, the features are encoded through 24 layers of Transformer encoder layers, and then the last four layers of outputs are extracted and fed into the CNN layer for convolution pooling. The CNN layer includes Conv convolution, ReLu activation function, and Pooling pooling operations. In the embodiment, the CNN captures local text features, and the RoBERTa captures global language information, which can effectively improve the classification effect of the model and extract the injury site classification result and the injury degree classification result as the first classification result.

[0030] At the same time, since the RoBERTa model has undergone a certain scale of pre-training, it can also complete the feature extraction process without a large amount of special processing. Only a small-scale fine-tuning according to the dataset is required to obtain good results in the two classification tasks of injury site classification and injury degree classification. The CNN layer can flexibly change the structure and number of layers of the CNN layer according to the training results, and the parameter adjustment is more convenient compared to the pure RoBERTa model. In the embodiment, the fine-tuning tasks are injury degree classification and site classification respectively. The injury degree classification aims to train the RoBERTa model to predict the injury degree received by the identified object, and this task is a 6-class classification task. The injury site classification aims to train the RoBERTa model to predict the injury sites involved in the disability performance, which are divided into 12 categories in total. At the same time, for a disability performance text, the identified object may have multiple parts injured at the same time. Therefore, this task includes a 6-class multi-classification task and a 12-class multi-classification task respectively.

[0031] In some embodiments of the present invention, the preset vectorized knowledge base is obtained by vectorizing the injury appraisal standard text. Figure 4 It is a schematic flowchart of constructing a preset vectorized knowledge base according to an embodiment of the present invention. As Figure 4 shown, the vectorization process of the injury appraisal standard text includes: S401. Construct an injury category index for each injury appraisal standard text according to the injury site category and the injury degree category; S402. Vectorize each injury appraisal standard text to obtain an appraisal standard vector, and construct a preset vectorized knowledge base according to the appraisal standard vector and the corresponding injury category index.

[0032] Specifically, in the process of constructing the vectorized knowledge base, in the embodiment, the "Standard for Appraising the Extent of Human Body Injury" is used as the injury appraisal standard text. First, it is necessary to structure the appraisal standard text. During the structuring process of the appraisal standard text, according to the chapter division therein, the injury parts are divided into 12 injury part categories: the head and spinal cord, the face and auricle, the auditory organ and hearing, the visual organ and vision, the neck, the chest, the abdomen, the pelvis and perineum, the spine and extremities, the hand, the body surface, and other parts. Each part category corresponds to 5 injury severity categories: first-degree serious injury, second-degree serious injury, first-degree minor injury, second-degree minor injury, and minor injury. For the circumstances that do not reach the minor injury standard, they are identified as "not constituting a minor injury" in the appraisal process. Therefore, there are a total of 6 appraisal possibilities. Then, the embodiment combines the injury part categories and the injury severity categories to construct a number of injury category indexes in two dimensions, such as "hand - second-degree minor injury", and corresponds each injury appraisal standard text to its respective category.

[0033] After structuring the appraisal standard, the embodiment vectorizes a number of injury appraisal standard texts corresponding to each index to obtain an appraisal standard vector, and then constructs a preset vectorized knowledge base together with its respective index.

[0034] In some embodiments of the present invention, Figure 5 is a schematic flowchart of the similarity matching of the vectorized knowledge base in the embodiment of the present invention. As Figure 5 shown, performing similarity matching on the disability text vector and the preset vectorized knowledge base to obtain a second classification result, including: S501. Determine the Euclidean distance between the disability text vector and each appraisal standard vector in the preset disability appraisal knowledge base based on the Euclidean distance method; S502. Sort the appraisal standard vectors according to the Euclidean distance, and determine the second classification result according to the similarity sorting result.

[0035] Specifically, the embodiment uses the semantic similarity retrieval reference result as the second retrieval enhancement generation method. After the embodiment vectorizes the text by the Chinese word embedding method, for the input content, the similarity between vectors can be calculated to retrieve the knowledge semantically similar to the input content in the constructed vectorized knowledge base. In the embodiment, the Euclidean distance is used to judge the similarity between vectors. The embedding vector of the input text is represented as , and it is necessary to find the appraisal standard vector in the vectorized knowledge base such that the and The Euclidean distance between The shortest, the Euclidean distance formula is expressed as:

[0036] Then, sort the obtained results, and use the content corresponding to the identification standard vector with the highest similarity, that is, the shortest Euclidean distance, as the second classification result.

[0037] In some embodiments of the present invention, the preset graph knowledge base is obtained by constructing identification standard triples for the injury identification standard text. Figure 6 It is a schematic flow chart of constructing the preset graph knowledge base in an embodiment of the present invention. As Figure 6 shown, constructing identification standard triples for the injury identification standard text includes: S601. Construct an injury category index for each injury identification standard text according to the injury part category and the injury degree category; S602. Extract the disability plot text according to the injury part and the injury manifestation in the injury identification standard text, form an identification standard triple with the disability plot text and the corresponding injury category index, and construct a preset graph knowledge base according to the identification standard triple.

[0038] Specifically, considering that during the vectorization process, there are differences in the text lengths of each part in the injury identification standard text. For example, the short ones only contain one or two short sentences, with a length of less than 50 words, while the long ones may contain dozens of sentences, with a total length reaching hundreds of words. Therefore, if the segmentation is performed according to a certain length limit, due to the different granularities of each dimension, the situation of "cutting too much" or "cutting too little" often occurs, which may lead to the retrieval result either including indexes of multiple dimensions or not being able to include all the standard texts under one dimension, ultimately resulting in chaotic retrieval results and the destruction of context coherence. In addition, since the "Standard for the Assessment of Human Injury" is a document with strong authority, the writing style therein is rigorous and unified, and has a high degree of homogeneity. Using this as the injury identification standard text is limited by the performance of the embedding model, and there may be a situation where the calculated embedding vector does not match the actual text. Therefore, it will affect the accuracy of the classification result obtained by vectorization matching. To improve the accuracy of the finally obtained prompt words, in addition to the vectorized knowledge base, the embodiment also constructs a graph knowledge base for retrieving and matching to obtain the third classification result.

[0039] Similar to constructing a vectorized knowledge base, similarly, to construct a graph knowledge base, it is necessary to first construct an injury category index for the injury identification standard text according to the injury site category and the injury degree category. Then, in the embodiment, the injury sites and injury manifestations in the injury identification standard text are extracted as entities, and together with the injury category index, they form an identification standard triple. For example, taking "fracture of carpal bone, metacarpal bone or phalange" as an example, three triples can be extracted and constructed: "carpal bone fracture - minor injury - hand", "metacarpal bone fracture - minor injury - hand", and "phalange fracture - minor injury - hand". Then, a preset graph knowledge base is constructed according to the obtained identification standard triples. This is used to avoid the defects of the vectorized knowledge base caused by text length differences and semantic similarity problems.

[0040] In some embodiments of the present invention, Figure 7 is a schematic flowchart of the retrieval and matching of the graph knowledge base in the embodiment of the present invention. As Figure 7 shown, the text of the disability description to be identified is retrieved and matched with the preset graph knowledge base to obtain a third classification result, including: S701. Extract the disability circumstances to be identified according to the injury site and injury manifestation in the text of the disability description to be identified; S702. Based on the TOP-N retrieval method, retrieve and match the disability circumstances to be identified with the identification standard triples in the preset graph knowledge base to obtain a third classification result.

[0041] Specifically, in the retrieval and matching of the graph knowledge base, in the embodiment, the injury site and injury manifestation in the text of the disability description to be identified are first extracted to form the disability circumstances to be identified. For example, "finger - fracture", and then the disability circumstances to be identified and the graph knowledge base are used with the TOP-N retrieval method to retrieve the triples in the graph knowledge base and return the closest identification standard triple as the third classification result. In the implementation process of TOP-N retrieval, the embodiment uses the FAISS (Facebook AI Similarity Search) method to perform similarity retrieval on the triples.

[0042] In some embodiments of the present invention, Figure 8 is a schematic flowchart of constructing a disability identification prompt word in the embodiment of the present invention. As Figure 8 shown, construct a disability identification prompt word according to the first classification result, the second classification result, and the third classification result, including: S801. Use the averaging method for the first classification result, the second classification result, and the third classification result to determine the injury category judgment result and the injury degree judgment result; S802. Construct a disability identification prompt word according to the text of the disability description to be identified, the injury category judgment result, the injury degree judgment result, and the preset prompt word template.

[0043] Specifically, after obtaining the three classification results based on the deep learning model, semantic similarity, and graph-based knowledge base, the embodiment uses the averaging method to calculate the classification mean of the three strategies, and takes the category with the highest frequency as the final retrieval result, which is returned to the large language model. For example, the judgment results of the three recall strategies are: "hand, skin; minor injury", "hand; first-degree minor injury", "skin; minor injury", then the retrieval result of "hand, skin; minor injury" is returned. Then, the retrieval result and the prompt word template are used to generate the final prompt word for asking questions to the large language model to obtain the disability appraisal result. For example, for "hand, skin; minor injury", the constructed prompt word is shown in Table 1: Table 1: Example Table of Disability Appraisal Prompt Words

[0044] Finally, the large language model performs disability appraisal in the Q&A mode on the above prompt words to obtain the disability appraisal result.

[0045] In addition, to verify the effectiveness of the proposed solution of the present invention, the embodiment also sets up a control experiment to verify the solution. The embodiment sets up five groups for comparison, namely: 1. Without using the retrieval enhancement strategy; 2. Using the vectorized knowledge base retrieval enhancement strategy; 3. Using the graph-based knowledge base retrieval enhancement strategy; 4. Using the deep learning retrieval enhancement strategy; 5. Using the comprehensive multiple retrieval enhancement strategies of the present solution.

[0046] In the classification view, there are four possibilities for the result of a single sample: True Positive (TP): The model predicts a positive example, and the true result is also a positive example.

[0047] True Negative (TN): The model predicts a negative example, and the true result is also a negative example.

[0048] False Positive (FP): The model predicts a positive example, and the true result is a negative example.

[0049] False Negative (FN): The model predicts a negative example, and the true result is a positive example.

[0050] For each category, the embodiment uses indicators such as accuracy, precision, recall rate, and F1 value as evaluation indicators.

[0051] Among them, the accuracy rate is the proportion that the judgment result of the large language model is completely consistent with the true appraisal result:

[0052] The precision rate is the proportion that among the samples predicted by the model as 1, the true category is 1:

[0053] The recall rate is the proportion where the true class is 1 and the model predicts 1:

[0054] The F1 value is the comprehensive result of precision and recall:

[0055] When evaluating all classes, the calculation method used in the embodiment is as follows: Macro average: First, count the index values for each class, and then calculate the arithmetic mean of the indices for all classes. The macro average does not consider the sample distribution, and its value can better reflect the influence of extreme values in the sample:

[0056]

[0057]

[0058] Weighted average: Determine the weights according to the distribution ratio of each class, and then sum them after weighting. The weighted average method takes into account the class imbalance, and its value is more easily affected by common classes:

[0059]

[0060]

[0061] Finally, the embodiment extracted a batch of data, conducted experiments under different strategies to verify the model performance, and obtained the results as shown in Table 2: Table 2: Model Performance Verification Result Table

[0062] It can be seen from the results that both the graph-based knowledge base enhancement strategy and the deep learning enhancement strategy can significantly improve the model performance. The comprehensive enhancement strategy of the embodiment of the present invention is further improved compared to using the graph-based knowledge base enhancement strategy and the deep learning enhancement strategy alone.

[0063] In summary, the present invention classifies the text of the disability description to be identified through three methods: feature convolution classification, vector similarity matching, and knowledge base retrieval matching, and uses the classification result as reference information for prompt word design, which can effectively improve the accuracy of prompt words, and thus improve the accuracy of disability identification of the large language model.

[0064] Based on the large language model disability assessment method provided by the present invention, the present invention also provides a large language model disability assessment device, as Figure 9 shown, including: A retrieval enhancement unit 901, configured to obtain a disability description text to be assessed, perform text vectorization on the disability description text to be assessed to obtain a disability text vector, perform feature convolution classification on the disability text vector to obtain a first classification result, perform similarity matching on the disability text vector and a preset vectorized knowledge base to obtain a second classification result, and perform retrieval matching on the disability description text to be assessed and a preset graphical knowledge base to obtain a third classification result; A question and answer assessment unit 902, configured to construct a disability assessment prompt word according to the first classification result, the second classification result, and the third classification result, input the disability assessment prompt word into the large language model for disability assessment in a question and answer mode, and obtain a disability assessment result.

[0065] The large language model disability assessment device 900 provided in the above embodiment can implement the technical solutions described in the above embodiment of the large language model disability assessment method. The specific implementation principles of the above modules or units can be referred to the corresponding content in the above embodiment of the large language model disability assessment method, and will not be elaborated here.

[0066] The present invention also provides an electronic device 1000, as Figure 10 shown, Figure 10 which is a schematic structural diagram of an embodiment of the electronic device provided by the present invention. The electronic device 1000 includes a processor 1001, a memory 1002, and a computer program stored in the memory 1002 and executable on the processor 1001. When the processor 1001 executes the program, the above-mentioned large language model disability assessment method is implemented.

[0067] As a preferred embodiment, the above-mentioned electronic device further includes a display 1003, configured to display the process of the processor 1001 executing the above-mentioned large language model disability assessment method.

[0068] Among them, the processor 1001 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 1001 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC). It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may also be a microprocessor, or the processor may also be any conventional processor, etc.

[0069] Among them, the memory 1002 can be, but is not limited to, a Random Access Memory (RAM), a Read Only Memory (ROM), a Secure Digital (SD) card, a Flash Card, etc. The memory 1002 is used to store programs. After receiving an execution instruction, the processor 1001 executes the programs. The method defined by the processes disclosed in any of the embodiments of the present invention can be applied to or implemented by the processor 1001.

[0070] Among them, the display 1003 can be an LED display screen, a liquid crystal display, a touch display, etc. The display 1003 is used to display various information of the electronic device 1000.

[0071] It can be understood that Figure 10 the structure shown is only a schematic structural diagram of the electronic device 1000, and the electronic device 1000 may further include more or fewer components than Figure 10 those shown. Figure 10 Each of the components shown in can be implemented by hardware, software, or a combination thereof.

[0072] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. A large language model disability identification method, characterized in that: include: Obtain a text description of the disability to be identified, perform text vectorization on the text description of the disability to be identified to obtain a disability text vector, perform feature convolution classification on the disability text vector to obtain a first classification result, perform similarity matching on the disability text vector and a preset vectorized knowledge base to obtain a second classification result, and perform search matching on the text description of the disability to be identified and a preset graph knowledge base to obtain a third classification result; A disability assessment prompt word is constructed according to the first classification result, the second classification result and the third classification result, and the disability assessment prompt word is input into a large language model to perform disability assessment in a question-and-answer mode to obtain a disability assessment result.

2. The large language model disability identification method according to claim 1, characterized in that: The step of performing text vectorization on the to-be-identified disability description text to obtain a disability text vector comprises: Based on the Chinese word embedding method, the text describing the disability to be identified is subjected to word decomposition embedding, position embedding and separation embedding to obtain word decomposition embedding features, position embedding features and separation embedding features; The word decomposition embedding feature, the position embedding feature and the separation embedding feature are combined to obtain a disabled text vector.

3. The large language model disability identification method according to claim 1, characterized in that: The first classification result includes an injury site classification result and an injury degree classification result. The first classification result is obtained by performing feature convolution classification on the disability text vector, including: Performing feature splitting, first feature encoding and first convolution pooling operations on the disability text vector in sequence, and then classifying the injury part to obtain an injury part classification result; The disability text vector is subjected to feature splitting, second feature encoding and second convolution pooling operations in sequence, and then the degree of injury is classified to obtain the degree of injury classification result.

4. The large language model disability identification method according to claim 1, characterized in that: The preset vectorized knowledge base is obtained by vectorizing the damage identification standard text, and the vectorizing the damage identification standard text includes: Construct the damage category index of each damage appraisal standard text according to the damage part category and damage degree category; Each of the damage identification standard texts is vectorized to obtain an identification standard vector, and a preset vectorized knowledge base is constructed based on the identification standard vector and the corresponding damage category index.

5. The large language model disability identification method according to claim 4, characterized in that: The performing similarity matching between the disabled text vector and a preset vectorized knowledge base to obtain a second classification result includes: Determine the Euclidean distance between the disability text vector and each identification standard vector in the preset disability identification knowledge base based on the Euclidean distance method; The identification standard vectors are sorted by similarity according to the Euclidean distance, and a second classification result is determined according to the similarity sorting result.

6. The large language model disability identification method according to claim 1, characterized in that: The preset graphical knowledge base is obtained by constructing an identification standard triplet for the damage identification standard text, and the identification standard triplet constructed for the damage identification standard text includes: Construct the damage category index of each damage appraisal standard text according to the damage part category and damage degree category; The disability details are extracted according to the injury location and injury manifestation in the injury identification standard text to obtain the disability details text, the disability details text and the corresponding injury category index are combined into an identification standard triple, and a preset graphical knowledge base is constructed according to the identification standard triple.

7. The large language model disability identification method according to claim 6, characterized in that: The third classification result is obtained by searching and matching the description text of the disability to be identified with a preset graphical knowledge base, including: Extracting the disability details according to the injury location and injury manifestation in the description text of the disability to be identified to obtain the disability details to be identified; Based on the TOP-N retrieval method, the disability to be identified and the identification standard triples in the preset graphical knowledge base are searched and matched to obtain the third classification result.

8. The large language model disability identification method according to claim 1, characterized in that: The step of constructing a disability assessment prompt word according to the first classification result, the second classification result and the third classification result includes: Determine an injury category judgment result and an injury degree judgment result by using an average method for the first classification result, the second classification result and the third classification result; The disability assessment prompt words are constructed according to the description text of the disability to be assessed, the injury category judgment result, the injury degree judgment result and a preset prompt word template.

9. A large language model disability assessment device, characterized in that: include: a retrieval enhancement unit, configured to obtain a description text of a disability to be identified, perform text vectorization on the description text of the disability to be identified to obtain a disability text vector, perform feature convolution classification on the disability text vector to obtain a first classification result, perform similarity matching on the disability text vector and a preset vectorized knowledge base to obtain a second classification result, and perform search matching on the description text of the disability to be identified and a preset graph knowledge base to obtain a third classification result; The question-and-answer identification unit constructs disability identification prompt words according to the first classification result, the second classification result and the third classification result, and inputs the disability identification prompt words into a large language model to perform disability identification in a question-and-answer mode to obtain a disability identification result.

10. An electronic device, characterized in that: comprising a memory and a processor, wherein: The memory is used to store computer programs; The processor, coupled to the memory, is used to execute a computer program to implement the steps in the large language model disability identification method as described in any one of claims 1 to 8.