Method and system for generating work injury auxiliary identification based on dual-channel retrieval enhancement

By constructing a dual-path retrieval enhanced generation model, and combining regulatory maps and historical cases, a multimodal data fusion assessment of work injury levels is achieved, which solves the problem of inaccurate assessment caused by single-modal data and improves the accuracy and efficiency of work injury identification.

CN121171540BActive Publication Date: 2026-04-24QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
Filing Date
2025-08-21
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing methods for assessing work-related injury disability levels rely on single-modal data, which cannot comprehensively consider all medical information of the injured party, leading to inaccurate assessment results and subjective biases, especially in complex or ambiguous cases where errors are prone to occur.

Method used

A dual-path retrieval-enhanced generation method is adopted. By constructing a structured knowledge graph of regulations and a semantic retrieval path of historical cases, and combining medical images, diagnostic texts and legal clauses, the data is encoded into a unified semantic vector space to achieve automated auxiliary assessment of work injury levels and to accurately determine disability levels by integrating multimodal data.

Benefits of technology

It improves the accuracy and efficiency of work-related injury assessment, and the generated results are both legally compliant and medically accurate, supporting the complete traceability of the judgment basis and enhancing the professionalism and interpretability of the assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121171540B_ABST
    Figure CN121171540B_ABST
Patent Text Reader

Abstract

The disclosure provides a work injury auxiliary identification method and system based on double-channel retrieval enhancement generation, relating to the technical fields of artificial intelligence and work injury auxiliary identification, comprising obtaining multi-modal data of injured person motion video, medical image and case text; fusing the semantic vectors of images, texts and videos into unified semantic representation vectors; inputting the semantic representation vectors into a double-channel retrieval enhancement generation model, after the semantic representation vectors are input into the double-channel retrieval enhancement generation model, they enter a regulation structured knowledge graph retrieval channel and a historical case semantic retrieval channel respectively, the regulation structured knowledge graph retrieval channel extracts the final representation of each node in the regulation knowledge graph, and performs regulation node path expansion to obtain a regulation graph matching basis chain; the historical case semantic retrieval channel calculates historical related cases through semantic similarity to obtain case core abstract information. The disclosure improves the auxiliary evaluation efficiency and enhances the explainability of the results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of artificial intelligence and work injury assistance identification technology, specifically to a work injury assistance identification method and system based on dual-pathway retrieval enhancement generation. Background Technology

[0002] The statements in this section are merely background information relating to this disclosure and do not necessarily constitute prior art.

[0003] With the rapid development of artificial intelligence technology, the biomedical field has ushered in revolutionary technological innovations, particularly in the area of ​​disability assistance assessment. The introduction and widespread application of Retrieval Enhancement Generation (RAG), Large Language Modeling (LLM), and multimodal models (CLIP, BLIP, etc.) have promoted the structuring, transparency, and simplification of the disability assessment process. The combination of these technologies not only improves data processing efficiency but also effectively enhances the accuracy and consistency of the assessment process, driving the intelligent transformation of the field of disability assistance assessment.

[0004] Traditional assessment methods typically rely on manual evaluation, and the results are prone to discrepancies when dealing with complex or ambiguous cases. In this process, assessors' understanding and application of standards are often subject to subjective bias, leading to inconsistent results. Existing auxiliary assessment methods for work-related injury disability levels mainly rely on single-modal data (images or text) to assist in determining the disability level, often neglecting the connections between different data sources. Especially when using RAG (Retrieval-Augmented Generation) technology, a single-modal RAG system is often employed. While this single-modal RAG system can obtain certain assessment results by independently analyzing text or image data, it still has the following drawbacks:

[0005] 1) Assessments based on single-modal data may not fully consider all medical information of the injured party. For example, while images or video data can reveal the specific circumstances of the injury, they cannot fully reflect the patient's medical history, symptom descriptions, and treatment effects, leading to insufficient accuracy in the assessment results.

[0006] 2) While unimodal text data can provide more comprehensive background information, it lacks accurate capture and understanding of image features. Therefore, unimodal RAG systems are prone to erroneous assessments when assisting in the evaluation of work-related injury disability levels due to isolated data and incomplete information, especially in complex cases or judgments with ambiguous boundaries.

[0007] 3) Existing RAG technologies mostly rely on single-modal data (images or text) for assessment, failing to effectively integrate key information from different data sources. A single data source cannot comprehensively reflect the disability status of the injured party, thus affecting the accuracy of the assessment results. Summary of the Invention

[0008] To address the aforementioned issues, this disclosure proposes a method and system for assisting in the assessment of work-related injuries based on dual-pathway retrieval enhancement. It constructs a dual-pathway retrieval enhancement generation structure that combines regulatory maps and similar cases. By encoding multimodal data such as medical images, diagnostic texts, and legal clauses into a unified semantic vector space, it achieves automated auxiliary assessment of work-related injury levels, assisting in the generation of accurate and reliable disability assessment levels, and significantly improving the efficiency and accuracy of the auxiliary assessment.

[0009] According to some embodiments, the present disclosure adopts the following technical solutions:

[0010] Work injury auxiliary identification methods based on dual-pathway retrieval enhancement include:

[0011] Acquire multimodal data including injury patient movement videos, medical images, and medical record texts, and perform preprocessing.

[0012] Feature extraction is performed on the preprocessed video frames, images, and text respectively, and the semantic vectors of images, text, and video are fused into a unified semantic representation vector using a self-attention mechanism;

[0013] The semantic representation vector is input into the dual-path retrieval enhanced generation model, and the output is the regulatory map matching basis chain and the core case summary information. The semantic representation vector, the regulatory map matching basis chain and the core case summary information are input into the RAG generation model, and the output is the disability level auxiliary discrimination result.

[0014] The semantic representation vector is input into the dual-path retrieval enhancement generation model and then enters the regulatory structured knowledge graph retrieval path and the historical case semantic retrieval path, respectively. The regulatory structured knowledge graph retrieval path extracts the final representation of each node in the regulatory knowledge graph and performs cosine similarity calculation with the semantic representation vector. Based on the similarity calculation result, the regulatory node path is expanded to obtain the regulatory graph matching basis chain. The historical case semantic retrieval path calculates relevant historical cases through semantic similarity to obtain the core summary information of the cases.

[0015] According to some embodiments, the present disclosure adopts the following technical solutions:

[0016] A work injury auxiliary assessment system based on dual-pathway retrieval enhancement includes:

[0017] The data acquisition module is used to acquire multimodal data, including videos of the injured person's movements, medical images, and medical records, and to perform preprocessing.

[0018] The multimodal fusion module is used to extract features from preprocessed video frames, images, and text respectively, and to fuse the semantic vectors of images, text, and video into a unified semantic representation vector using a self-attention mechanism.

[0019] The auxiliary discrimination module is used to input the semantic representation vector into the dual-path retrieval enhanced generation model and output the regulatory map matching basis chain and the core case summary information. The semantic representation vector, the regulatory map matching basis chain and the core case summary information are input into the RAG generation model and output the disability level auxiliary discrimination result.

[0020] The semantic representation vector is input into the dual-path retrieval enhancement generation model and then enters the regulatory structured knowledge graph retrieval path and the historical case semantic retrieval path, respectively. The regulatory structured knowledge graph retrieval path extracts the final representation of each node in the regulatory knowledge graph and performs cosine similarity calculation with the semantic representation vector. Based on the similarity calculation result, the regulatory node path is expanded to obtain the regulatory graph matching basis chain. The historical case semantic retrieval path calculates relevant historical cases through semantic similarity to obtain the core summary information of the cases.

[0021] According to some embodiments, the present disclosure adopts the following technical solutions:

[0022] A work injury auxiliary assessment system based on dual-pathway retrieval enhancement includes:

[0023] The data acquisition module is used to acquire multimodal data, including videos of the injured person's movements, medical images, and medical records, and to perform preprocessing.

[0024] The multimodal fusion module is used to extract features from preprocessed video frames, images, and text respectively, and to fuse the semantic vectors of images, text, and video into a unified semantic representation vector using a self-attention mechanism.

[0025] The auxiliary discrimination module is used to input the semantic representation vector into the dual-path retrieval enhanced generation model and output the regulatory map matching basis chain and the core case summary information. The semantic representation vector, the regulatory map matching basis chain and the core case summary information are input into the RAG generation model and output the disability level auxiliary discrimination result.

[0026] The semantic representation vector is input into the dual-path retrieval enhancement generation model and then enters the regulatory structured knowledge graph retrieval path and the historical case semantic retrieval path, respectively. The regulatory structured knowledge graph retrieval path extracts the final representation of each node in the regulatory knowledge graph and performs cosine similarity calculation with the semantic representation vector. Based on the similarity calculation result, the regulatory node path is expanded to obtain the regulatory graph matching basis chain. The historical case semantic retrieval path calculates relevant historical cases through semantic similarity to obtain the core summary information of the cases.

[0027] According to some embodiments, the present disclosure adopts the following technical solutions:

[0028] A computer program product includes a computer program that, when executed by a processor, implements the work injury auxiliary identification method based on dual-path retrieval enhancement.

[0029] According to some embodiments, the present disclosure adopts the following technical solutions:

[0030] A non-transitory computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the work injury auxiliary identification method based on dual-path retrieval enhancement.

[0031] According to some embodiments, the present disclosure adopts the following technical solutions:

[0032] An electronic device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the work injury auxiliary identification method based on dual-path retrieval enhancement.

[0033] Compared with the prior art, the beneficial effects of this disclosure are as follows:

[0034] The work injury auxiliary assessment method disclosed herein, based on dual-path retrieval enhancement generation, can construct a dual-path retrieval enhancement generation model without the need for additional external equipment. The dual-path retrieval enhancement generation model is a dual-path retrieval enhancement generation structure that combines regulatory maps and similar cases. By encoding multimodal data such as medical images, diagnostic texts, and legal clauses into a unified semantic vector space, and then using dual-path hybrid retrieval and effective fusion of multimodal data for auxiliary judgment, it achieves automated auxiliary assessment of work injury level, assists in generating accurate and reliable disability assessment levels, and greatly improves the efficiency and accuracy of auxiliary assessment.

[0035] This disclosed method for work-related injury auxiliary identification based on dual-path retrieval enhancement combines relevant work-related injury regulations and standards with structured historical cases to construct a dynamic knowledge graph. It automates the construction of the knowledge graph using a large language model, employs a BIO-annotated named entity recognition model for entity extraction, and utilizes a fine-tuned BART-Relation model for relation extraction to generate structured triples. The confidence level of the triples is calculated using softmax normalization, and a confidence threshold of 0.8 is set for quality filtering. Standardized mapping of legal provision numbers establishes a unique node identification system. Secondly, a conflict detection mechanism is constructed, automatically identifying conflicts between old and new regulations by comparing version number fields, triggering a manual review process. Finally, a Diff-based incremental update algorithm is used to maintain the knowledge graph, and version tag management and obsolete clause marking ensure the timeliness of the knowledge base. This technical solution achieves efficient knowledge graph construction, automatic conflict detection, and dynamic updating, providing accurate and reliable structured knowledge support for work-related injury identification.

[0036] This disclosed method for work-related injury auxiliary identification based on dual-pathway retrieval enhancement generation employs a dual-pathway retrieval enhancement generation mechanism during the inference phase. The dual-pathway retrieval enhancement generation model includes a legal structured knowledge graph retrieval path and a historical case semantic retrieval path. These paths can retrieve highly relevant information from the legal knowledge graph and historical case database, respectively. The legal structured knowledge graph retrieval path uses a graph attention network (GAT) to enhance node semantics, dynamically modeling inter-node dependencies through learnable attention weights to generate fused node representations. Combined with cosine similarity calculation and path scoring mechanisms, it outputs the optimal legal matching path chain. The historical case semantic retrieval path uses a ColBERT model to achieve fine-grained semantic matching, accurately recalling Top-K relevant historical cases through token-level maximum similarity calculation (MaxSim) and field weighting mechanisms. This technical solution, through dual-pathway collaborative retrieval, ensures both the accuracy of legal citations and the relevance of case references, enabling the generated identification results to possess both legal rigor and practical guidance, significantly improving the professionalism and interpretability of work-related injury level determination.

[0037] This disclosed method for work-related injury auxiliary assessment based on dual-pathway retrieval enhancement utilizes a large model (GPT-4o) in the auxiliary judgment generation stage. This model enables high-quality disability level auxiliary assessment without fine-tuning. To improve inference quality, a structured prompt design strategy is adopted. A prompt input is constructed, and a pre-trained language model generates supported disability level auxiliary assessment results and explanations. A multimodal feature fusion template is constructed, structurally organizing a unified semantic representation vector with the legal basis chain and case summaries obtained from dual-pathway retrieval. Top-K cases are presented in three dimensions: "injury similarity - diagnostic matching degree - level reference value." This scheme achieves dynamic context management, multi-evidence weighting, and terminology verification through a RAG architecture, ensuring that the generated results comply with legal norms, possess medical accuracy, and support complete traceability of the judgment basis. Attached Figure Description

[0038] The accompanying drawings, which form part of this disclosure, are used to provide a further understanding of this disclosure. The illustrative embodiments of this disclosure and their descriptions are used to explain this disclosure and do not constitute an undue limitation of this disclosure.

[0039] Figure 1 This is an overall flowchart of the work injury auxiliary identification method based on dual-path retrieval enhancement according to an embodiment of the present disclosure.

[0040] Figure 2 This is an implementation architecture diagram of the work injury auxiliary identification method based on dual-path retrieval enhancement according to an embodiment of this disclosure;

[0041] Figure 3 This is a flowchart illustrating the multimodal data feature extraction and fusion process according to an embodiment of the present disclosure.

[0042] Figure 4 This is a system processing flowchart of the work injury auxiliary identification method based on dual-path retrieval enhancement according to an embodiment of the present disclosure;

[0043] Figure 5 This is a flowchart of the data processing for the regulatory structured knowledge graph retrieval pathway, as described in this embodiment of the disclosure.

[0044] Figure 6 This is a flowchart of the semantic retrieval path data processing for historical cases, as described in this embodiment of the disclosure.

[0045] Figure 7 This is a flowchart illustrating the application of a large-model-assisted generation reasoning method for work-related injury auxiliary identification based on dual-pathway retrieval enhancement in this embodiment of the present disclosure. Detailed Implementation

[0046] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.

[0047] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0048] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this disclosure. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0049] Terminology Explanation

[0050] Multimodal RAG (Retrieval Enhanced Generation): A technical architecture that combines multi-source data retrieval with generative reasoning. By integrating multimodal inputs such as text, images, and videos, and retrieving legal knowledge bases and historical cases, it generates legally sound expert conclusions.

[0051] CLIP (Contrastive Language–Image Pre-training) is a multimodal model proposed by OpenAI that can simultaneously understand images and text and map them into a shared semantic vector space (joint embedding space).

[0052] GAT is a graph neural network (GNN) model that uses an attention mechanism to assign different weights to the neighboring nodes of each node in the graph, thereby achieving finer-grained aggregation of structure-aware information.

[0053] SBERT is a structural improvement on BERT (Bidirectional Encoder Representations from Transformers), enabling it to efficiently generate fixed-length semantic vectors at the sentence level for calculating text similarity or retrieval tasks.

[0054] ColBERT is a dense semantic retrieval framework designed to perform large-scale semantic retrieval tasks efficiently and accurately. It employs a late interaction mechanism to achieve more efficient vector retrieval while preserving the contextual semantic information encoded by BERT.

[0055] LLM (Large Language Model) is a deep learning-based natural language processing (NLP) model. Its core feature is that it has billions to trillions of parameters. Through pre-training on large-scale text corpora, it has powerful language understanding, generation and reasoning capabilities.

[0056] Example 1

[0057] One embodiment of this disclosure provides a work injury auxiliary identification method based on dual-pathway retrieval enhancement, the steps of which include:

[0058] Step 1: Acquire multimodal data including the injured person's movement videos, medical images, and medical records, and perform preprocessing.

[0059] Step 2: Extract features from the preprocessed video frames, images, and text respectively, and use a self-attention mechanism to fuse the semantic vectors of images, text, and video into a unified semantic representation vector;

[0060] Step 3: Input the semantic representation vector into the dual-path retrieval enhancement generation model, and output the regulatory map matching basis chain and the core case summary information. Input the semantic representation vector, the regulatory map matching basis chain and the core case summary information into the RAG generation model, and output the disability level auxiliary discrimination result.

[0061] The semantic representation vector is input into the dual-path retrieval enhancement generation model and then enters the regulatory structured knowledge graph retrieval path and the historical case semantic retrieval path, respectively. The regulatory structured knowledge graph retrieval path extracts the final representation of each node in the regulatory knowledge graph and performs cosine similarity calculation with the semantic representation vector. Based on the similarity calculation result, the regulatory node path is expanded to obtain the regulatory graph matching basis chain. The historical case semantic retrieval path calculates relevant historical cases through semantic similarity to obtain the core summary information of the cases.

[0062] As one embodiment, the work injury auxiliary assessment method based on dual-pathway retrieval enhancement generation disclosed herein constructs a dual-pathway retrieval enhancement generation model. This model includes two processing pathways: a regulatory structured knowledge graph retrieval pathway and a historical case semantic retrieval pathway. By encoding medical images, case texts, and injured person's movement videos into a unified semantic vector space, and combining the regulatory structured knowledge graph retrieval pathway and the historical case semantic retrieval pathway into a dual-pathway RAG structure, the automated assessment of work injury levels is achieved, assisting in the generation of accurate and reliable disability grades. The specific implementation process is as follows:

[0063] Step 1: Acquire multimodal data including the injured person's movement video, medical images, and medical record text, and perform preprocessing;

[0064] Specifically, the system first acquires handwritten scanned PDFs, electronic medical records (XML / HL7) or accident reports (Word-PDF) through a multimodal data acquisition module. Then, it uses PaddleOCR and a medical terminology error correction model to perform text recognition, achieving a recognition rate of 97.2%. The system also acquires medical images (PACS system DICOM files with standardized tags extracted by pydicom) and video data (mp4 / avi format at 25fps, with keyframes extracted based on a ΔPose ≥ 15% motion threshold).

[0065] Furthermore, the preprocessing of multimodal data includes:

[0066] 1) Text preprocessing: The text preprocessing and entity extraction module uses MedSpaCy to perform clinical syntactic analysis and segmentation to identify term phrases (such as "left knee flexion disorder"), and uses UMLS (Unified Medical Language System) and ICD-10-CM to standardize and map entities such as anatomical structures, injury descriptions, functional status, disability grades, and assessment criteria.

[0067] 2) Image preprocessing: Image preprocessing includes grayscale conversion, size normalization, and noise reduction.

[0068] 3) Video preprocessing: First, the original video is sampled and preprocessed frame by frame, including frame rate unification, posture smoothing filtering and outlier removal, to obtain stable human posture trajectory video frames.

[0069] Step 2: Extract features from the preprocessed video frames, images, and text respectively, and use a self-attention mechanism to fuse the semantic vectors of images, text, and video into a unified semantic representation vector;

[0070] Specifically, step 21: First, feature extraction is performed on the preprocessed video frames, images, and text respectively. The specific process is as follows:

[0071] Step 211: The image is processed using a CLIP (ViT-B / 32) based feature extraction model, and Grad-CAM is used to highlight the damaged area. A fracture recognition binary classifier, finely tuned on the RSNA 2018 dataset, is also introduced to improve the bone injury recognition rate, including:

[0072] The image was processed using the CLIP-ViT (ViT-B / 32) model to extract global semantic features and perform binary classification of fractures. When the classification result was a fracture and the confidence score exceeded a set threshold (set to 0.85), a heatmap of the injury area was generated using Grad-CAM, and the corresponding bounding box was extracted. Finally, the location, classification label, confidence score, and semantic embedding vector of the injury area were uniformly encapsulated into a structured output, for example: {"region": "right knee joint", "bbox": [x1, y1, x2, y2], "label": "fracture", "confidence": 0.91, "embedding":[0.132, 0.432, ...]}, which yielded the image semantic vector.

[0073] Step 212: The text is used to extract features using the BERT model to generate a text semantic vector, including: the text is generated by a medical fine-tuned version of the BERT model to generate semantic embeddings, i.e., text semantic vectors.

[0074] Step 213: For video data, after frame-by-frame sampling and preprocessing, the human pose trajectory is obtained. Based on the human pose trajectory, a dynamic feature sequence of key points changing over time is constructed, and a two-layer LSTM network is used to encode the dynamic feature sequence to output the video semantic vector, including:

[0075] First, the original video is sampled and preprocessed frame by frame to obtain a stable human posture trajectory. Then, the MediaPipe framework is used to extract key points of the human skeleton and generate a 33-dimensional temporal coordinate point sequence (each frame contains the three-dimensional coordinate information of 33 key points).

[0076] Furthermore, a time series features (TSF) sequence of key points is constructed, including velocity, acceleration, displacement curves, and joint angle changes.

[0077] Subsequently, a two-layer LSTM network was used to model and classify the TSF, identifying the restricted patterns of key actions (such as "inability to fully bend the knee" and "dragging while walking"), which are the video semantic vectors.

[0078] As one embodiment, the two-layer LSTM network model is trained using a labeled motion anomaly dataset, and the output consists of several movement disorder categories and their confidence levels. Simultaneously, the rate of change of angle Δθ / Δt for key joints such as the knee and ankle joints is calculated. This is compared frame-by-frame with a standard movement template (derived from a medical rehabilitation movement library). If the angle change of a certain joint is below a set threshold (e.g., 15° / s) for several consecutive frames, it is determined that the joint has a functional disorder.

[0079] Step 22: The semantic vectors of images, text, and videos are fused into a unified semantic representation vector using a self-attention mechanism. This is achieved through a customized CrossModalEncoder module with a Transformer variant architecture. Its core customized features include:

[0080] 1) Multimodal differentiation is achieved through a modal identifier embedding layer;

[0081] 2) Frame-level learnable positional coding is used to accurately model the timing of MediaPipe key points in video data;

[0082] 3) Attention mask matrix with medical prior constraints;

[0083] 4) Output a fixed 768-dimensional semantic space to adapt to subsequent SBERT / GAT processing, thereby solving the technical challenges of multimodal semantic alignment and medical decision reliability in work injury identification.

[0084] By employing the self-attention mechanism in the Transformer architecture, the semantic vectors of images, text, and videos are fused into a unified semantic vector representation through the aforementioned customized CrossModalEncoder module. Specifically, the semantic vectors of images, text, and videos are input into the CrossModalEncoder module, and each vector is uniformly projected onto the same dimensional space (e.g., 768 dimensions) for fusion.

[0085] During the fusion process, modality type embedding and positional encoding are introduced before each modal input to indicate the information source and maintain temporal order dependencies. Then, the three modal vectors are concatenated into an input sequence and fed into the CrossModalEncoder module to perform multi-head self-attention computation, capturing significant interactions and contextual dependencies between modalities. The fused output is a unified semantic representation vector V. input .

[0086] Step 3: Input the semantic representation vector into the dual-path retrieval enhancement generation model, and output the regulatory graph matching basis chain and the core case summary information;

[0087] Specifically, the dual-path retrieval enhancement generation model includes a regulatory structured knowledge graph retrieval path and a historical case semantic retrieval path. The regulatory structured knowledge graph retrieval path includes a graph attention network and a bidirectional sentence embedding model, while the historical case semantic retrieval path includes a query encoder and a document encoder.

[0088] Further, step 31: After the semantic representation vector is input into the dual-path retrieval enhancement generative model, it enters the regulatory structured knowledge graph retrieval path and the historical case semantic retrieval path respectively. The processing in the regulatory structured knowledge graph retrieval path includes:

[0089] Using semantic representation vector V input As input, the semantics of the knowledge graph nodes are first initialized. The text description of each entity node or edge in the knowledge graph is input into the fine-tuned SBERT model to obtain the context embedding vector of each token. The initial node semantic vector is then generated through Mean Pooling. This serves as the input to the GAT graph neural network. Fine-tuning SBERT aims to enhance its understanding of semantics in specific domains (such as legal provisions and case law), thereby improving the quality of semantic similarity representations between nodes and laying the foundation for subsequent graph neural network processing.

[0090] As one example, the knowledge graph construction process includes:

[0091] (1) A knowledge graph structure is constructed using a large language model (such as GPT-4o). The full text of the disability assessment standard and related case documents are fine-tuned to achieve automatic identification and relationship extraction of key medical and legal entities. Entity identification adopts the BIO annotation mode to construct a named entity recognition (NER) model. The relationship extraction part is based on the fine-tuned BART-Relation structure, and the output is a structured triple (e.g., ["right knee", "loss of function", "level nine disability"]). The extraction results are normalized by the softmax function, the confidence of the triple is calculated, and 0.8 is set as the confidence threshold. For the extracted legal article number, the system further completes the standardization mapping operation and binds it to the graph node (e.g., "GB / T 16180-2022 Article 4.2.3" → node_8745) to realize the automatic construction of the knowledge graph and the allocation of unique node identifiers.

[0092] (2) Conflict detection and handling is carried out by comparing the version number field (such as valid_from, valid_to) in the legal clause knowledge graph to detect the conflict between the extracted knowledge. For example, it can identify the situation where the same injury is classified differently under the old and new versions of the legal clauses, trigger the manual review mechanism and push it to the queue to be processed.

[0093] (3) Obtain a knowledge graph, which consists of entity nodes (E), relation edges (R), and legal clause nodes (L). Entity nodes represent standardized medical terms (such as anatomical locations encoded in ICD-10) and are stored as JSON structures containing feature vectors. Relation edges define medical associations between entities (such as "causing") and carry confidence weights and timeliness markers. Legal clause nodes (L) map legal provisions (such as GB / T standards) and are dynamically maintained through version tags and repeal attributes. A diff-based incremental update algorithm is used to automatically add new nodes and relationships. After updating, the nodes are tagged (such as "v2025.1") and the change record is retained. For repealed clause nodes, the attribute "deprecated": true is set and a link to the alternative clause is displayed.

[0094] Furthermore, As input to the GAT graph neural network, the GAT graph neural network dynamically models the contextual dependencies of each node by assigning learnable attention weights to its neighboring nodes. Its core mechanism is that for each node in the graph... For its adjacent nodes The connections between them are scored using an attention-based scoring function. Attention weights are obtained through Softmax normalization. The details are as follows:

[0095]

[0096] in, For attention weights, For learnable transformation matrices, It is a non-linear activation function. for The set of neighbors of a node This is the transpose of the attention vector. The initial features of neighboring nodes are given by [the following]. The weighted neighbor node representation yields the structure enhancement vector:

[0097]

[0098] Finally, the initial node semantic vector is concatenated. With structural enhancement vector Generate fused node representation , fusion node representation Used for subsequent path retrieval and semantic matching.

[0099] Furthermore, path retrieval and semantic matching are performed using fused node representations, including:

[0100] First, the semantic representation vector V is...input Integration of nodes with the legal knowledge graph Cosine similarity calculations are performed to initially screen relevant nodes.

[0101] Subsequently, path expansion is performed on these nodes in the graph, and the candidate paths are scored and ranked based on the semantic relation weights of the edges, specifically as follows:

[0102]

[0103] For each triple in each path : It is a node The final representation (obtained via SBERT+GAT); This indicates the semantic relevance of the node to the query; Representing an edge Relationship weights; It is the product of two factors, measuring the importance of the edge (or path segment). Then, for all edges in the entire path... The summation is the total score for the path.

[0104] Finally, the legal path chain with the highest score is selected as the legal graph matching basis chain, providing interpretable and context-relevant legal basis support for the generative reasoning module.

[0105] Step 32: The historical case semantic retrieval pathway constructs a case library based on a three-element structure of "injury description + diagnosis conclusion + disability level," and uniformly maps it to a high-dimensional dense vector space. A dual-encoding retrieval mechanism is implemented using the ColBERT (Contextualized Late Interaction over BERT) model. The specific implementation process is as follows:

[0106] First, the current case query progress (including injury summary, diagnosis conclusion, preliminary assessment level, etc.) is input into the query encoder, and the candidate historical cases (including complete disability description, diagnosis, clause citation, etc.) are input into the document encoder. Both share the pre-trained BERT model parameters. ColBERT does not directly calculate the global vector similarity of the entire sentence, but instead encodes the query and document (candidate historical cases) into word-level vector sequences by decomposing the token-level vector sequence, thus obtaining: The system processes each query vector With document vector set Calculate the dot product similarity, take the maximum score, and finally all The sum of the maximum scores is used as the overall matching score:

[0107]

[0108] in, The BERT output token vector set for the query; A collection of token vectors for a document; Represents the dot product, measuring the first product. i The query token and the first j Similarity of document tokens.

[0109] Furthermore, to ensure fairness and consistency in dimensionality distribution during matching, L2 normalization is applied to all candidate case vectors D, using this formula... For each token vector, the modulus is standardized. Simultaneously, weighting factors for different fields (such as injury severity, diagnosis, and grade) can be introduced, where the weighting factor ratio is... , ( The formula is (1.0, 1.5, 0.8). This controls the relative contribution of various information types to the final matching score.

[0110] Finally, the system retrieves the Top-K most relevant historical cases from the candidate case set based on the matching score, and obtains the core summary information of the cases. The specific Top-K high-return cases are as follows:

[0111]

[0112] Top-k high-return cases serve as contextual supplements and judgment support for generative reasoning modules (such as GPT-4o), effectively improving the professionalism, rationality, and interpretability of work injury level determination.

[0113] Step 4: Input the semantic representation vector, the regulatory graph matching basis chain, and the core case summary information into the RAG generation model, and output the disability level auxiliary discrimination result;

[0114] Specifically, after obtaining the regulatory map matching chain and core case summary information, the multimodal fusion representation and the initial input are merged at a certain ratio and used as input to the RAG generation model. The RAG generation model calls the large model (GPT-4o), which can perform high-quality disability level auxiliary determination without fine-tuning. To improve the reasoning quality, a structured prompt design strategy is adopted to construct the prompt input. The model outputs the disability level auxiliary prediction, the explanation reason, and the applicable regulatory clauses.

[0115] Example 2

[0116] One embodiment of this disclosure provides a work injury auxiliary identification system based on dual-pathway retrieval enhancement, comprising:

[0117] The data acquisition module is used to acquire multimodal data, including videos of the injured person's movements, medical images, and medical records, and to perform preprocessing.

[0118] The multimodal fusion module is used to extract features from preprocessed video frames, images, and text respectively, and to fuse the semantic vectors of images, text, and video into a unified semantic representation vector using a self-attention mechanism.

[0119] The auxiliary discrimination module is used to input the semantic representation vector into the dual-path retrieval enhanced generation model and output the regulatory map matching basis chain and the core case summary information. The semantic representation vector, the regulatory map matching basis chain and the core case summary information are input into the RAG generation model and output the disability level auxiliary discrimination result.

[0120] The semantic representation vector is input into the dual-path retrieval enhancement generation model and then enters the regulatory structured knowledge graph retrieval path and the historical case semantic retrieval path, respectively. The regulatory structured knowledge graph retrieval path extracts the final representation of each node in the regulatory knowledge graph and performs cosine similarity calculation with the semantic representation vector. Based on the similarity calculation result, the regulatory node path is expanded to obtain the regulatory graph matching basis chain. The historical case semantic retrieval path calculates relevant historical cases through semantic similarity to obtain the core summary information of the cases.

[0121] Example 3

[0122] One embodiment of this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the work injury auxiliary identification method based on dual-path retrieval enhancement.

[0123] Example 4

[0124] One embodiment of this disclosure provides a non-transitory computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, they implement the work injury auxiliary identification method based on dual-path retrieval enhancement.

[0125] Example 5

[0126] One embodiment of this disclosure provides an electronic device, including a processor, a memory, and a computer program; wherein the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the work injury auxiliary identification method based on dual-path retrieval enhancement.

[0127] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0128] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0129] While the specific embodiments of this disclosure have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of this disclosure are still within the scope of protection of this disclosure.

Claims

1. A work-related injury auxiliary identification method based on dual-path retrieval enhancement, characterized in that, include: Acquire multimodal data including injury patient movement videos, medical images, and medical record texts, and perform preprocessing. Feature extraction is performed on the preprocessed video frames, images, and text respectively, and the semantic vectors of images, text, and video are fused into a unified semantic representation vector using a self-attention mechanism; The semantic representation vector is input into the dual-path retrieval enhanced generation model, and the output is the regulatory map matching basis chain and the core case summary information. The semantic representation vector, the regulatory map matching basis chain and the core case summary information are input into the RAG generation model, and the output is the disability level auxiliary discrimination result. The semantic representation vector is input into the dual-path retrieval enhancement generation model and then enters the legal structured knowledge graph retrieval path and the historical case semantic retrieval path, respectively. The legal structured knowledge graph retrieval path extracts the final representation of each node in the legal knowledge graph and performs cosine similarity calculation with the semantic representation vector. Based on the similarity calculation result, the legal node path is expanded to obtain the legal graph matching basis chain. The historical case semantic retrieval path calculates relevant historical cases through semantic similarity to obtain the core summary information of the cases. The legal structured knowledge graph retrieval pathway includes a graph attention network and a bidirectional sentence embedding model. First, the text description of each entity node or edge in the knowledge graph is input into the bidirectional sentence embedding model to obtain the context embedding vector for each token. An initial node semantic vector is generated through average pooling. This initial node semantic vector is then input into the graph attention network. By assigning learnable attention weights to the neighboring nodes of each node in the graph, contextual dependencies are dynamically modeled, and the neighbor node representations are weighted to obtain a structure enhancement vector. The initial node semantic vector and the structure enhancement vector are concatenated to generate a fused node representation. Cosine similarity is calculated between the semantic representation vector and the fused node representation to initially screen relevant nodes. Subsequently, path expansion is performed based on relevant nodes in the knowledge graph, and candidate paths are scored and ranked according to the semantic relationship weights of the edges. The legal path chain with the highest score is selected as the legal graph matching basis chain.

2. The work-related injury auxiliary identification method based on dual-path retrieval enhancement as described in claim 1, characterized in that, The feature extraction process for preprocessed video frames, images, and text includes: extracting semantic vectors for each of the three modalities (images, text, and video frames) using pre-trained models; extracting image semantic vectors using the CLIP-ViT model; extracting text semantic vectors using the BERT model; and for video data, sampling and preprocessing the video data frame by frame to obtain human pose trajectories, constructing dynamic feature sequences of key points changing over time based on the human pose trajectories, encoding the dynamic feature sequences using a two-layer LSTM network, and outputting video semantic vectors.

3. The work-related injury auxiliary identification method based on dual-pathway retrieval enhancement as described in claim 1, characterized in that, The method of using a self-attention mechanism to fuse semantic vectors of images, text, and videos into a unified semantic representation vector includes: during the fusion process, a modality identifier embedding and position encoding are introduced before each modality input to indicate the source of information and maintain temporal order dependency. Then, the semantic vectors of images, text, and videos are concatenated into an input sequence and sent to the self-attention mechanism module to perform multi-head self-attention calculation, capturing significant interaction relationships and contextual dependencies between modalities, and outputting the fused semantic representation vector.

4. The work-related injury auxiliary identification method based on dual-pathway retrieval enhancement as described in claim 1, characterized in that, The processing steps of the historical case semantic retrieval pathway include: first, inputting the current case query progress into the query encoder and the candidate historical cases into the document encoder; encoding the query progress and candidate historical cases into sub-word level vector sequences to obtain query vectors and document vectors; calculating the dot product similarity between each query vector and the document vector set; taking the maximum score; finally, summing all the maximum scores as the overall matching score; and introducing different field weighting factors to control the relative contribution of various information to the final matching score; and recalling the Top-K most relevant historical cases from the candidate historical cases as the core case summary information based on the matching score.

5. The work-related injury auxiliary identification method based on dual-pathway retrieval enhancement as described in claim 1, characterized in that, A knowledge graph structure is constructed using a large language model. The full text of the disability assessment standard and related case documents are fine-tuned. Entity recognition adopts the BIO annotation mode to build a named entity recognition model. Relation extraction is based on the fine-tuned BART-Relation structure, and structured triples are output. The extraction results are normalized by the softmax function, and the confidence of the triples is calculated. For the extracted legal article numbers, a standardized mapping operation is further performed to bind them to the graph nodes, realizing the automatic construction of the knowledge graph and the allocation of unique node identifiers.

6. A work injury auxiliary identification system based on dual-path retrieval enhancement, specifically implementing the work injury auxiliary identification method based on dual-path retrieval enhancement as described in any one of claims 1-5, characterized in that, include: The data acquisition module is used to acquire multimodal data, including videos of the injured person's movements, medical images, and medical records, and to perform preprocessing. The multimodal fusion module is used to extract features from preprocessed video frames, images, and text respectively, and to fuse the semantic vectors of images, text, and video into a unified semantic representation vector using a self-attention mechanism. The auxiliary discrimination module is used to input the semantic representation vector into the dual-path retrieval enhanced generation model and output the regulatory map matching basis chain and the core case summary information. The semantic representation vector, the regulatory map matching basis chain and the core case summary information are input into the RAG generation model and output the disability level auxiliary discrimination result. The semantic representation vector is input into the dual-path retrieval enhancement generation model and then enters the regulatory structured knowledge graph retrieval path and the historical case semantic retrieval path, respectively. The regulatory structured knowledge graph retrieval path extracts the final representation of each node in the regulatory knowledge graph and performs cosine similarity calculation with the semantic representation vector. Based on the similarity calculation result, the regulatory node path is expanded to obtain the regulatory graph matching basis chain. The historical case semantic retrieval path calculates relevant historical cases through semantic similarity to obtain the core summary information of the cases.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the work injury auxiliary identification method based on dual-path retrieval enhancement as described in any one of claims 1-5.

8. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement the work injury auxiliary identification method based on dual-path retrieval enhancement as described in any one of claims 1-5.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to perform the work injury auxiliary identification method based on dual-path retrieval enhancement as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Intelligent trademark law question and answer method for enhancing big language model reasoning through knowledge graph

    CN117807202A

  • Conference record data searching method and system based on AI

    CN120448597A