An Automatic Labeling Method and System for Injury Tags

By processing the forensic injury test videos and text reports and automatically labeling the injury labels, the problem of low marking efficiency in the existing technology is solved, and efficient and accurate injury labeling is achieved.

CN119339292BActive Publication Date: 2025-06-17SHANGHAI CITY PUDONG NEW DISTRICT ZHOUPU HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411442443.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2025-06-17
Estimated Expiration
2044-10-16

AI Technical Summary

Technical Problem

In the prior art, the labeling efficiency of forensic injury test reports is low and cannot meet actual needs.

Method used

By obtaining forensic injury test video stream data and text injury test report, video stream data is used for image segmentation and injury feature extraction, combined with large language models to understand the text report, and matching analysis and automatic labeling of injury labels are carried out.

Benefits of technology

It realizes automatic labeling of injury labels, improves the efficiency of labeling forensic injury test results, and ensures the accuracy of labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339292B_ABST
    Figure CN119339292B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of forensic identification. A method and system for automatically annotating injury labels are provided. The method includes: obtaining first video stream data of a forensic doctor examining an injured person, and splitting the first video stream data into multiple segments of images according to the examination area; extracting injury characteristics from each segment of image to obtain corresponding first injury labels and corresponding injury positions; receiving the written injury report generated by the forensic doctor, using a large language model to perform semantic understanding on the written injury report, and extracting several second injury labels; performing matching analysis on the second injury labels and the first injury labels, and if a matching result is obtained, determining the corresponding second injury label as the target injury label; in the target image, annotating each target injury label at the injury position corresponding to the matching first injury label. The present invention realizes automatic and accurate annotation of injury labels, greatly improving the efficiency of annotating the results of forensic injury identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of forensic identification, and more particularly, to an automatic injury label annotation method and system. Background Art

[0002] Forensic injury examination refers to the activity in which a forensic doctor conducts an injury appraisal on an injured individual according to certain standards and procedures. This activity is very important for determining the nature, degree, and consequences of the injury and is commonly used in criminal cases, civil disputes, industrial accident cases, etc. When issuing an injury report, in order to enhance the intuitiveness of the report, in addition to describing it in words, it is also necessary for a forensic doctor or other staff to label the injury results one by one in the image of the injured person. Obviously, the manual annotation method is relatively inefficient and cannot meet the actual needs. Summary of the Invention

[0003] To address this, the present invention provides an automatic injury label annotation method, system, electronic device, computer storage medium, and computer program product to solve the above technical problems.

[0004] The present invention discloses an automatic injury label annotation method, which includes the following steps:

[0005] Obtain the first video stream data of the forensic doctor's injury examination work on the injured person, and divide the first video stream data into multiple segments of images according to the injury area;

[0006] Extract injury characteristics from each segment of the image to obtain the corresponding first injury label and the corresponding injury location;

[0007] Receive the written injury report generated by the forensic doctor, use a large language model to perform semantic understanding on the written injury report, and extract a number of second injury labels;

[0008] Perform a matching analysis on the second injury label and the first injury label. If a matching result is obtained, determine the corresponding second injury label as the target injury label;

[0009] In the target image, label each of the target injury labels at the injury location corresponding to the matching first injury label; wherein, the target image is any frame of high-definition image in the first video stream data.

[0010] In some embodiments, the target image is determined by the following method:

[0011] Obtain the second video stream data including the entire body area of the injured person, and extract the external characteristics and behavioral characteristics of the injured person according to the second video stream data;

[0012] Predict the injury area distribution information of the injured person based on the external characteristics and the behavioral characteristics, and determine the category attribute of the target image according to the injury area distribution information, where the category attribute includes a full-body image and a non-full-body image;

[0013] If the determined category attribute of the target image is a full-body image, then intercept a high-definition image frame containing the full-body image of the injured person in the second video stream data, and intercept several high-definition image frames corresponding to the second injury label in the first video stream data;

[0014] If the determined category attribute of the target image is a non-full-body image, then intercept several high-definition image frames corresponding to the second injury label in the first video stream data.

[0015] In some embodiments, the determining the category attribute of the target image according to the injury area distribution information includes:

[0016] Determine the two injury points with the farthest distance according to each injury point in the injury area distribution information, and draw a connected area on the body image of the injured person based on these two injury points;

[0017] Calculate the ratio of the area of the connected area to the area of the body image of the injured person. If the ratio is greater than the threshold, determine that the category attribute of the target image is a full-body image; otherwise, determine that the category attribute of the target image is a non-full-body image.

[0018] In some embodiments, the splitting the first video stream data into multiple segments of images according to the injury examination area includes:

[0019] Identify the injury examination actions of the forensic doctor and the body area of the injured person in the first video stream data, and determine the injury examination area according to the injury examination actions and the body area of the injured person;

[0020] When there is a change more than a preset distance outside the injury examination area, determine that there is a switch of the injury examination area, and determine the moment of a preset duration before the switching moment as the splitting node;

[0021] Split the first video stream data into multiple segments of images according to each splitting node.

[0022] In some embodiments, use a large language model to perform semantic understanding on the written injury report, and extract several second injury labels, including:

[0023] Obtain an injury report template from the database, and determine the injury record area according to the injury report template;

[0024] Identify the injury record area in the written injury report, use a large language model to semantically understand the text content in the injury record area, and extract a number of the second injury labels.

[0025] In some embodiments, when performing a matching analysis between the second injury label and the first injury label, if no matching result is obtained, then:

[0026] Obtain the injury location corresponding to the second injury label, query whether there is a transmissive medical image stored in the database corresponding to the injury location. If so, determine a third injury label corresponding to the injury location according to the transmissive medical image;

[0027] If the second injury label matches the third injury label, determine the second injury label as the target injury label;

[0028] If the database does not store a transmissive medical image corresponding to the injury location or the second injury label does not match the third injury label, output a manual confirmation prompt message.

[0029] The present invention also discloses an automatic injury label annotation system, which includes an acquisition module, an injury extraction module, a report analysis module, a comparison analysis module, and an injury annotation module;

[0030] The acquisition module is used to acquire the first video stream data of the forensic doctor's injury examination work on the injured person, and divide the first video stream data into multiple segments of images according to the examination area;

[0031] The injury extraction module is used to extract injury features from each segment of the image, and obtain the corresponding first injury label and the corresponding injury location;

[0032] The report analysis module is used to receive the written injury report generated by the forensic doctor, use a large language model to semantically understand the written injury report, and extract a number of second injury labels;

[0033] The comparison analysis module is used to perform a matching analysis between the second injury label and the first injury label. If a matching result is obtained, determine the corresponding second injury label as the target injury label;

[0034] The injury annotation module is used to annotate each of the target injury labels at the injury location corresponding to the matching first injury label in the target image; wherein, the target image is any frame of high-definition image in the first video stream data.

[0035] The present invention also discloses an electronic device, comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the method described in any one of the preceding paragraphs.

[0036] The present invention also discloses a computer storage medium storing a computer program, which is executed by a processor to implement the method described in any one of the preceding paragraphs.

[0037] The present invention also discloses a computer program product, which implements the method described in any one of the preceding paragraphs when running on a terminal.

[0038] By processing both the forensic injury video and the forensic injury written report, and performing matching analysis on the obtained injury labels respectively, the present invention verifies the confidence level of the automatically generated injury labels through the mutual verification of the injury labels, and then realizes the automatic annotation of the injury labels, which can greatly improve the efficiency of the annotation of forensic injury identification results on the basis of ensuring the accuracy of injury annotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 is a schematic flowchart of a method for automatically annotating injury labels disclosed in an embodiment of the present invention.

[0041] Figure 2 is a schematic structural diagram of a system for automatically annotating injury labels disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The following specific embodiments illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0043] In addition, the technical features involved in different implementation manners of the present application described below can be combined with each other as long as they do not conflict with each other.

[0044] Such asFigure 1 As shown in Figure 1 , an embodiment of the present invention discloses an automatic annotation method for injury labels. The method includes the following steps:

[0045] Obtain the first video stream data of the forensic doctor's injury examination of the injured person, and segment the first video stream data into multiple segments of images according to the injury examination area;

[0046] Extract injury characteristics from each segment of the image to obtain the corresponding first injury label and the corresponding injury location;

[0047] Receive the written injury examination report generated by the forensic doctor, use a large language model to perform semantic understanding on the written injury examination report, and extract a number of second injury labels;

[0048] Perform a matching analysis on the second injury label and the first injury label. If a matching result is obtained, determine the corresponding second injury label as the target injury label;

[0049] In the target image, label each target injury label at the injury location corresponding to the matching first injury label; where the target image is any frame of high-definition image in the first video stream data.

[0050] During the process of the forensic doctor's injury examination of the injured person, a high-definition camera is used to record the video data of the forensic doctor's injury examination process. According to the different injury examination areas of the forensic doctor (each injury examination area corresponds to one injury), the entire video data can be segmented into multiple small segments, and each small segment corresponds to the injury examination process of one area. Using image recognition technology, injury characteristics can be extracted from each small segment of the image, and accordingly, the corresponding first injury label and the injury location corresponding to the first injury label can be determined. Then, after the forensic doctor completes the injury examination operation, a written injury examination report will be manually generated, and a large language model is used to perform semantic understanding on the written injury examination report, and a number of second injury labels corresponding to the injury examination results determined by the forensic doctor can be extracted. Finally, perform a matching analysis on each second injury label and each first injury label. If there is a matching result, it means that the second injury label analyzed semantically is associated with a certain first injury label extracted from the image, that is, it means that the confidence level of the second injury label is high enough. At this time, label it at the corresponding position in the target image, and this position is the injury location corresponding to the matching first injury label. Injury labels include calf fracture, forehead wound 3 cm, back burn area about 3 cm², etc.

[0051] Therefore, the present invention processes both the forensic injury identification video and the forensic injury identification text report, performs matching analysis on the respectively obtained injury labels, thereby confirming the confidence of the automatically generated injury labels through mutual verification of the injury labels, and then realizes automatic labeling of the injury labels. This can greatly improve the efficiency of labeling of forensic injury identification results while ensuring the accuracy of injury labeling.

[0052] Among them, the large language model can be any one of the BERT model, T5 (Text-to-Text Transfer Transformer) model, PaLM (Pathways Language Model) model, and LLaMA (Large Language Model Meta AI) model, and the present invention does not make any limitation. In addition, the large language model can generally obtain a proprietary model suitable for specific problem recognition, classification, prediction, etc. by fine-tuning the training data set of a specific scenario. The construction of its training data set and the fine-tuning training method are mature existing technologies, and the present invention will not repeat them here.

[0053] In some embodiments, the target image is determined by:

[0054] Acquire second video stream data containing the entire body area of ​​the injured person, and extract the external features and behavioral features of the injured person according to the second video stream data;

[0055] Predicting the injury area distribution information of the injured person according to the external features and the behavioral features, and determining the category attribute of the target image according to the injury area distribution information, wherein the category attribute includes a full-body image and a non-full-body image;

[0056] If the determined category attribute of the target image is a full-body image, a high-definition image frame containing a full-body image of the injured person is intercepted from the second video stream data, and a plurality of high-definition image frames corresponding to the second injury label are intercepted from the first video stream data;

[0057] If the determined category attribute of the target image is a non-full body image, a number of high-definition image frames corresponding to the second injury label are captured from the first video stream data.

[0058] In this embodiment, the target image is a high-definition image in the captured video stream data, which contains an image of the injured person and needs to be labeled with an injury label. The labeled target image will be used as part of the injury assessment report. Therefore, the target image should correspond to the injury distribution of the injured person.

[0059] Specifically, before a forensic doctor conducts an injury examination, second video stream data covering all body areas of the injured person is obtained first. From this, the external characteristics and behavioral characteristics of the injured person can be extracted. The external characteristics refer to wounds, bruises, blood, dressings, fixed splints (fracture conditions), etc. that can directly represent the injuries of the injured person. The behavioral characteristics include walking postures, hand movements (such as supporting a fractured right hand with the left hand), and the helping actions of other people towards the injured person, etc., which can indirectly represent the injuries of the injured person. By analyzing the external characteristics and behavioral characteristics of the injured person, it is possible to generally know where the injuries are on the injured person and the distribution of the injuries.

[0060] The external characteristics and behavioral characteristics can be processed through a preset injury prediction model. The injury prediction model can thereby predict the possible injuries and their position distributions on the injured person, and then output the injury area distribution information. The injury prediction model preferably uses a Transformer with a self-attention mechanism for construction and is fully pre-trained using a corresponding dataset. The dataset contains multiple groups of data. Each group of data includes a video data segment of the injured person, the external characteristics and behavioral characteristics extracted from this video data, and the injury labels marked at the corresponding positions on the body of the injured person in this video data.

[0061] Based on the injury distribution information of the injured person, the category attribute of the target image can be determined, namely, a full-body image or a non-full-body image. The meaning of the full-body image category is that the target image should contain at least one full-body image of the injured person. The meaning of the non-full-body image category is that the target image does not need to contain a full-body image of the injured person and only contains local images of each injury area of the injured person. For the full-body image category, a high-definition image frame containing the full-body image of the injured person needs to be intercepted from the second video stream data because it is easier to have image frames with only the full-body image of the injured person in the second video stream data, while the presence of the forensic doctor in the first video stream data makes it difficult to obtain image frames with only the full-body image of the injured person. For the non-full-body image category, only multiple high-definition image frames need to be intercepted from the first video stream data, and these high-definition image frames correspond to the previously determined second injury labels. Similarly, for the full-body image category, multiple high-definition image frames also need to be intercepted from the first video stream data, so that full-body images and local images can be added to the injury examination report to improve the injury display effect of the injury examination report.

[0062] It should be noted that the object of the forensic doctor's injury examination may also be a non-living object. In this case, only its external characteristics can be obtained, that is, the injury area distribution information of the non-living object is directly determined only based on the external characteristics.

[0063] In some embodiments, the determining the category attribute of the target image according to the injury area distribution information includes:

[0064] Based on each injury point in the injury area distribution information, determine the two injury points with the farthest distance, and draw a connected area on the body image of the injured person based on these two injury points;

[0065] Calculate the ratio of the area of the connected area to the area of the body image of the injured person. If this ratio is greater than the threshold, determine that the category attribute of the target image is a full-body image; otherwise, determine that the category attribute of the target image is a non-full-body image.

[0066] In this embodiment, the injury area distribution information includes each injury point and its distribution point in the body image of the injured person. Determine the two injury points with the farthest distance from them, for example, determine them according to the orientation from bottom to top or from top to bottom. Draw a connected area on the body image of the injured person based on these two injury points with the farthest distance, and this connected area is the injury distribution area. Then, calculate the ratio of the area of the injury distribution area to the area of the body image of the injured person. The larger this ratio is, the more injuries the injured person has and the more dispersed they are distributed at multiple points on the body; on the contrary, it means that the injuries of the injured person are less and more concentrated at certain points on the body. When this ratio is greater than the threshold, it is determined that all or main injury labels need to be marked in the full-body image, that is, a general view of the injuries needs to be provided. At this time, determine that the category attribute of the target image is a full-body image; otherwise, it is determined that there is no need to provide a general view of the injuries, and only the injury map of the specific area needs to be provided. At this time, determine that the category attribute of the target image is a non-full-body image.

[0067] In some embodiments, the splitting of the first video stream data into multiple segments according to the injury inspection area includes:

[0068] Identify the injury inspection actions of the forensic doctor and the body area of the injured person in the first video stream data, and determine the injury inspection area according to the injury inspection actions and the body area of the injured person;

[0069] When there is a change more than a preset distance outside the injury inspection area, determine that there is a switch in the injury inspection area, and determine the moment of a preset duration before this switching moment as the splitting node;

[0070] Split the first video stream data into multiple segments according to each splitting node.

[0071] In this embodiment, when the forensic doctor examines the injury, he can use image recognition technology to extract the forensic doctor's injury examination action (mainly hand action) and the image of the injured person's body area in the image frame, dynamically track the forensic doctor's injury examination action and the image of the injured person's body area, and analyze the relative position relationship between the area corresponding to the injury examination action (located on the image of the injured person's body area) and the injured person's body area, thereby determining whether the injury examination area has changed beyond the preset distance, thereby determining whether the forensic doctor has switched to the next point of other injuries. The moment of the preset time before the above-identified switching moment is determined as a segmentation node, and the preset time is, for example, 2s. According to these segmentation nodes, the first video stream data can be segmented into multiple image segments.

[0072] The preset distance can also be determined according to the type of body region where the injury inspection area is located. The human body is roughly divided into the head region, the internal organs region, and the limbs region from top to bottom. Since there are many organs in the head, more detailed distinctions are required during injury inspection, such as eye injuries, ear injuries, back of the head injuries, neck injuries, etc., and these organs are distributed at a close distance, so the preset distance of the head region is set to the minimum; while there are fewer types of organs in the limbs region, so the preset distance of the limbs region is set to the maximum; in addition, the number of organs in the internal organs region is between the head and the limbs, so the preset distance of the internal organs region is set to a medium value.

[0073] It should be noted that injury assessment actions refer to actions related to injury assessment, such as picking, tweezing, measuring, etc., applied to the injured person's body, and do not involve non-direct injury assessment actions such as checking records after the injury, looking for professional equipment, etc. This can reduce the probability of misjudgment of switching injury assessment areas.

[0074] In some embodiments, a large language model is used to perform semantic understanding on the text injury assessment report to extract a number of second injury labels, including:

[0075] Obtaining an injury report template from a database, and determining an injury record area according to the injury report template;

[0076] An injury record area in the text injury assessment report is identified, and a large language model is used to perform semantic understanding on the text content in the injury record area to extract a plurality of the second injury labels.

[0077] In this embodiment, the database stores an injury report template, which includes filling specifications for the injury report, for example, in the form of annotations added to the corresponding position of the injury report template. By semantically understanding these annotations, the injury record area in the injury report template can be determined. Therefore, the text content corresponding to the injury record area in the text injury report produced by the forensic doctor is extracted, and the semantic understanding is performed by the large language model to extract a number of second injury labels.

[0078] Among them, there may be multiple versions of the injury inspection report templates in the database. The corresponding injury inspection report template can be selected according to the type of case involved by the injured person, and the latest injury inspection report template should be selected.

[0079] In some embodiments, when performing matching analysis on the second injury label and the first injury label, if no matching result is obtained, then:

[0080] Obtain the injury location corresponding to the second injury label, and query whether there is a transmissive medical image corresponding to this injury location stored in the database. If so, determine a third injury label corresponding to this injury location based on this transmissive medical image;

[0081] If the second injury label matches the third injury label, then determine the second injury label as the target injury label;

[0082] If the database does not store a transmissive medical image corresponding to this injury location or the second injury label does not match the third injury label, then output a manual confirmation prompt message.

[0083] In this embodiment, if the second injury label hits a first injury label, it indicates that the second injury label is credible. If the second injury label does not hit any first injury label, it indicates that the second injury label obtained by the large language model extraction may be incorrect. In this regard, the present invention further queries whether there is a transmissive medical image corresponding to this injury location stored in the database. If there is and the third injury label corresponding to this transmissive medical image matches the second injury label, it indicates that the second injury label is credible, and at this time it is still determined as the target injury label.

[0084] It should be noted that for internal injuries such as fractures, forensic doctors generally cannot determine them through external observation, but need to confirm them through transmissive medical images such as X-rays and CTs taken, and these transmissive medical images are stored in the database. Obviously, for these internal injuries, the first injury label determined based on the external characteristics of the injuries in the first video stream data is inevitably incorrect. That is to say, the reason for the second injury label not hitting the result may be the error of the first injury label.

[0085] As Figure 2 shown, an embodiment of the present invention also discloses an automatic injury label annotation system, and the system includes an acquisition module, an injury extraction module, a report analysis module, a comparison analysis module, and an injury annotation module;

[0086] The obtaining module is configured to obtain first video stream data of a forensic doctor's injury examination of the injured person, and segment the first video stream data into multiple segments of images according to the injury examination areas;

[0087] The injury extraction module is configured to extract injury characteristics from each segment of image to obtain corresponding first injury labels and corresponding injury positions;

[0088] The report analysis module is configured to receive the written injury examination report generated by the forensic doctor, use a large language model to perform semantic understanding on the written injury examination report, and extract and obtain a number of second injury labels;

[0089] The comparison and analysis module is configured to perform matching analysis on the second injury labels and the first injury labels. If a matching result is obtained, the corresponding second injury label is determined as the target injury label;

[0090] The injury annotation module is configured to annotate each of the target injury labels at the injury positions corresponding to the matched first injury labels in the target image; wherein, the target image is any frame of high-definition image in the first video stream data.

[0091] An embodiment of the present invention also discloses an electronic device, including: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, where the processor executes the computer program to implement the method as described in the foregoing embodiment.

[0092] An embodiment of the present invention also discloses a computer storage medium, where the computer storage medium stores a computer program, and the computer program is executed by a processor to implement the method as described in the foregoing embodiment.

[0093] An embodiment of the present invention also discloses a computer program product, which implements the method as described in the foregoing embodiment when running on a terminal.

[0094] Each component embodiment of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the device according to the embodiments of the present invention. The present invention can also be implemented as a device or device program (for example, a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0095] As used herein, "one embodiment", "an embodiment", or "one or more embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. In addition, it should be noted that examples of the phrase "in one embodiment" herein do not necessarily all refer to the same embodiment.

[0096] In the specification provided herein, a number of specific details are set forth. However, it will be understood that embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure an understanding of this specification.

[0097] It should be noted that the above embodiments are illustrative of the present invention and not restrictive thereof, and alternative embodiments may be designed by those skilled in the art without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.

[0098] In addition, it should also be noted that the language used in this specification has been principally selected for readability and teaching purposes rather than for the purpose of interpreting or limiting the subject matter of the present invention. Thus, many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure of the present invention is illustrative, not restrictive, and the scope of the present invention is defined by the appended claims.

Claims

1. A method for automatically labeling injury labels, characterized in that: The method comprises the following steps: Acquire first video stream data of a forensic doctor performing injury examination on an injured person, and divide the first video stream data into multiple image segments according to the injury examination area; Extract injury features from each image segment to obtain the corresponding first injury label and the corresponding injury location; receiving a text injury report generated by a forensic doctor, using a large language model to perform semantic understanding on the text injury report, and extracting a plurality of second injury labels; Performing a matching analysis on the second injury condition label and the first injury condition label, and if a matching result is obtained, determining the corresponding second injury condition label as a target injury condition label; In a target image, each target injury label is marked at the injury position corresponding to the matched first injury label; wherein the target image is any frame of high-definition image in the first video stream data; The target image is determined by: Acquire second video stream data containing the entire body area of ​​the injured person, and extract the external features and behavioral features of the injured person according to the second video stream data; Predicting the injury area distribution information of the injured person according to the external features and the behavioral features, and determining the category attribute of the target image according to the injury area distribution information, wherein the category attribute includes a full-body image and a non-full-body image; If the determined category attribute of the target image is a full-body image, a high-definition image frame containing a full-body image of the injured person is intercepted from the second video stream data, and a plurality of high-definition image frames corresponding to the second injury label are intercepted from the first video stream data; If the determined category attribute of the target image is a non-full body image, a number of high-definition image frames corresponding to the second injury label are captured from the first video stream data.

2. The method for automatically labeling injury labels according to claim 1, characterized in that: The determining the category attribute of the target image according to the injury area distribution information includes: According to each injury point in the injury area distribution information, two injury points with the farthest distance are determined, and a connected area is drawn on the body image of the injured person based on the two injury points; The ratio of the area of ​​the connected region to the area of ​​the injured person's body image is calculated. If the ratio is greater than a threshold, the category attribute of the target image is determined to be a full-body image; otherwise, the category attribute of the target image is determined to be a non-full-body image.

3. The method for automatically labeling injury labels according to claim 2, characterized in that: The step of dividing the first video stream data into a plurality of image segments according to the injury detection area includes: Identifying the forensic doctor's injury examination action and the injured person's body area in the first video stream data, and determining the injury examination area according to the injury examination action and the injured person's body area; When the injury detection area changes beyond a preset distance, it is determined that there is a switch in the injury detection area, and a moment of a preset time before the switching moment is determined as a segmentation node; The first video stream data is divided into a plurality of image segments according to each of the segmentation nodes.

4. The method for automatically labeling injury labels according to claim 1, characterized in that: A large language model is used to perform semantic understanding on the text injury assessment report, and several second injury labels are extracted, including: Obtaining an injury report template from a database, and determining an injury record area according to the injury report template; An injury record area in the text injury assessment report is identified, and a large language model is used to perform semantic understanding on the text content in the injury record area to extract a plurality of the second injury labels.

5. The method for automatically labeling injury labels according to claim 1, characterized in that: The second injury condition label is matched and analyzed with the first injury condition label, and if no matching result is obtained, then: Obtaining the injury position corresponding to the second injury label, querying whether a transmission medical image corresponding to the injury position is stored in a database, and if so, determining a third injury label corresponding to the injury position according to the transmission medical image; If the second injury condition label matches the third injury condition label, determining the second injury condition label as the target injury condition label; If the database does not store the transmission medical image corresponding to the injury position or the second injury label does not match the third injury label, a manual confirmation prompt message is output.

6. An automatic injury labeling system, the system comprising an acquisition module, an injury extraction module, a report analysis module, a comparison analysis module, and an injury labeling module; The acquisition module is used to acquire first video stream data of the forensic doctor performing injury examination on the injured person, and divide the first video stream data into multiple image segments according to the injury examination area; The injury extraction module is used to extract injury features from each image segment to obtain a corresponding first injury label and a corresponding injury position; The report analysis module is used to receive the text injury report generated by the forensic doctor, use the large language model to perform semantic understanding on the text injury report, and extract a plurality of second injury labels; The comparison analysis module is used to perform matching analysis on the second injury condition label and the first injury condition label, and if a matching result is obtained, determine the corresponding second injury condition label as a target injury condition label; The injury labeling module is used to label each target injury label at the injury position corresponding to the matched first injury label in the target image; wherein the target image is any frame of high-definition image in the first video stream data; The target image is determined by: Acquire second video stream data containing the entire body area of ​​the injured person, and extract the external features and behavioral features of the injured person according to the second video stream data; Predicting the injury area distribution information of the injured person according to the external features and the behavioral features, and determining the category attribute of the target image according to the injury area distribution information, wherein the category attribute includes a full-body image and a non-full-body image; If the determined category attribute of the target image is a full-body image, a high-definition image frame containing a full-body image of the injured person is intercepted from the second video stream data, and a plurality of high-definition image frames corresponding to the second injury label are intercepted from the first video stream data; If the determined category attribute of the target image is a non-full body image, a number of high-definition image frames corresponding to the second injury label are captured from the first video stream data.

7. An electronic device comprising: At least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the method according to any one of claims 1 to 5.

8. A computer storage medium storing a computer program, characterized in that: The computer program is executed by a processor to implement the method according to any one of claims 1 to 5.

9. A computer program product, characterized in that: When the computer program product runs on a terminal, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Method for automatically labeling label, electronic equipment and storage medium

    CN113283509A

  • Video labeling method and device, computer equipment and storage medium

    CN118264870A