Processing Method of Training Samples, Related Devices, Storage Medium and Program Product
By adjusting the number ratio of reference description text in the training sample, the problem of uneven label analysis object in the report generation model is solved, and the accuracy of the model and the quality of the detection report are improved.
Patent Information
- Application Number
- CN202111682891.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-12-31
AI Technical Summary
During the training process, the existing report generation model has low accuracy due to the uneven number of training samples of different labels between each analysis object.
By obtaining the reference description text in the training sample, the target reference description text is determined based on the sample discard rate and label ratio, so that the reference description text number ratio of the first label and the second label of each analysis object is within a preset interval, and a second training sample is constructed to optimize the report generation model.
The sample size ratio equalization between each analysis object is achieved, the accuracy and consistency of the report generation model is improved, and the generated detection report is more accurate.
Smart Images

Figure CN116432017B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a method for processing training samples, related devices, storage media, and program products. Background Art
[0002] The rapid development of medical technologies is inseparable from the assistance of computer science technologies. Nowadays, the automatic generation of medical test reports in the medical industry can also be realized based on computer science technologies. Under normal circumstances, the automatically generated medical test reports can be used to assist doctors in corresponding medical diagnoses, thereby reducing the workload of doctors during medical diagnoses and improving the patient's medical treatment speed. When generating medical test reports, relevant report generation models are usually used to perform text descriptions on medical test images (such as: X-ray films, CT films, etc.). However, the existing report generation models usually have the following problems during the training process: the proportion of the number of training samples with different labels among the various analysis objects included in the training images is unbalanced, which in turn leads to a low accuracy of the medical test reports generated by using the report generation model in actual applications. Therefore, how to balance the quantity ratio of training data with different labels among various analysis objects has become a current research hotspot. Summary of the Invention
[0003] Embodiments of this application provide a method for processing training samples, related devices, storage media, and program products, which can improve the proportion balance of training samples with different labels.
[0004] On the one hand, embodiments of this application provide a method for processing training samples, including:
[0005] Obtaining a first training sample, where the first training sample includes a plurality of training images and a reference test report for each training image, each training image includes a plurality of analysis objects, and the reference test report includes reference description texts for each analysis object;
[0006] Traversing the reference description texts of each analysis object among the plurality of analysis objects of the first training sample, and determining a target reference description text among the plurality of reference description texts based on the sample discard rate of each analysis object and the label of each reference description text among the plurality of reference description texts of each analysis object, and the ratio of the number of reference description texts with a first label to the number of reference description texts with a second label in the target reference description text of each analysis object is within a preset interval;
[0007] Determine a second training sample based on the target reference description text of each analysis object, the reference detection report to which each target reference description text belongs, and the multiple training images; the second training sample includes the multiple training images and the training detection report of each training image, and the training detection report includes the target reference description text in the reference detection report corresponding to the training image, and the second training sample is used to optimize the model of the report generation model.
[0008] On the one hand, an embodiment of the present application further provides a processing device for training samples, including:
[0009] An acquisition unit, configured to acquire a first training sample, where the first training sample includes multiple training images and the reference detection report of each training image, each training image includes multiple analysis objects, and the reference detection report includes the reference description text of each analysis object;
[0010] A traversal unit, configured to traverse the multiple reference description texts of each analysis object among the multiple analysis objects of the first training sample, and determine a target reference description text among the multiple reference description texts based on the sample discard rate of each analysis object and the label of each reference description text among the multiple reference description texts of each analysis object, and the ratio of the number of reference description texts with the first label to the number of reference description texts with the second label in the target reference description text of each analysis object is within a preset interval;
[0011] A training sample determination unit, configured to determine a second training sample based on the target reference description text of each analysis object, the reference detection report to which each target reference description text belongs, and the multiple training images; the second training sample includes the multiple training images and the training detection report of each training image, and the training detection report includes the target reference description text in the reference detection report corresponding to the training image, and the second training sample is used to optimize the model of the report generation model.
[0012] On the one hand, an embodiment of the present application further provides a computer device, including:
[0013] A processor, the processor is adapted to implement one or more computer programs;
[0014] A computer storage medium, the computer storage medium stores one or more computer programs, and the one or more computer programs are adapted to be loaded and executed by the processor:
[0015] Obtain a first training sample, where the first training sample includes a plurality of training images and a reference detection report for each training image. Each training image includes a plurality of analysis objects, and the reference detection report includes reference description texts for each analysis object. Traverse the multiple reference description texts of each analysis object among the multiple analysis objects of the first training sample, and based on the sample discard rate of each analysis object and the labels of each reference description text among the multiple reference description texts of each analysis object, determine a target reference description text among the multiple reference description texts. The ratio of the number of reference description texts with a first label to the number of reference description texts with a second label in the target reference description text of each analysis object is within a preset range. Based on the target reference description text of each analysis object, the reference detection report to which each target reference description text belongs, and the multiple training images, determine a second training sample. The second training sample includes the multiple training images and a training detection report for each training image. The training detection report includes the target reference description text in the reference detection report corresponding to the training image. The second training sample is used to optimize the model of the report generation model.
[0016] On the one hand, an embodiment of the present application further provides a computer storage medium, which stores one or more computer programs, and the one or more computer programs are adapted to be loaded and executed by a processor:
[0017] Obtain a first training sample, where the first training sample includes a plurality of training images and a reference detection report for each training image. Each training image includes a plurality of analysis objects, and the reference detection report includes reference description texts for each analysis object. Traverse the multiple reference description texts of each analysis object among the multiple analysis objects of the first training sample, and based on the sample discard rate of each analysis object and the labels of each reference description text among the multiple reference description texts of each analysis object, determine a target reference description text among the multiple reference description texts. The ratio of the number of reference description texts with a first label to the number of reference description texts with a second label in the target reference description text of each analysis object is within a preset range. Based on the target reference description text of each analysis object, the reference detection report to which each target reference description text belongs, and the multiple training images, determine a second training sample. The second training sample includes the multiple training images and a training detection report for each training image. The training detection report includes the target reference description text in the reference detection report corresponding to the training image. The second training sample is used to optimize the model of the report generation model.
[0018] On the one hand, an embodiment of the present application further provides a computer program product or a computer program. The computer program product includes a computer program, and the computer program is adapted to be loaded and executed by a processor:
[0019] Obtain a first training sample, where the first training sample includes a plurality of training images and a reference detection report for each training image. Each training image includes a plurality of analysis objects, and the reference detection report includes reference description texts for each analysis object; traverse the reference description texts of each analysis object among the plurality of analysis objects in the first training sample, and based on the sample discard rate of each analysis object and the labels of each reference description text among the plurality of reference description texts of each analysis object, determine a target reference description text among the plurality of reference description texts. The ratio of the number of reference description texts with a first label to the number of reference description texts with a second label in the target reference description text of each analysis object is within a preset interval; based on the target reference description text of each analysis object, the reference detection report to which each target reference description text belongs, and the plurality of training images, determine a second training sample; the second training sample includes the plurality of training images and a training detection report for each training image. The training detection report includes the target reference description texts in the reference detection report corresponding to the training image. The second training sample is used to optimize the model of the report generation model.
[0020] In the embodiments of the present application, the first training sample includes a plurality of training images and a reference detection report for each training image. Each training image includes a plurality of analysis objects, each reference detection report includes reference description texts for each analysis object, and the label of the reference description text is a first label or a second label. The computer device can determine the target reference description text of each analysis object from each reference detection report based on the label of the reference description text and the sample discard rate of each analysis object, so that the ratio of the number of reference description texts with a first label to the number of reference description texts with a second label in the target reference description text of each analysis object is within a preset interval, achieving the balance between the positive and negative sample quantity ratios of each analysis object. Therefore, after optimizing the model of the report generation model using the second training sample, the generated detection report of the optimized report generation model is more accurate. Description of the Drawings
[0021] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1a It is a schematic diagram of a training image and an analysis object provided by the embodiments of the present application;
[0023] Figure 1bIt is a schematic diagram of a process for generating a detection report provided by an embodiment of the present application;
[0024] Figure 2 It is a schematic flowchart of a method for processing training samples provided by an embodiment of the present application;
[0025] Figure 3 It is a schematic flowchart of another method for processing training samples provided by an embodiment of the present application;
[0026] Figure 4 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application;
[0027] Figure 5 It is a schematic flowchart of a model optimization method provided by an embodiment of the present application;
[0028] Figure 6 It is a schematic diagram of the comparison result of a detection report provided by an embodiment of the present application;
[0029] Figure 7 It is a schematic diagram of the structure of a device for processing training samples provided by an embodiment of the present application;
[0030] Figure 8 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0031] The embodiments of this application are based on Artificial Intelligence (AI) technology and propose a method for processing training samples. Using the processed training samples obtained by this method to optimize the report generation model can enable the optimized report generation model to generate more accurate detection reports. Among them, artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. It can be understood that artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, processing technologies for large training samples, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. In the embodiments of this application, computer vision technology and natural language processing technology in artificial intelligence technology are mainly utilized.
[0032] In an actual model optimization scenario, before optimizing the report generation model, it is necessary to collect training samples in the corresponding medical field. Training samples usually include training images and reference detection reports of the training images. The training images can include multiple analysis objects, and the training images can be medical detection images in any medical discipline. Medical detection images include images of at least one physiological structure, and the number of physiological structures included in the same type of medical detection image is the same. Exemplarily, medical disciplines can be, for example, radiology and ultrasonography; training images can include but are not limited to: radiology detection images (such as: X-ray films, XR-ray films, CT chest films, etc.), ultrasonic detection images (such as: B-ultrasound images, color Doppler ultrasound images, etc.), Magnetic Resonance Imaging (MRI), etc. The analysis object can be the physiological structure in the training image. Then, exemplarily, the multiple analysis objects included in the training image can be all the physiological structures in the training image or part of the physiological structures in the training image. Please combine Figure 1a for understanding. Figure 1a The image marked by 10 in Figure 1a can be understood as a training image, and the analysis object can be Figure 1aThe physiological structure marked by 102 and the physiological structure marked by 103 can also be used as analysis objects. It should be noted that when there are multiple training images in the training sample, the analysis objects included in each training image are the same. That is to say, it can be understood that the number of analysis objects included in each training image is also the same.
[0033] Among them, the understanding of "the analysis objects included in each training image are the same" can be combined with the following example: Suppose there are 3 medical detection images in the training sample, and each medical detection image includes 3 physiological structures (physiological structure a, physiological structure b, physiological structure c); then, when these 3 medical detection images are used as training images, the physiological structures determined as analysis objects in each medical detection image are the same. For example, the physiological structure a and the physiological structure b in each medical detection image are determined as analysis objects, or the physiological structure b and the physiological structure c in each medical detection image are determined as analysis objects, etc. It can be seen that the number of analysis objects between each training image is the same (both are 2).
[0034] The reference detection report in the training sample can be used to describe the health information of each analysis object in the training image. Specifically, the reference detection report can include multiple reference description texts. In practical applications, the reference description texts in the reference detection report can be used to describe the health information of the analysis object. That is to say, the role of the reference description text can be understood as: explaining the health status of the analysis object in the form of text description. A reference description text can include one or more description statements. And specifically, in a reference detection report, an analysis object can correspond to at least one reference description text, and the health status of the analysis object indicated by each reference description text in these at least one reference description texts is the same (that is: healthy state, or non-healthy state). That is to say, suppose the training image includes analysis object A, and analysis object A is in a non-healthy state. Then, when there are multiple reference description texts of analysis object A in the reference detection report of this training image, each of these multiple reference description texts can only be used to indicate that the analysis object is in a non-healthy state. In addition, each reference description text also has a label, and the label can include any one of the first label and the second label. In this case, the reference detection report with the first label can be used to describe that the corresponding analysis object is in a healthy state, and the reference detection report with the second label can be used to describe that the corresponding analysis object is in a non-healthy state. It can be understood that when an analysis object corresponds to at least one reference description text in a reference detection report, the labels of each reference description text in these at least one reference description texts are the same.
[0035] For the convenience of clearly elaborating the method provided by this application in subsequent embodiments, unless otherwise specified, the description is given by taking one reference description text corresponding to one analysis object in the same reference detection report as an example. In addition, for the convenience of understanding, the reference description text of the first label in this application can be understood as the positive sample of the corresponding analysis object, and the reference description text of the second label can be understood as the negative sample of the corresponding analysis object.
[0036] In the process of collecting the above training samples, the embodiments of this application also found that: the difficulty of collecting the reference description text of each analysis object is different, and the reference description text of the first label of each analysis object is easier to collect than the reference description text of the second label of the same analysis object. This results in the number of reference description texts of the first label of each analysis object in the collected training samples being much larger than the number of reference description texts of the second label, for example: the number of collected reference description texts of the first label is usually more than twice the number of reference description texts of the second label. In addition, the ratio of the number of reference description texts of the first label to the number of reference description texts of the second label varies greatly among different analysis objects. In this case, if the computer device directly uses the collected training samples to optimize the report generation model, it will lead to inconsistent accuracies of the descriptions of each analysis object and a large deviation between the accuracies when the optimized report generation model generates the detection report of the medical detection image, thereby resulting in a low accuracy of the entire detection report and low practicality of the optimized report generation model. Similarly, for the convenience of the description of subsequent embodiments, the "ratio of the number of reference description texts of the first label to the number of reference description texts of the second label" is hereinafter understood as: the ratio of the number of positive and negative samples (abbreviation: positive and negative sample ratio).
[0037] To solve the above problems, the present application proposes a method for processing training samples based on artificial intelligence technology, which can be executed by a computer device. Among them, the computer device can be a terminal device or a server. Specifically, the terminal device may include, but is not limited to: smart phones, tablet computers, laptop computers, desktop computers, vehicle-mounted terminals, smart TVs, etc.; various clients (applications, APPs) can run inside the terminal device, such as multimedia playback clients, social clients, browser clients, information flow clients, education clients, and so on. The server may include, but is not limited to: independent physical servers, server clusters or distributed systems composed of multiple physical servers, and cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network, content delivery network), and big data and artificial intelligence platforms, etc. The present application does not limit this.
[0038] In a specific embodiment, the general principle of the processing method for the training samples can be as follows: The computer device first obtains a first training sample, which includes a plurality of training images and the reference detection report for each training image. After the computer device obtains the first training sample, the computer device can determine the target reference description text for the analysis object from the reference detection report of the training image based on the sample discard rate of each analysis object in each training image and the labels of each reference description text (i.e., the first label or the second label) in the reference detection report of the training image. Since there are multiple training images, the computer device can determine multiple target reference description texts for each analysis object based on the first training sample. These multiple target reference description texts include the reference description texts with the first label and the reference description texts with the second label of the analysis object. Further, after the computer device determines the multiple target reference description texts corresponding to each analysis object, the computer device can construct a second training sample according to the determined target reference description texts, the labels corresponding to the target reference description texts, and all the training images in the first training sample. Among them, since the computer device determines the multiple target reference description texts for each analysis object based on the sample discard rate, and the sample discard rate in the embodiments of the present application is mainly used to probabilistically discard the reference description texts with the first label of each analysis object in the reference detection report, the number of reference description texts with the first label in the multiple target reference description texts of each analysis object can be reasonably reduced, and further, the ratio of the number of reference description texts with the first label to the number of reference description texts with the second label is within a preset interval (i.e., the ratio of the number of positive and negative samples of each analysis object is within a preset interval). It should be noted that if the ratio of the number of positive and negative samples of each analysis object is within a preset interval, it means that the ratio of the number of positive and negative samples of each analysis object is in a balanced state. In this case, if the computer device uses the second training sample to optimize the report generation model, a detection report with higher accuracy can be generated through the optimized report generation model.
[0039] Since the second training sample determined in the embodiments of the present application includes the target reference description text of each analysis object, it can be understood that the report generation model optimized by using the second training sample of the embodiments of the present application can generate the description text of each analysis object. Exemplarily, in the embodiments of the present application, the general process of the optimized report generation model generating a detection report can be referred to Figure 1b . As Figure 1bAs shown in the figure, the computer device can adopt an optimized report generation model to generate a description text for each analysis object in the medical detection image. Further, the computer device can adopt the optimized report generation model to obtain a complete medical detection report based on the description text of each analysis object. The obtained medical detection report can be used to assist the corresponding medical staff in medical diagnosis. In the subsequent embodiments, the model optimization method of the report generation model and the method for the computer device to generate a detection report by using the optimized report generation model will be elaborated in detail, and this application will not elaborate here.
[0040] Please refer to Figure 2 as shown in the figure Figure 2 is a schematic flowchart of a method for processing training samples proposed by an embodiment of the present application based on the above principle. The method indicated by this flowchart can be executed by the computer device mentioned above. As Figure 2 shown in the figure, the method includes steps S201 - S203:
[0041] S201, obtain a first training sample. The first training sample includes multiple training images and a reference detection report for each training image. Each training image includes multiple analysis objects, and the reference detection report includes a reference description text for each analysis object.
[0042] In a specific embodiment, the training images in the first training sample may include multiple analysis objects, and the number of analysis objects included in each training image is the same. Then, optionally, the number of reference description texts included in the reference detection report may be the same as the number of analysis objects included in the training image. In this case, one analysis object corresponds to one reference description text, and one reference description text has one label. Optionally, the number of reference description texts included in the reference detection report may also be more than the number of analysis objects in the training image. In this case, one analysis object may correspond to one or more reference description texts, and the labels of each of these one or more reference description texts are the same. For the convenience of description, unless otherwise specified, in the following embodiments, the case where one analysis object in the reference detection report corresponds to one reference description text is used as an example to elaborate in detail on the method provided by this application.
[0043] S202, traverse the multiple reference description texts of each analysis object among the multiple analysis objects in the first training sample, and determine a target reference description text from the multiple reference description texts based on the sample discard rate of each analysis object and the label of each reference description text among the multiple reference description texts of each analysis object.
[0044] Among them, traversal can be understood as obtaining or processing one by one. Since one analysis object corresponds to one reference description text in the reference detection report of this application, and the first training sample includes multiple reference detection reports, it can be understood that each analysis object in this application corresponds to multiple reference description texts. Optionally, when the computer device traverses the multiple reference description texts of each analysis object in the first training sample, it can be understood that: during the traversal process, the computer device first traverses the multiple analysis objects to determine the analysis object for which the target reference description text needs to be determined currently, and then, the computer device can traverse the multiple reference description texts corresponding to the determined analysis object to sequentially determine whether each reference description text in the multiple reference description texts is used as the target reference description text. In this case, when the computer device determines the target reference description text among the multiple reference description texts, the multiple reference description texts can refer to the multiple reference description texts corresponding to each analysis object. Optionally, when the computer device traverses the multiple reference description texts of each analysis object in the first training sample, it can also be understood that: during the traversal process, the computer device sequentially obtains the reference detection reports of each training image, and then analyzes each reference description text in the multiple reference description texts included in each reference detection report to determine whether the corresponding reference description text is used as the target reference description text. In this case, when the computer device determines the target reference description text among the multiple reference description texts, the multiple reference description texts can refer to the multiple reference description texts included in each reference detection report.
[0045] In practical applications, when the computer device determines whether any reference description text is used as the target reference description text, it can be determined based on the label of the reference description text and the sample discard rate of the analysis object corresponding to the reference description text. Among them, the sample discard rate is mainly used to constrain the computer device to determine whether the reference description text with the label of the first label is discarded according to the probability indicated by the sample discard rate. If the reference description text with the first label is not discarded, the reference description text is used as the target reference description text.
[0046] S203. Determine a second training sample based on the target reference description text of each analysis object, the reference detection report to which each target reference description text belongs, and multiple training images.
[0047] Among them, the second training sample includes a plurality of training images and the training detection reports of each training image. The training detection report includes the target reference detection text in the reference detection report corresponding to the training image. It is not difficult to understand that the target reference description text can be specifically understood as: the reference description text retained in the reference detection report. In the embodiments of the present application, it can be understood that one training image corresponds to one reference detection report and one training detection report, and the training detection report corresponding to any one of the training images is determined according to the reference detection report corresponding to the training image. In this case, the number of target reference detection texts in the training detection report of any one training image is less than or equal to the number of reference detection texts in the reference detection report of any one training image. For example, assume that the reference detection report corresponding to training image A is Report 1, and Report 1 includes a plurality of reference description texts; the training detection report corresponding to training image A is Report 2, and Report 2 includes at least one target reference description text. Then, it can be understood that the target reference description texts in Report 2 are determined based on the reference description texts in Report 1, and the reference detection report to which any target reference description text in Report 2 belongs is Report 1.
[0048] In the embodiments of the present application, the first training sample includes a plurality of training images and the reference detection reports of each training image. Each training image includes a plurality of analysis objects, and each reference detection report includes the reference description text of each analysis object. The label of the reference description text is the first label or the second label. The computer device determines the target reference description text of each analysis object from each reference detection report based on the label of the reference description text and the sample discard rate of each analysis object, so that the computer device can construct a second training sample based on the plurality of training images in the first training sample and the target reference description text of each analysis object. Among them, since the reference description texts with the first label are usually more than those with the second label, and the sample discard rate is mainly used to probabilistically discard the reference description texts with the first label, then, it can be understood that the ratio of the number of reference description texts with the first label to the number of reference description texts with the second label in the target reference description text of each analysis object determined in the embodiments of the present application can all be within a preset interval, so that the difference in the ratio of the number of positive and negative samples of each analysis object is small. Then, in this case, after optimizing the report generation model with the second training sample, the optimized report generation model has similar sensitivity to each analysis object, so that the detection report generated by the optimized report generation model is more accurate.
[0049] Based on the above description of Figure 2 the present application embodiments also provide a method for processing training samples. This method can be executed by the computer device mentioned above, please refer to Figure 3, the method includes steps S301 - S304:
[0050] S301, obtain a first training sample, the first training sample includes a plurality of training images and a reference detection report for each training image, each training image includes a plurality of analysis objects, and the reference detection report includes reference description text for each analysis object.
[0051] In the actual training sample collection scenario, since each reference detection report of the first training sample includes the reference description text for each analysis object, and medical detection reports usually only describe the health information of analysis objects in a non - healthy state, it can be seen that it is usually impossible to directly collect the reference detection reports of training images when collecting training samples. However, experiments have shown that to enable the optimized report generation model to generate high - quality (i.e., complete and highly accurate) medical detection reports, it is usually necessary to use training samples with the reference description text for each analysis object to optimize the report generation model. Based on this, in the embodiments of this application, in order to enable the optimized report generation model to generate high - quality medical detection reports, the computer device can preferentially select detection reports including the reference description text for each analysis object as the reference detection reports in the first training sample.
[0052] Then, it can be understood that the computer device in the embodiments of this application can directly collect the target detection reports of training images, and further, the computer device can fill the collected target detection reports to obtain the reference detection reports for the corresponding training images. Among them, the target detection report includes the reference description text for at least one analysis object, that is to say, the target detection report may only include the description text for some analysis objects in the training image. And the target detection report can include a plurality of detection statements, and each detection statement includes the object identifier for at least one analysis object. It should be noted that the object identifier is used to uniquely identify the corresponding analysis object, that is to say, there is a one - to - one correspondence between the object identifier and the analysis object, that is: one analysis object corresponds to one object identifier, and one object identifier also only corresponds to one analysis object. In practical applications, the object identifier can be a text character (such as: the name of the analysis object), a numeric character (such as: the number of the analysis object), or any other form of identifier, and the embodiments of this application do not limit this.
[0053] The principle of a computer device obtaining a reference detection report for each training image (i.e., the principle of the computer device filling in the target detection report) is elaborated below. It should be noted that in the following elaboration, specifically, one analysis object corresponds to one reference description text in the reference detection report. Based on this, the principle of the computer device obtaining a reference detection report for each training image can be as follows: After the computer device obtains the target detection report for each training image, further, the computer device can determine the number of reference description texts in the target detection report of each training image. If the computer device determines that the number of reference description texts in the target detection report is less than the number of analysis objects in each training image, it means that the target detection report only provides text descriptions for the health information of some of the analysis objects. Then, in this case, the computer device can fill in the text of the target detection report to obtain the reference detection report for each training image, so that there is a corresponding description text for the health information of each analysis object in the training image. It can be understood that correspondingly, if the computer device determines that the number of reference description texts in the target detection report is equal to the number of analysis objects in each training image, it means that there are already description texts for the health information of each analysis object in the target detection report. Then, in this case, the computer device can directly use this target detection report as the reference detection report for the corresponding training image.
[0054] Among them, the specific method for the computer device to determine the number of reference description texts in the target detection report of each training image can be as follows: The computer device first obtains each detection statement in the target detection report. Then, the computer device can obtain all the object identifiers included in each detection statement, and obtain the weight of each object identifier among all the object identifiers. Further, the computer device can determine the target object identifier with the largest weight from each detection statement, so that the computer device can generate a reference description text based on the detection statements with the same target object identifier. That is to say, among the reference description texts generated by the computer device based on all the detection statements in the target detection report, the target object identifiers corresponding to different reference description texts are different. For example, when N (N is a positive integer) different target object identifiers are determined in the target detection report, the computer device can generate N reference description texts based on all the detection statements in the target detection report. Further, after the computer device generates the reference description text based on the determined target object identifier, the computer device can obtain the number of the generated reference description texts, so as to determine the number of reference description texts in the target detection report.
[0055] In a specific embodiment, the weight of each object identifier can be determined by the computer device based on the target detection report corresponding to each training image in the first training sample, and its specific determination method can be combined with Figure 4The structure of the computer device shown will be described. As Figure 4 shown, the computer device may include a knowledge graph-based data processing module (KGDP). The structure of the KGDP can be exemplarily referred to as shown by the structure labeled 401 in Figure 4 . It should be noted that the KGDP can be used to fill the target detection report. Therefore, in order to clearly elaborate on the method proposed in the embodiments of the present application in subsequent embodiments, hereinafter, the KGDP (i.e., the knowledge graph-based data processing module) is referred to as: the report filling module. Also, as Figure 4 shown, the computer device may further include a task-aware report generation module (TRG). The structure of the task-aware report generation module can be exemplarily referred to as shown by the structure labeled 402 in Figure 4 . In the embodiments of the present application, the task-aware report generation module can be used to generate the detection report of the medical detection image. Therefore, similarly, in order to clearly elaborate on the method proposed in the embodiments of the present application in subsequent embodiments, hereinafter, the task-aware report generation module is referred to as: the report generation module.
[0056] Among them, since the report filling module is used to fill the text of the target detection report that needs text filling, it can be understood that the computer device can use the report filling module to determine the weight of each object identifier. Specifically, the computer device can first obtain the target detection report of each training image from the database, so that the computer device can further obtain the object identifiers included in each detection statement in each target detection report. Exemplarily, when the computer device obtains the object identifier, it can be implemented by using the report extractor in the report filling module. And during the process of the computer device obtaining the object identifier, the computer device can also count the number of times each object identifier appears in the target detection report of each training image to implement the statistics of the total number of times each object identifier appears in all the target detection reports corresponding to the first training sample. Further, when the computer device obtains the total number of times each object identifier appears, the computer device can set the weight for the object identifier according to the number of times the object identifier appears. Exemplarily, the computer device can arrange the object identifiers in ascending order of the number of times they appear, and in the arranged object identifier sequence, set the weights for each object identifier in descending order of the weights.
[0057] The following uses specific examples to elaborate in detail on the method of setting weights for object identifiers by a computer device. Suppose the first training sample includes 3 training images, namely training image 1, training image 2, and training image 3; each training image includes 3 analysis objects, namely analysis object a, analysis object b, and analysis object c, and each training image corresponds to a target detection report. Then, if training image 1 corresponds to target detection report 1, training image 2 corresponds to target detection report 2, and training image 3 corresponds to target detection report 3. Then, all the target detection reports corresponding to the first training sample can be target detection report 1, target detection report 2, and target detection report 3. Further, if the computer device counts the number of times the object identifiers appear in each target detection report, it is found that in target detection report 1, the object identifier A of analysis object a appears 10 times, the object identifier B of analysis object b appears 3 times, and the object identifier C of analysis object c appears 0 times; in target detection report 2, the object identifier A of analysis object a appears 1 time, the object identifier B of analysis object b appears 4 times, and the object identifier C of analysis object c appears 5 times; in target detection report 3, the object identifier A of analysis object a appears 0 times, the object identifier B of analysis object b appears 0 times, and the object identifier C of analysis object c appears 4 times. Then, in this case, it is not difficult to understand that the total number of times the object identifier A appears in all the target detection reports corresponding to the first training sample is 10 + 1 + 0 = 11 times, the total number of times the object identifier B appears in all the target detection reports corresponding to the first training sample is 3 + 4 + 0 = 7 times, and the total number of times the object identifier C appears in all the target detection reports corresponding to the first training sample is 0 + 5 + 4 = 9 times. It can be further understood that after the computer device arranges the object identifiers in ascending order of the total number of appearances, the obtained object identifier sequence can be [object identifier B, object identifier C, object identifier A]. In this case, the computer device can set weights for each object identifier in the object identifier sequence in descending order of weight. That is to say, the computer device can set the weight of object identifier B to be the largest among the three object identifiers, the weight of object identifier C to be the second largest, and the weight of object identifier A to be the smallest among the three object identifiers. It can be understood that after the computer device sets weights for each object identifier, if the object identifiers are arranged in descending order of weight, the following arrangement order can be obtained: object identifier B, object identifier C, object identifier A.
[0058] In addition, it should be noted that in practical applications, the computer device can also determine the physiological structures that need to be concerned in the medical detection images through the report filling module, so that the computer device can determine the analysis objects in the training images. That is to say, the multiple analysis objects in the training images can be some or all of the physiological structures included in the training images. Based on this, the following elaborates in detail on the method by which the computer device determines the analysis objects in the training images.
[0059] In practical applications, when determining the analysis objects of the training images, the computer device can first analyze multiple target detection reports in the database based on the above-mentioned method of obtaining the total number of occurrences of the object identifiers by the computer device, so as to obtain the total number of occurrences of the structure identifiers of each physiological structure in the training images in the multiple target detection reports. Further, the computer device can sequentially determine multiple analysis objects from the physiological structures included in the training images according to the total number of occurrences. Exemplarily, the computer device can first arrange the structure identifiers of each physiological structure in descending order of the total number of occurrences, and then the computer device can select the first M (M is a positive integer) structure identifiers from the arranged structure identifiers, so that the computer device can determine M analysis objects based on the physiological structures corresponding to the M structure identifiers. Of course, in other implementation manners, the computer device can also arrange the structure identifiers of each physiological structure in ascending order of the total number of occurrences and then determine the analysis objects, or the computer device can directly obtain the structure identifier with the most occurrences, the structure identifier with the second most occurrences,... until the computer device determines M analysis objects without arranging the structure identifiers and directly based on the total number of occurrences of the structure identifiers.
[0060] The following details the specific implementation method of filling the target detection report to obtain the reference detection report during the process of a computer device obtaining the reference detection report of the training image. Specifically, the computer device can first obtain all the analysis objects corresponding to the target detection report. Exemplarily, when the computer device obtains all the analysis objects corresponding to the target detection report, it can first determine the analysis objects described by each reference description text in the target detection report. After the computer device determines the analysis objects described by each reference description text, the computer device can use the analysis objects described by each reference description text in the target detection report as the analysis objects corresponding to the target detection report, so that the computer device can obtain all the analysis objects corresponding to the target detection report. Then, further, the computer device can determine the target analysis objects other than the obtained analysis objects among the multiple analysis objects of each training image, so that after the computer device determines the target analysis objects, it can obtain the reference description text of each target analysis object from the description text database, and thus the computer device can fill the reference description text of each target analysis object into the target detection report to obtain the reference detection report of each training image.
[0061] Among them, the description text library can be pre-constructed by the computer device. Exemplarily, the computer device can use Figure 4 the structure marked by 401 in (i.e., the report filling module) to construct the description text library. The following details the specific method of the computer device constructing the description text library in combination with the report filling module. Specifically, after the computer device sets weights for each object identifier, it can group each detection statement included in multiple target detection reports in the database to obtain the description text of each analysis object. It should be noted that the number of description texts of each analysis object is at least one. Further, the computer device can use the description text of each analysis object to construct a knowledge graph (i.e., the description text library) so that the computer device can obtain the reference description text of the corresponding analysis object from the description text library. Exemplarily, when the computer device groups each detection statement, it can first obtain the object identifier with the largest weight in each detection statement as the target object identifier of the detection statement. Further, the computer device can add the detection statements with the same target weight identifier to the same detection statement group. It can be understood that one detection statement group corresponds to one analysis object. Based on this, then, the computer device can generate at least one description text of the corresponding analysis object based on the detection statements included in each detection statement group.
[0062] S302. Determine the sample discard rate of each analysis object among the multiple analysis objects based on the sample balance parameter.
[0063] Among them, the parameter value of the sample balance parameter can be preset. Exemplarily, the sample balance parameter can be represented by α. In a specific embodiment, the sample balance parameter can be used to determine the sample discard rate of each analysis object. Since the sample balance rate is mainly used to probabilistically discard the reference description text of the first label, the ratio of the number of positive and negative samples in each analysis can be within a preset range. That is to say, the sample balance parameter can be used to constrain the ratio of the number of positive and negative samples of each analysis object in the second training sample. Further, it can be understood that the sample balance parameter can be used to determine the preset range. Exemplarily, the preset range of this application can be [0, α].
[0064] The following details the method for the computer device to determine the sample discard rate of any analysis object. In a specific embodiment, the computer device can first obtain the label of each reference description text in multiple reference description texts of any analysis object. Further, after the computer device obtains the label of each reference description text, the computer device can determine the number of reference description texts of the first label and the number of reference description texts of the second label in the multiple reference description texts of the analysis object. It can be understood that the multiple reference description texts of the analysis object refer to all the reference description texts of the analysis object in the first training sample. Further, the computer device can determine the sample discard rate of the analysis object according to the sample balance parameter, the number of reference description texts of the first label, and the number of reference description texts of the second label.
[0065] Among them, any analysis object in the training image can be called the i-th (i is a positive integer) analysis object among multiple analysis objects. Then, when the computer device determines the sample discard rate of the i-th analysis object, it can first obtain the number of all reference description texts of the second label corresponding to the i-th analysis object in the first training sample (denoted by Na), and obtain the number of all reference description texts of the first label corresponding to the i-th analysis object in the first training sample (denoted by Nn). Further, the computer device can obtain the ratio between the number of reference description texts of the second label and the number of reference description texts of the first label, that is: And the computer device can further perform a multiplication operation on the obtained ratio and the sample balance parameter to obtain the multiplication operation result, that is, obtain: After the computer device obtains the multiplication operation result, the computer device can obtain the sample discard rate constraint parameter value (e.g., 1), and the sample discard rate constraint parameter value can be preset. It should be noted that the computer device can also obtain the sample discard rate constraint parameter value during or before the computer device calculates the multiplication operation result. The embodiments of the present application do not specifically limit the execution order of related steps. Further, the computer device can use the difference between the discard rate constraint parameter value and the multiplication operation result as the first candidate discard rate, that is: use as the first candidate discard rate. In addition, the computer device can also obtain the second candidate discard rate, and the second candidate discard rate can also be preset (e.g., 0). Then, when the computer device obtains the first candidate discard rate and the second candidate discard rate, the computer device can use the largest candidate discard rate among the first candidate discard rate and the second candidate discard rate as the sample discard rate of the i-th analysis object. Based on the above description, it can be understood that the computer device can exemplarily use the method shown in Equation 1 to determine the sample discard rate of the i-th analysis object.
[0066]
[0067] where p i is the sample discard rate of the i-th analysis object; 1 represents the sample discard rate constraint parameter value; α represents the sample balance parameter (exemplarily, the sample balance parameter in the embodiments of the present application can be 2); Na represents the number of reference description texts of the first label in the multiple reference description texts corresponding to the i-th analysis object; Nn represents the number of reference description texts of the second label in the multiple reference description texts corresponding to the i-th analysis object; 0 represents the second candidate discard rate. It should be noted that in other embodiments, the second candidate discard rate, the sample balance parameter, and the sample discard rate constraint parameter value can also be other values, and the embodiments of the present application do not specifically limit this.
[0068] S303. Traverse the multiple reference description texts of each analysis object among the multiple analysis objects of the first training sample. If the label of any reference description text is the first label, determine the text retention rate corresponding to the reference description text based on the sample discard rate of each analysis object, and when the text retention rate is a non-zero value, use the reference description text as the target reference description text.
[0069] Among them, the text retention rate can be understood as the probability that the computer device uses the reference description text of the first tag as the target reference description text. Moreover, in the embodiments of the present application, the text retention rate of the reference description text of the first tag is determined by the computer device based on the sample discard rate of the analysis object corresponding to the reference description text of the first tag. Specifically, for the reference description text of any first tag of the i-th analysis object, the computer device can first use a random value generation function to generate the text reference retention rate of the reference description text under the constraint of the sample discard rate. Exemplarily, the text reference retention rate can be represented by Rand(p i ) and 0 < Rand(p i ) ≤ 1. Then, further, the computer device can obtain the text retention rate constraint parameter value, which can be preset, and exemplarily, the text retention rate constraint parameter value can be 1. After the computer device obtains the text retention constraint parameter value, the computer device can further determine the text retention rate of the relevant reference description text of the i-th analysis object based on the text retention rate constraint parameter value.
[0070] Among them, when the computer device determines the text retention rate of the reference description text, specifically, it can use the difference between the text retention rate constraint parameter value and the text reference retention rate as the text retention rate corresponding to the reference description text. That is to say, the computer device can use 1 - Rand(p i ) as the text retention rate of the i-th reference description text. Since the text retention rate constraint parameter value can be 1 and 0 < Rand(p i ) ≤ 1, then, that is to say, there can be a reference description text with a text retention rate of 0 in the i-th analysis object; if the text retention rate of any reference description text is 0, the computer device will discard the reference description text. It can be seen that the text retention rate constraint parameter value can also constrain the text retention rate corresponding to the corresponding reference description text. Correspondingly, it can be understood that if the text retention rate of any reference description text is a non-zero value, the computer device can determine the any reference description text as the target reference description text.
[0071] In addition, during the process of the computer device traversing multiple reference description texts of each analysis object among multiple analysis objects of the first training sample, if the computer device determines that the label of any reference description text is the second label, the computer device can directly determine the reference description text as the target reference description text. Based on the above determination method of the target reference description text, it can be seen that by using the method provided in the embodiments of the present application, the computer device can reduce the difference between the number of reference description texts of the first label and the number of reference description texts of the second label, so that the number of reference description texts of the first label and the number of reference description texts of the second label are within a preset range. It should also be noted that in the embodiments of the present application, among multiple reference description texts of the i-th analysis object, each reference description text corresponds to a text retention rate, and the text retention rates between different reference description texts are not necessarily the same, because the text retention rate of the reference description text is related to the text reference retention rate generated by the random generation function.
[0072] S304. Determine a second training sample based on the target reference description text of each analysis object, the reference detection report to which each target reference description text belongs, and multiple training images.
[0073] Among them, based on the foregoing description, it is not difficult to understand that the target reference description text of each analysis object is determined from multiple reference description texts of the analysis object, and each reference description text has a corresponding reference detection report, which also makes each target reference description text have a corresponding reference detection report. For example, assume that the reference description text a exists in the reference detection report 1. Then, if the computer device uses the reference description text a as the target reference description text, the reference detection report to which the target reference description text belongs is the reference detection report 1.
[0074] Based on this, in the process of the computer device determining the second training sample, the computer device can first determine, based on the reference detection report to which each target reference description text belongs, the target reference description text belonging to any reference detection report among the target reference description texts of multiple analysis objects. That is to say, the computer device can determine, from all the target reference description texts, the target reference description texts belonging to the same reference detection report. Further, the computer device can generate a training detection report for the training image corresponding to this reference detection report. It can be understood that the training detection report includes the determined target reference description texts. That is to say, the computer device can perform text fusion on each target reference description text belonging to the same reference detection report to obtain a training detection report, and the training image corresponding to this training detection report is the training image corresponding to this reference detection report. Then, based on this, it can be understood that the computer device can generate a training detection report for each training image through the above method, so that the computer device can construct a second training sample based on multiple training images and the training detection reports generated for each training image.
[0075] In the embodiments of the present application, the first training sample includes multiple training images and the reference detection reports of each training image. Each training image includes multiple analysis objects, and each reference detection report includes the reference description text of each analysis object. That is to say, each analysis object corresponds to multiple reference description texts, where the label of the reference description text is the first label or the second label. Since the computer device can determine the text retention rate of the reference description text with the first label from the multiple reference description texts of the analysis object based on the sample discard rate of each analysis object, and when the text retention rate is a non-zero value, use the reference description text with the first label as the target reference description text, which ensures that the number of reference description texts with the first label in the second training sample is less than the number of reference description texts with the first label in the first training sample. Moreover, by directly using the reference description text with the second label as the target reference description text, when the number of reference description texts with the second label is already small, the number of reference description texts with the second label is ensured to remain unchanged, so that the ratio of the number of reference description texts with the first label corresponding to each analysis object to the number of reference description texts with the second label can be reduced, reducing the difference in the number of positive and negative samples of each analysis object. In addition, since the sample discard rate of each analysis object in the present application is determined based on the same sample balance parameter, among all the target reference description texts obtained by the computer device based on the sample discard rate, the ratio of the number of positive and negative samples of each analysis object (that is, the ratio of the number of reference description texts with the first label corresponding to each analysis object to the number of reference description texts with the second label) is within the preset interval, so that the ratio of the number of positive and negative samples of each object in the second training sample is balanced. It can be understood that the optimized report generation model obtained by optimizing the report generation model with the second training sample can generate a relatively complete and accurate detection report.
[0076] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of a model optimization method provided by an embodiment of the present application. The method indicated by this flowchart can be executed by the above-mentioned computer device, or by other model training devices other than the above computer device. For the convenience of description, the following takes the computer device to optimize the report generation model as an example to Figure 5 detail the Figure 5 related steps. As
[0077] shown, the method includes steps S501 - S503:
[0078] Among them, the second training sample includes multiple training images, and there is at least one target reference description text in the training detection report of each training image. When the label of the target reference description text is the first label, the target reference description text corresponds to a text retention rate. In this case, during the process of the computer device obtaining the training detection report, it can also obtain the text retention rate corresponding to the target reference description text of the first label.
[0079] S502. Use the report generation model to perform text description processing on each training image to obtain the predicted detection report of each training image.
[0080] Among them, the structure of the report generation model can continue as Figure 4 shown in the structure marked by 402 in Figure 4 . Next, continue to combine
[0081] {v1, v2,..., vn} = f V (X) Equation 2
[0082] Among them, {v1, v2,..., vn} refers to the extracted visual features, n is the dimension of the visual features. fv represents the visual feature extractor, and X represents the training image. Then, f V(X) can be understood as: using a visual feature extractor to extract the visual features of training image X. It should be noted that the visual feature extractor can be any standard CNN (Convolutional Neural Network) structure, such as: VGG (Visual Geometry Group, deep convolutional neural network), DenseNet (Densely Connected Convolutional Networks), or ResNet (Residual Neural Network). Among them, it has been found through research that ResNet-101 trained on ImageNet (a large visual database for visual object recognition software research) is widely used and has better feature extraction effects. Therefore, in this application, ResNet-101 trained on ImageNet is preferentially used as the visual feature extractor.
[0083] Further, after the computer device extracts the visual features of the training image, the computer device can use the encoder in the report generation model to encode the visual features of the training image to obtain the intermediate image features corresponding to the corresponding training image. Furthermore, the computer device can use the decoder in the report generation model to decode the intermediate image features to obtain the predicted detection report of the corresponding training image. Exemplarily, the computer device can obtain the intermediate image features of the training image based on the method shown in Equation 3, and the computer device can obtain the predicted detection report of the corresponding training image based on the method shown in Equation 4.
[0084] {h1, h2,..., hn} = f E ({v1, v2,..., vn}) Equation 3
[0085] where {h1, h2,..., hn} are the intermediate image features; f E represents the encoder, and {v1, v2,..., vn} refers to the extracted visual features.
[0086] {y1, y2,..., ym} = f D ({h1, h2,..., hn}) Equation 4
[0087] where {y1, y2,..., ym} represents the predicted detection report, and y1, y2, and ym represent the predicted description texts of different analysis objects. It can be seen that the predicted detection report includes the predicted description texts of at least one analysis object, and in practical applications, one predicted description text corresponds to one target reference description text in the training detection report of each training image. f Ddenotes a decoder, and {h1, h2, …, hn} are intermediate image features, f D ({h1, h2, …, hn}) represents decoding the intermediate image features using the decoder. In practical applications, please refer to Figure 4 the structure marked by 402 in [reference]. Specifically, the decoder can be a multi-head decoder in a Transformer (a natural language processing model) structure. The "Head" in the decoder can be understood as a text description module. In this application, a text description module can be used to generate a predicted description text of an analysis object. That is to say, there can be multiple text description modules in the decoder of this application. Exemplarily, the number of text description modules in the decoder can be the number of analysis objects in the training images. It can be understood that in this case, after the computer device decodes the intermediate image features using the decoder in the report generation model, it can obtain the predicted description text of each analysis object, so that the computer device can generate a predicted detection report based on all the generated predicted description texts.
[0088] S503. Based on each target reference description text included in the training detection report of each training image, and the predicted description text corresponding to each target reference description text, optimize the report generation model to obtain an optimized report generation model.
[0089] In a specific application, the computer device can first obtain each target reference description text in any training detection report, and then the computer device can obtain the predicted description text corresponding to each target reference description text. Further, the computer device can determine the loss value corresponding to each target reference description text according to the difference information (such as cross-entropy) between each target reference description text and the predicted description text corresponding to the target reference description text, so as to obtain the target loss value of the report generation model based on the loss value of each target reference description text. Thereafter, the computer device can perform a model optimization on the report generation model in the direction of reducing the target loss value. It should be noted that the optimized report generation model in the embodiments of this application refers to a report generation model in which all model parameters converge. That is to say, in the embodiments of this application, the optimized report generation model can be: the computer device optimizes the report generation model according to each training image in the second training sample and the training detection report and predicted detection report of each training image.
[0090] In a feasible implementation manner, the way for the computer device to calculate the cross-entropy between the training detection report and the predicted detection report can be as shown in Equation 5:
[0091]
[0092] where LCE (R, Y) represents the cross-entropy, q represents the number of target reference description texts, q is a positive integer, represents the cross-entropy between the target reference description text and the predicted description text, l represents the number of text words included in the target reference description text, l is a positive integer, r ij represents the j-th text word in the target reference description text of the i-th analysis object, r ij log(y ij ) represents the cross-entropy corresponding to the j-th word.
[0093] In another feasible implementation, in order to further improve the effect when optimizing the report generation model, such as: accelerating the convergence rate of the report generation model, or, enhancing the sensitivity of the report generation model to the health status of each analysis object, this application combines the sample discard rate of each analysis object and the calculation method of the target loss value of the report generation model in the above implementation manner, and proposes the automatic balance loss function (Auto-balance loss, ABL) shown in Equation 6. The computer device can use this automatic balance loss function to calculate the target loss value of the report generation model.
[0094]
[0095] Among them, L ABL (R, Y) is the target loss value of the report generation model determined by the computer device using the automatic balance loss function, Ni represents the number of analysis objects in each training image; r i represents the reference description text of the i-th analysis object in the reference detection report. Since in the same training image, one analysis object corresponds to one reference description text, for the sake of convenience of description, the reference description text of the i-th analysis object is hereinafter referred to as: the i-th reference description text. In this case, max(δ(r i ), 1 - Rand(p i )) can be understood as the loss value coefficient corresponding to the i-th reference description text, and this loss value coefficient is related to the label of the i-th reference description text. Among them, when the label of the i-th reference description text is the second label, δ(r i ) = 1; when the label of the i-th reference description text is the first label, δ(r i ) = 0. At this time, the loss value of the reference description text will be determined by the value of 1 - Rand(p i ). Since 1 - Rand(p i ) can represent the text retention rate of this reference text, that is to say, when the label of the i-th reference description text is the first label, the computer device will use this reference description text as the target reference description text when the text retention rate of this reference description text is non-zero.
[0096] It should be noted that in this application, the probability of Rand(p i ) generating 1 is pi, and 0 < Rand(p i ) ≤ 1. Then, that is to say, there is a probability of pi that the reference description text is discarded. Then, it can be further understood that when 1 - Rand(p i ) is not equal to 0, the computer device will generate a target loss value using the loss value of the reference description text. Based on this method, it can be ensured that all the reference description texts of the second label are used to calculate the loss value, so that all the reference description texts of the second label are used to train the report generation model, improving the sensitivity of the report generation model when the analysis object is in a non-healthy state, thus improving the accuracy of the report generation model when generating the description text of the corresponding analysis object, and further improving the accuracy of the detection report generated by the report generation model.
[0097] Among them, based on the foregoing, when the computer device optimizes the model of the report generation model using the second training sample, it can adjust the model parameters of the report generation model once by using one training image and the training detection report of the training image at a time (that is: one model optimization). In practical applications, optimizing the model of the report generation model can be understood as adjusting the model parameters based on the learning rate of the report generation model. Then, in order to improve the optimization effect of the report generation model, the computer device can adjust the learning rate of the report generation model after adjusting the model parameters once, so that the report generation model can reach the convergence state at a suitable rate while ensuring a relatively high accuracy of the generated predicted description text. It can also be understood that the learning rate can be used to constrain the adjustment range and speed of the model parameters when the computer device adjusts the model parameters based on the target loss value. Exemplarily, in this application, the initial learning rate of the visual feature extractor in the report generation model can be set to 5x10 -5 , and the initial learning rate of the decoder can be set to 1x10 -4 . And, exemplarily, the computer device can gradually reduce the learning rates of the visual feature extractor and the decoder according to the poly strategy shown in Equation 7 (a learning rate adjustment strategy).
[0098]
[0099] Among them, lr represents the adjusted learning rate, base_lr is the baseline learning rate (i.e., the initial learning rate), epoch is the number of iterations (i.e., the number of times to optimize the model parameters of the report generation model); num_epoch is the maximum number of iterations; power controls the shape of the curve, which is used to control the shape of the learning rate change curve. If power < 1, the learning rate change curve bulges, indicating that the learning rate decreases slowly first and then quickly; correspondingly, if power > 1, the learning rate change curve is concave, indicating that the learning rate decreases quickly first and then slowly. Exemplarily, power = 0.8 in the embodiments of the present application.
[0100] In specific applications, the report generation model optimized based on the model optimization method proposed in the present application can generate a relatively complete detection report. Moreover, the optimized report generation model of the present application is more sensitive to the analysis object in an unhealthy state, and can obtain a more accurate description text of the analysis object, so that the detection report generated by the present application can assist the medical diagnosis of medical staff. To verify the effect of the present application, we take the example of generating a detection report of a chest X-ray using the report generation model to compare the performance of the optimized report generation model obtained in the present application with multiple report generation models in the prior art. Among them, during the comparison process, the computer device selects chest X-rays from two mainstream data sets (IU-Xray and MIMIC-CXR) respectively to generate detection reports, and then compares the performance of different report generation models based on the generated detection reports. Among them, when using each report generation model to generate a detection report of the chest X-ray in MIMIC-CXR (Thoracic X-ray Imaging Database), the comparison results between the present application and the prior art can be shown in Table 1:
[0101] Table 1
[0102]
[0103]
[0104] Among them, Methods represents the specific method adopted when generating the detection report (or understood as: the specific report generation model). Among them, the evaluation metrics include: (1) BLEU (Bilingual Evaluation Understudy), which is a machine translation task evaluation metric used to evaluate the similarity between the text generated by a machine and the human annotation. Specifically, the BLEU metric can include four types: BLEU-1, BLEU-2, BLEU-3, and BLEU-4. (2) METEOR (Metric for Evaluation of Translation with Explicit Ordering), which is a metric used to evaluate the matching degree between synonyms, roots, and affixes. (3) ROUGE (Recall-Oriented Understudy for Gisting Evaluation), which is a set of evaluation metrics designed for automatic summarization in machine translation and article summarization. It is a pure recall-based similarity measurement method, specifically achieved by comparing word sequences and word pairs among the overlapping N words, mainly examining the sufficiency and authenticity of image annotation. Based on Table 1, it can be seen that when generating the detection report for the detection images in MIMIC-CXR, the indicators of the optimized report generation model in the embodiments of the present application are superior to those of each report generation model in the prior art in all aspects. It can thus be seen that the optimized report generation model in the embodiments of the present application has better comprehensive performance.
[0105] When generating the detection report for the chest X-ray images in IU-Xray (X-ray dataset), the comparison results between the present application and the prior art can be shown in Table 2 as follows:
[0106] Table 2
[0107]
[0108] Similarly, based on Table 2, it can be seen that when generating the detection report for the detection images in IU-Xray, the optimized report generation model in the embodiments of the present application also has better performance. The detailed description of the analysis process in the present application is not elaborated.
[0109] In addition, the embodiments of the present application also compare the detection reports. Specifically, taking the method of the embodiments of the present application and the BaseLine method as an example for comparison, assuming that there are multiple analysis objects: analysis object A, analysis object B, and analysis object C, then the comparison results can be as Figure 6 shown. It can be seen that the detection report of the method proposed in the embodiments of the present application can generate a more complete and more detailed description text.
[0110] Based on the above description of Figure 5 it should be noted that in a specific embodiment, after the computer device obtains the optimized report generation model, the computer device can use the optimized report generation model to generate a detection report for any image to be detected. Specifically, the computer device can obtain the image to be detected, the image type of the image to be detected is the same as that of the training image, and the image to be detected can include multiple objects to be analyzed. Further, the computer device can perform feature extraction processing on the image to be detected to obtain the image features of the image to be detected. Then, the computer device can use the optimized report generation model to perform text description on each object to be analyzed to obtain the target description text of each object to be analyzed, so that the computer device can further perform text fusion processing on the target description text of each object to be analyzed to obtain the detection report of the image to be detected. It should be noted that the specific implementation principle of the computer device for generating a medical detection report can refer to the relevant description in step S502, which will not be elaborated in this application.
[0111] In the embodiment of the present application, the first training sample includes multiple training images and the reference detection report of each training image. Each training image includes multiple analysis objects, and each reference detection report includes the reference description text of each analysis object. The label of the reference description text is the first label or the second label. The computer device determines the target reference description text of each analysis object from each reference detection report based on the label of the reference description text and the sample discard rate of each analysis object, so that the computer device can construct a second training sample based on the multiple training images in the first training sample and the target reference description text of each analysis object. Since the reference description text of the first label is usually more than that of the second label, and the sample discard rate is mainly used to probabilistically discard the reference description text of the first label, it can be understood that the ratio of the number of reference description texts of the first label to the number of reference description texts of the second label in the target reference description text of each analysis object determined in the embodiment of the present application can be within a preset interval, so that the difference in the ratio of the number of positive and negative samples of each analysis object is small. In this case, after optimizing the report generation model with the second training sample, the optimized report generation model has similar sensitivity to each analysis object, so that the detection report generated by the optimized report generation model is more accurate.
[0112] Based on the above description of the processing method of the training sample, the embodiment of the present application also discloses a processing device for the training sample. The processing device for the training sample can be a computer program (including program code) running in the above-mentioned computer device. The processing device for the training sample can execute asFigure 2 , Figure 3 and Figure 5 the methods shown. Please refer to Figure 7 , the processing device for the training samples may at least include: an acquisition unit 701, a traversal unit 702, and a training sample determination unit 703.
[0113] The acquisition unit 701 is configured to acquire a first training sample, where the first training sample includes a plurality of training images and a reference detection report for each training image, each training image includes a plurality of analysis objects, and the reference detection report includes a reference description text for each analysis object;
[0114] The traversal unit 702 is configured to traverse the plurality of reference description texts of each analysis object among the plurality of analysis objects of the first training sample, and based on the sample discard rate of each analysis object and the label of each reference description text among the plurality of reference description texts of each analysis object, determine a target reference description text among the plurality of reference description texts, where the ratio of the number of reference description texts with the first label to the number of reference description texts with the second label in the target reference description text of each analysis object is within a preset interval;
[0115] The training sample determination unit 703 is configured to determine a second training sample based on the target reference description text of each analysis object, the reference detection report to which each target reference description text belongs, and the plurality of training images; the second training sample includes the plurality of training images and a training detection report for each training image, and the training detection report includes the target reference description text in the reference detection report corresponding to the training image, and the second training sample is used to optimize the model of the report generation model.
[0116] In one implementation manner, the traversal unit 702 may specifically further perform:
[0117] If the label of any reference description text is the first label, determine the text retention rate corresponding to the any reference description text based on the sample discard rate of each analysis object, and when the text retention rate is a non-zero value, use the any reference description text as the target reference description text;
[0118] If the label of any reference description text is the second label, directly determine the any reference description text as the target reference description text.
[0119] In another implementation manner, the training sample determination unit 702 may specifically be further configured to perform:
[0120] Generate the text reference retention rate corresponding to the any reference description text under the constraint of the sample discard rate by using a random value generation function;
[0121] Obtain the text retention rate constraint parameter value, and use the difference between the text retention rate constraint parameter value and the reference text retention rate as the text retention rate corresponding to any one of the reference description texts.
[0122] In yet another embodiment, the processing device for the training samples may further include a sample discard rate determination unit 704, and the sample discard rate determination unit 704 may be specifically configured to perform:
[0123] Determine the sample discard rate of each analysis object among the multiple analysis objects based on the sample balance parameter;
[0124] Among them, the determination method of the sample discard rate of any one analysis object includes:
[0125] Obtain the label of each reference description text among the multiple reference description texts of any one analysis object;
[0126] Among the multiple reference description texts of any one analysis object, determine the number of reference description texts with the first label and the number of reference description texts with the second label;
[0127] Determine the sample discard rate of any one analysis object according to the sample balance parameter, the number of reference description texts with the first label, and the number of reference description texts with the second label.
[0128] In yet another embodiment, the sample discard rate determination unit 704 may be specifically configured to perform:
[0129] Obtain the ratio between the number of reference description texts with the second label and the number of reference description texts with the first label, and perform a multiplication operation on the obtained ratio and the sample balance parameter to obtain a multiplication operation result;
[0130] Obtain the sample discard rate constraint parameter value, and use the difference between the discard rate constraint parameter and the multiplication operation result as the first candidate discard rate;
[0131] Obtain the second candidate discard rate, and use the largest candidate discard rate among the first candidate discard rate and the second candidate discard rate as the sample discard rate of any one analysis object.
[0132] In yet another embodiment, the training sample determination unit 703 may be specifically configured to perform:
[0133] Based on the reference detection report to which each target reference description text belongs, determine the target reference description texts belonging to any one reference detection report among the target reference description texts of the multiple analysis objects;
[0134] A training detection report for generating a training image corresponding to any one of the reference detection reports, the training detection report including a determined target reference description text;
[0135] Based on the multiple training images and the training detection reports respectively generated for each training image, construct the second training sample.
[0136] In another embodiment, the obtaining unit 701 may specifically be configured to perform:
[0137] Obtain the target detection report of each training image, the target detection report including the reference description text of at least one analysis object;
[0138] If the number of reference description texts in the target detection report is less than the number of analysis objects in each training image, perform text filling on the target detection report to obtain the reference detection report of each training image;
[0139] If the number of reference description texts in the target detection report is equal to the number of analysis objects in each training image, use the target detection report as the reference detection report of each training image.
[0140] In another embodiment, the obtaining unit 701 may also specifically be configured to perform:
[0141] Obtain all analysis objects corresponding to the target detection report;
[0142] Among the multiple analysis objects, determine target analysis objects other than the obtained analysis objects;
[0143] Obtain the reference description text of each target analysis object from the description text database, and fill the reference description text of each target analysis object into the target detection report to obtain the reference detection report of each training image.
[0144] In another embodiment, the target detection report includes multiple detection statements, and each detection statement includes the object identifier of at least one analysis object; when determining the number of reference description texts in the target detection report, the obtaining unit 701 may specifically be configured to perform:
[0145] Obtain the weight of each object identifier among all object identifiers included in each detection statement, and determine the target object identifier with the largest weight in each detection statement;
[0146] Generate a reference description text according to the detection statements with the same target object identifier, and obtain the number of generated reference description texts; among the generated reference description texts, the target object identifiers corresponding to different reference description texts are different.
[0147] In yet another embodiment, the processing device for the training samples further includes a model optimization unit 705, and the model optimization unit 705 may be specifically configured to perform:
[0148] Obtain each training image in the second training sample, and the training detection report of each training image;
[0149] Perform text description processing on each training image by using the report generation model to obtain a predicted detection report for each training image, where the predicted detection report includes predicted description texts of the at least one analysis object, and one predicted description text corresponds to one target reference description text in the training detection report of each training image;
[0150] Optimize the report generation model based on each target reference description text included in the training detection report of each training image and the predicted description text corresponding to each target reference description text to obtain an optimized report generation model.
[0151] In yet another embodiment, the processing device for the training samples further includes a report generation unit 706, and the report generation unit 706 may be specifically configured to perform:
[0152] Obtain an image to be detected, where the image to be detected includes a plurality of objects to be analyzed;
[0153] Perform feature extraction processing on the image to be detected to obtain image features of most of the images to be detected;
[0154] Use the optimized report generation model to perform text description on each object to be analyzed to obtain a target description text for each object to be analyzed;
[0155] Perform fusion processing on the target description texts of each object to be analyzed to obtain a detection report for the image to be detected.
[0156] According to a specific embodiment of the embodiments of the present application, Figure 2 、 Figure 3 And Figure 5 Each step involved in the method shown can be executed by each unit in the processing device for the training samples shown in Figure 7 For example, Figure 2 The step S201 shown can be executed by the acquisition unit 701 of the training sample processing device shown in Figure 7 ; the step S202 can be executed by the traversal unit 702 of the training sample processing device shown in Figure 7 ; the step S203 can be executed by Figure 7It is executed by the training sample determination unit 703 of the training sample processing device shown. Again, Figure 3 In the method shown, step S301 can be executed by Figure 7 the acquisition unit 701 of the training sample processing device shown; step S302 can be executed by Figure 7 the sample discard rate determination unit 704 of the training sample processing device shown; step S303 can be executed by Figure 7 the traversal unit 702 of the training sample processing device shown; step S304 can be executed by Figure 7 the training sample determination unit 703 of the training sample processing device shown. Again, Figure 5 In the method shown, steps S501 to S503 can all be executed by Figure 7 the model optimization unit 705 of the training sample processing device shown.
[0157] According to another embodiment of the present application, Figure 7 Each unit in the training sample processing device shown is divided based on logical functions. The above-mentioned each unit can be separately or all combined into one or several other units to form, or some of them can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. In other embodiments of the present application, the above-mentioned training sample processing device may also include other units. In actual applications, these functions can also be assisted by other units and can be realized by multiple units collaborating.
[0158] According to another embodiment of the present application, it can be achieved by running a computer program (including program code) capable of executing the steps involved in the methods shown in Figure 2 , Figure 3 or Figure 5 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct the training sample processing device shown in Figure 7 and to implement the training sample processing method of the embodiments of the present application. The computer program can be recorded on, for example, a computer storage medium, loaded into the above-mentioned general computing device through the computer storage medium, and run therein.
[0159] In an embodiment of the present application, the first training sample includes multiple training images and a reference detection report for each training image. Each training image includes multiple analysis objects, and each reference detection report includes a reference description text for each analysis object. The label of the reference description text is the first label or the second label. The processing device of the training sample can determine the target reference description text of each analysis object from each reference detection report based on the label of the reference description text and the sample discard rate of each analysis object, so that for each analysis object, the ratio of the number of reference description texts with the first label to the number of reference description texts with the second label in the target reference description text is within a preset interval, achieving the balance between the positive and negative sample quantity ratios of each analysis object. Therefore, after optimizing the report generation model using the second training sample, the generated detection report of the optimized report generation model is more accurate.
[0160] Based on the related descriptions of the above method embodiments and device embodiments, the embodiments of the present application also provide a computer device. Please refer to Figure 8 . The computer device includes at least a processor 801 and a computer storage medium 802, and the processor 801 and the computer storage medium 802 of the computer device can be connected through a bus or other means.
[0161] Among them, the aforementioned computer storage medium 802 is a memory device in the computer device, used to store programs and data. It can be understood that the computer storage medium 802 here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer storage medium 802 provides a storage space, and this storage space stores the operating system of the computer device. And in this storage space, there is also stored one or more computer programs suitable for being loaded and executed by the processor 801. These computer programs can be one or more program codes. It should be noted that the computer storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one storage medium located far from the aforementioned processor. The processor 801 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the computer device, and it is suitable for implementing one or more computer programs, specifically suitable for loading and executing one or more computer programs to implement the corresponding method flow or corresponding function.
[0162] In one embodiment, one or more computer programs stored in the computer storage medium 802 can be loaded and executed by the processor 801 to implement the above related Figure 2 , Figure 3 andFigure 5 The corresponding method steps in the illustrated method embodiments. In a specific implementation, one or more computer programs in the computer storage medium 802 can be loaded and executed by the processor 801 to perform the following steps:
[0163] Obtain a first training sample, where the first training sample includes a plurality of training images and a reference detection report for each training image. Each training image includes a plurality of analysis objects, and the reference detection report includes reference description texts for each analysis object;
[0164] Traverse the reference description texts of each analysis object among the plurality of analysis objects in the first training sample, and based on the sample discard rate of each analysis object and the labels of each reference description text among the plurality of reference description texts of each analysis object, determine a target reference description text among the plurality of reference description texts. The ratio of the number of reference description texts with the first label to the number of reference description texts with the second label in the target reference description text of each analysis object is within a preset interval; based on the target reference description text of each analysis object, the reference detection report to which each target reference description text belongs, and the plurality of training images, determine a second training sample;
[0165] The second training sample includes the plurality of training images and a training detection report for each training image. The training detection report includes the target reference description texts in the reference detection report corresponding to the training image. The second training sample is used to optimize the model of the report generation model.
[0166] In one implementation manner, the processor 801 can specifically be used to load and execute:
[0167] If the label of any reference description text is the first label, determine the text retention rate corresponding to the any reference description text based on the sample discard rate of each analysis object, and when the text retention rate is a non-zero value, use the any reference description text as the target reference description text;
[0168] If the label of any reference description text is the second label, directly determine the any reference description text as the target reference description text.
[0169] In yet another implementation manner, the processor 801 can specifically be used to load and execute:
[0170] Generate the text reference retention rate corresponding to the any reference description text under the constraint of the sample discard rate by using a random value generation function;
[0171] Obtain the text retention rate constraint parameter value, and use the difference between the text retention rate constraint parameter value and the text reference retention rate as the text retention rate corresponding to any reference description text.
[0172] In another embodiment, the processor 801 may be specifically configured to load and execute:
[0173] Determine the sample discard rate of each analysis object among the multiple analysis objects based on the sample balance parameter;
[0174] Among them, the determination method of the sample discard rate of any analysis object includes:
[0175] Obtain the label of each reference description text among the multiple reference description texts of any analysis object;
[0176] Among the multiple reference description texts of any analysis object, determine the number of reference description texts with the first label and the number of reference description texts with the second label;
[0177] According to the sample balance parameter, the number of reference description texts with the first label, and the number of reference description texts with the second label, determine the sample discard rate of any analysis object.
[0178] In another embodiment, the processor 801 may be specifically configured to load and execute:
[0179] Obtain the ratio of the number of reference description texts with the second label to the number of reference description texts with the first label, and perform a multiplication operation on the obtained ratio and the sample balance parameter to obtain a multiplication operation result;
[0180] Obtain the sample discard rate constraint parameter value, and use the difference between the discard rate constraint parameter and the multiplication operation result as the first candidate discard rate;
[0181] Obtain the second candidate discard rate, and use the largest candidate discard rate among the first candidate discard rate and the second candidate discard rate as the sample discard rate of any analysis object.
[0182] In another embodiment, the processor 801 may be specifically configured to load and execute:
[0183] Based on the reference detection report to which each target reference description text belongs, determine the target reference description texts belonging to any reference detection report among the target reference description texts of the multiple analysis objects;
[0184] Generate a training detection report for the training image corresponding to any reference detection report, where the training detection report includes the determined target reference description texts;
[0185] Construct the second training sample based on the multiple training images and the training detection reports respectively generated for each training image.
[0186] In yet another embodiment, the processor 801 may be specifically configured to load and execute:
[0187] Obtain the object detection report for each training image, where the object detection report includes reference description texts of at least one analysis object;
[0188] If the number of reference description texts in the object detection report is less than the number of analysis objects in each training image, perform text filling on the object detection report to obtain the reference detection report for each training image;
[0189] If the number of reference description texts in the object detection report is equal to the number of analysis objects in each training image, use the object detection report as the reference detection report for each training image.
[0190] In yet another embodiment, the processor 801 may be specifically configured to load and execute:
[0191] Obtain all the analysis objects corresponding to the object detection report;
[0192] Among the multiple analysis objects, determine the target analysis objects other than the obtained analysis objects;
[0193] Obtain the reference description texts of each target analysis object from the description text database, and fill the reference description texts of each target analysis object into the object detection report to obtain the reference detection report for each training image.
[0194] In yet another embodiment, the object detection report includes multiple detection statements, and each detection statement includes the object identifier of at least one analysis object; the processor 801 may be specifically configured to load and execute:
[0195] Obtain the weight of each object identifier among all the object identifiers included in each detection statement, and determine the target object identifier with the largest weight in each detection statement;
[0196] Generate reference description texts according to the detection statements with the same target object identifier, and obtain the number of the generated reference description texts; among the generated reference description texts, the target object identifiers corresponding to different reference description texts are different.
[0197] In yet another embodiment, the processor 801 may be specifically configured to load and execute:
[0198] Obtain each training image in the second training sample and the training detection report of each training image;
[0199] Use the report generation model to perform text description processing on each training image to obtain a predicted detection report for each training image. The predicted detection report includes predicted description texts of the at least one analysis object, and one predicted description text corresponds to one target reference description text in the training detection report of each training image;
[0200] Based on each target reference description text included in the training detection report of each training image and the predicted description text corresponding to each target reference description text, optimize the report generation model to obtain an optimized report generation model.
[0201] In another embodiment, the processor 801 may specifically be configured to load and execute:
[0202] Obtain an image to be detected, where the image to be detected includes a plurality of objects to be analyzed;
[0203] Perform feature extraction processing on the image to be detected to obtain image features of most of the images to be detected;
[0204] Use the optimized report generation model to perform text description on each object to be analyzed to obtain a target description text for each object to be analyzed;
[0205] Perform fusion processing on the target description texts of each object to be analyzed to obtain a detection report of the image to be detected.
[0206] In the embodiments of the present application, the first training sample includes a plurality of training images and a reference detection report for each training image. Each training image includes a plurality of analysis objects, and each reference detection report includes a reference description text for each analysis object. The label of the reference description text is the first label or the second label. The computer device can determine the target reference description text of each analysis object from each reference detection report based on the label of the reference description text and the sample discard rate of each analysis object, so that the ratio of the number of reference description texts with the first label to the number of reference description texts with the second label in the target reference description text of each analysis object is within a preset interval, achieving the balance between the positive and negative sample quantity ratios of each analysis object. Therefore, after optimizing the report generation model using the second training sample, the detection report generated by the optimized report generation model is more accurate.
[0207] The present application also provides a computer storage medium, in which one or more computer programs corresponding to the above-mentioned processing method of training samples are stored. When one or more processors load and execute the one or more computer programs, the description of the processing method of training samples in the embodiments can be implemented, which will not be elaborated here. The description of the beneficial effects of the same method will not be elaborated here. It can be understood that the computer program can be deployed on one or more devices capable of communicating with each other for execution.
[0208] It should be noted that according to one aspect of the present application, a computer program product or a computer program is also provided. The computer program product includes a computer program, and the computer program is stored in a computer storage medium. The processor in the computer device reads the computer program from the computer storage medium and then executes the computer program, so that the computer device can execute the above Figure 2 , Figure 3 and Figure 5 methods provided in various optional ways in the embodiments of the processing method of the training samples shown.
[0209] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer storage medium. When the computer program is executed, it can include the processes of the embodiments of the processing method of the training samples as described above. Among them, the computer storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0210] It can be understood that the above-disclosed is only a partial embodiment of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of the above embodiments and the equivalent changes made according to the claims of the present application still fall within the scope covered by the invention.
Claims
1. A method for processing training samples, characterized in that, Including: Obtain a first training sample, where the first training sample includes a plurality of training images and a reference detection report for each training image. Each training image includes a plurality of analysis objects, and the reference detection report includes reference description texts for each analysis object; Traverse the reference description texts of each analysis object among the plurality of analysis objects in the first training sample, and based on the sample discard rate of each analysis object and the labels of each reference description text among the plurality of reference description texts of each analysis object, determine a target reference description text among the plurality of reference description texts. The ratio of the number of reference description texts with the first label to the number of reference description texts with the second label in the target reference description text of each analysis object is within a preset range; Based on the target reference description text of each analysis object, the reference detection report to which each target reference description text belongs, and the plurality of training images, determine a second training sample; The second training sample includes the plurality of training images and a training detection report for each training image. The training detection report includes the target reference description text in the reference detection report corresponding to the training image. The second training sample is used to optimize the model of the report generation model.
2. The method according to claim 1, wherein The determining a target reference description text among the plurality of reference description texts based on the sample discard rate of each analysis object and the labels of each reference description text among the plurality of reference description texts of each analysis object includes: If the label of any reference description text is the first label, determine the text retention rate corresponding to the any reference description text based on the sample discard rate of each analysis object, and when the text retention rate is a non-zero value, use the any reference description text as the target reference description text; If the label of any reference description text is the second label, directly determine the any reference description text as the target reference description text.
3. The method according to claim 2, wherein The determining the text retention rate corresponding to the any reference description text based on the sample discard rate of each analysis object includes: Use a random value generation function to generate a text reference retention rate corresponding to the any reference description text under the constraint of the sample discard rate; Obtain a text retention rate constraint parameter value, and use the difference between the text retention rate constraint parameter value and the text reference retention rate as the text retention rate corresponding to the any reference description text.
4. The method according to any one of claims 1 to 3, characterized in that, Before determining the target reference description text of each analysis object, the method further includes: Determine the sample discard rate of each analysis object among the plurality of analysis objects based on a sample balance parameter; Wherein, the determination method of the sample discard rate of any analysis object includes: Obtain the labels of each reference description text among the plurality of reference description texts of the any analysis object; Among the plurality of reference description texts of the any analysis object, determine the number of reference description texts with the first label and the number of reference description texts with the second label; According to the sample balance parameter, the number of reference description texts with the first label, and the number of reference description texts with the second label, determine the sample discard rate of the any analysis object.
5. The method according to claim 4, characterized in that, Determining the sample discard rate of any one of the analysis objects according to the sample balance parameter, the number of reference description texts of the first label, and the number of reference description texts of the second label includes: Obtaining the ratio between the number of reference description texts of the second label and the number of reference description texts of the first label, and performing a multiplication operation on the obtained ratio and the sample balance parameter to obtain a multiplication operation result; Obtaining the value of the sample discard rate constraint parameter, and taking the difference between the discard rate constraint parameter and the multiplication operation result as the first candidate discard rate; Obtaining a second candidate discard rate, and taking the largest candidate discard rate among the first candidate discard rate and the second candidate discard rate as the sample discard rate of any one of the analysis objects.
6. The method according to claim 1, wherein Determining the second training sample based on the target reference description text of each analysis object, the reference detection report to which each target reference description text belongs, and the multiple training images includes: Based on the reference detection report to which each target reference description text belongs, determining the target reference description texts belonging to any one reference detection report among the target reference description texts of the multiple analysis objects; Generating a training detection report for the training image corresponding to any one reference detection report, where the training detection report includes the determined target reference description texts; Based on the multiple training images and the training detection reports generated corresponding to each training image, constructing the second training sample.
7. The method according to claim 1 or 6, characterized in that, The method for obtaining the reference detection report of each training image in the first training sample includes: Obtaining the target detection report of each training image, where the target detection report includes the reference description texts of at least one analysis object; If the number of reference description texts in the target detection report is less than the number of analysis objects in each training image, filling the target detection report with text to obtain the reference detection report of each training image; If the number of reference description texts in the target detection report is equal to the number of analysis objects in each training image, taking the target detection report as the reference detection report of each training image.
8. The method according to claim 7, wherein The filling the target detection report with text to obtain the reference detection report of each training image includes: Obtaining all the analysis objects corresponding to the target detection report; Among the multiple analysis objects, determining the target analysis objects other than the obtained analysis objects; Obtaining the reference description texts of each target analysis object from the description text database, and filling the reference description texts of each target analysis object into the target detection report to obtain the reference detection report of each training image.
9. The method according to claim 7, wherein The target detection report includes multiple detection statements, and each detection statement includes the object identifier of at least one analysis object; the method for determining the number of reference description texts in the target detection report includes: Obtaining the weight of each object identifier among all the object identifiers included in each detection statement, and determining the target object identifier with the largest weight in each detection statement; Generate reference description texts based on detection statements with the same target object identifier, and obtain the number of generated reference description texts; among the generated reference description texts, the target object identifiers corresponding to different reference description texts are different.
10. The method according to claim 1, wherein The method further includes: Obtain each training image in the second training sample, and the training detection report of each training image; Perform text description processing on each training image using the report generation model to obtain a predicted detection report for each training image, where the predicted detection report includes predicted description texts of at least one analysis object, and one predicted description text corresponds to one target reference description text in the training detection report of each training image; Based on each target reference description text included in the training detection report of each training image, and the predicted description text corresponding to each target reference description text, optimize the report generation model to obtain an optimized report generation model.
11. The method according to claim 10, wherein The method further includes: Obtain an image to be detected, where the image to be detected includes multiple objects to be analyzed; Perform feature extraction processing on the image to be detected to obtain image features of most of the images to be detected; Use the optimized report generation model to perform text description on each object to be analyzed to obtain a target description text for each object to be analyzed; Perform fusion processing on the target description texts of each object to be analyzed to obtain a detection report for the image to be detected.
12. A processing device for training samples, characterized in that, Includes: An acquisition unit for acquiring a first training sample, where the first training sample includes multiple training images and a reference detection report for each training image, each training image includes multiple analysis objects, and the reference detection report includes a reference description text for each analysis object; A traversal unit for traversing multiple reference description texts of each analysis object among multiple analysis objects in the first training sample, and determining a target reference description text among the multiple reference description texts based on the sample discard rate of each analysis object and the label of each reference description text among the multiple reference description texts of each analysis object, where the ratio of the number of reference description texts with the first label to the number of reference description texts with the second label in the target reference description text of each analysis object is within a preset interval; A training sample determination unit for determining a second training sample based on the target reference description text of each analysis object, the reference detection report to which each target reference description text belongs, and the multiple training images; The second training sample includes the multiple training images and a training detection report for each training image, and the training detection report includes the target reference description text in the reference detection report corresponding to the training image, and the second training sample is used to optimize the report generation model.
13. A computer device, characterized in that, Includes: A processor, where the processor is adapted to implement one or more computer programs; A computer storage medium storing one or more computer programs, where the one or more computer programs are adapted to be loaded and executed by the processor to perform the processing method of the training sample according to any one of claims 1-11.
14. A computer storage medium, characterized in that, The computer storage medium stores one or more computer programs, and the one or more computer programs are adapted to be loaded and executed by a processor to perform the processing method of the training sample according to any one of claims 1-11.
15. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program is adapted to be loaded and executed by a processor to perform the processing method of the training sample according to any one of claims 1-11.
Citation Information
Patent Citations
Model training method and device and storage medium
CN110163234A
System and method for inputting images or labels into electronic devices
US20160292148A1