Subject analysis method and apparatus, computer device and nonvolatile storage medium
Through the multimodal dialogue language model, feature extraction and data processing of coronary heart disease detection images are generated and analyzed, and the inefficiency problem caused by the single coronary heart disease detection results in the prior art is solved, and more efficient diagnosis and treatment results output is achieved.
Patent Information
- Application Number
- PCT/CN2024/094963
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2024-05-23
- Publication Date
- 2025-08-07
AI Technical Summary
In the prior art, the detection of coronary heart disease mainly focuses on classification tasks, resulting in a single output result, and then analyzes the single test result manually, resulting in a low efficiency in determining the diagnosis and treatment results of coronary heart disease.
A multimodal dialogue language model is adopted, including the training visual structure, attention structure and language structure, feature extraction and data processing of the image to be detected, initial diagnostic results are generated, and the initial diagnostic results are further processed through the language structure to generate object analysis results.
The efficiency of the diagnosis and treatment results of coronary heart disease is improved. Through the application of the multimodal dialogue language model, the index information contained in the image to be detected can be reflected in a variety of preset indicators and the treatment plan corresponding to the disease information can be output.
Smart Images

Figure CN2024094963_07082025_PF_FP_ABST
Abstract
Description
Object analysis method, device, computer equipment and non-volatile storage medium
[0001] Related applications
[0002] This application claims priority to Chinese patent application number 202410129641.0, filed on January 30, 2024, entitled “Object Analysis Method, Apparatus, Computer Equipment and Storage Medium,” the entire text of which is incorporated herein by reference. Technical Field
[0003] The present application relates to the field of artificial intelligence technology, and in particular to an object analysis method, apparatus, computer equipment, non-volatile storage medium, and computer program product. Background Art
[0004] With the development of artificial intelligence technology, the use of artificial intelligence technology to screen and analyze diseases can reduce unnecessary invasive examinations and improve the efficiency and accuracy of disease diagnosis.
[0005] In related technologies, a multimodal neural network model is used to process ECG signals and medical records to obtain test results, including health, coronary heart disease, and arrhythmia. Specifically, the ECG signals and medical records are first preprocessed, and then a multimodal feature extraction module is used to extract and fuse features from the ECG signals, time-frequency images, vector cardiograms, and medical record text to obtain fused features. Finally, the fused features are predicted based on the classification module to obtain the test results. Doctors then analyze the test results and medical records to obtain the diagnosis and treatment results.
[0006] However, in related technologies, the detection technology for coronary heart disease mainly focuses on classification tasks, resulting in a single output result. The single test result is then analyzed manually to determine the diagnosis and treatment results, resulting in low efficiency in determining the diagnosis and treatment results of coronary heart disease.
[0007] Summary of the Invention
[0008] Based on this, it is necessary to provide an object analysis method, apparatus, computer equipment, computer-readable storage medium and computer program product to address the above technical issues.
[0009] In a first aspect, the present application provides an object analysis method, comprising the following steps:
[0010] Obtaining an image to be detected corresponding to the target object and a trained multimodal dialogue language model, wherein the trained multimodal dialogue language model includes a trained visual structure, a trained attention structure, and a trained language structure;
[0011] Extracting features of the image to be detected based on the trained visual structure to obtain image features corresponding to the image to be detected;
[0012] Performing data processing on the image features based on the trained attention structure to obtain an initial diagnosis result corresponding to the target object;
[0013] The initial diagnosis result is subjected to data processing based on the trained language structure to obtain an object analysis result corresponding to the target object.
[0014] In one embodiment, the data processing of the initial diagnosis result based on the trained language structure to obtain the object analysis result corresponding to the target object includes: data processing of the initial diagnosis result according to the trained language structure to obtain a medical order text; and constructing the object analysis result corresponding to the target object based on the initial diagnosis result and the medical order text.
[0015] In one embodiment, before obtaining the image to be detected corresponding to the target object and the trained multimodal dialogue language model, the method further includes:
[0016] Acquire an image sample data set, wherein the image sample data set includes image sample data, reference description text corresponding to the image sample data, and reference analysis results corresponding to the type of the image sample data;
[0017] Training a visual model based on the image sample dataset and a momentum contrast learning framework to obtain a pre-trained visual model;
[0018] Training the visual structure of the multimodal dialogue language model based on the output of the pre-trained visual model in a knowledge distillation manner to obtain the trained visual structure;
[0019] Constructing a multi-branch attention structure according to preset indicators, and training the attention structure based on the image sample data and the reference description text corresponding to the image sample data to obtain the trained attention structure;
[0020] Based on the reference description text corresponding to the image sample data and the reference analysis result corresponding to the type of the image sample data, the language structure of the multimodal dialogue language model is trained to obtain the trained language structure.
[0021] In one embodiment, the acquiring of the image sample data set includes: acquiring initial image data; and screening the initial image data according to the clarity of the initial image data to obtain image sample data.
[0022] The step of training the visual model based on the image sample dataset and the momentum contrast learning framework to obtain the pre-trained visual model includes:
[0023] Based on the momentum contrastive learning framework and the image sample data, generating positive sample pairs and negative sample pairs;
[0024] The visual model is trained according to the first loss function, the positive sample pairs, and the negative sample pairs to obtain the pre-trained visual model.
[0025] In one embodiment, the first loss function is a normalized mutual information loss function, and the pre-trained visual model includes a query encoder and a target encoder.
[0026] The pre-trained visual model is trained according to the first loss function, the positive sample pairs, and the negative sample pairs to obtain the pre-trained visual model, including:
[0027] Performing data processing on the positive sample pairs according to the query encoder and the target encoder respectively to obtain a first feature vector group corresponding to the query encoder and a second feature vector group corresponding to the target encoder;
[0028] performing data processing on the negative sample pair according to the query encoder to obtain a third feature vector group corresponding to the negative sample pair;
[0029] The visual model is trained based on the normalized mutual information loss function, the first eigenvector group, the second eigenvector group, and the third eigenvector group to obtain the pre-trained visual model.
[0030] In one embodiment, the visual structure includes a backbone network and a first bypass.
[0031] The method of knowledge distillation is used to train the visual structure of the multimodal dialogue language model based on the output of the pre-trained visual model to obtain a trained visual structure, including:
[0032] Training the visual structure of the multimodal dialogue language model based on a preset training method, and during the training process, using the backbone network and the first bypass as student models and the pre-trained visual model as a teacher model;
[0033] Measuring the first output result of the teacher model and the second output result of the student model according to the target measurement function to obtain the similarity between the teacher model and the student model;
[0034] The backbone network and the first bypass corresponding to the student model are trained using a second loss function and the similarity to obtain a trained visual structure.
[0035] In one embodiment, the attention structure includes a plurality of second bypass paths;
[0036] The method of constructing a multi-branch attention structure according to preset indicators, and training the attention structure based on the image sample data and the reference description text corresponding to the image sample data to obtain a trained attention structure includes:
[0037] Classifying the image sample data based on preset indicators to obtain multiple types of target image sample sets, reference description text corresponding to each target image sample in the target image sample sets, and analysis results corresponding to each target sample in each type of the target image sample sets;
[0038] For each of the second bypass paths, determining a target type corresponding to each of the second bypass paths, and performing data processing on the target image sample set and the reference description text of the target type according to the second bypass paths to obtain a visual vector and a first semantic code, respectively;
[0039] The attention structure is trained based on the contrast loss of the visual vector and the first semantic code to obtain a trained attention structure.
[0040] In one embodiment, the training of the language structure of the multimodal dialogue language model based on the reference description text corresponding to the image sample data and the reference analysis result corresponding to the type of the image sample data to obtain the trained language structure includes:
[0041] Performing data processing on the target image sample according to the trained attention structure to obtain an initial diagnosis result;
[0042] For each type of the target image sample set, performing data processing on the initial diagnosis results corresponding to the target image samples according to the language structure to obtain a second semantic code;
[0043] Based on the contrast loss between the second semantic encoding and the third semantic encoding corresponding to the reference description text, the language structure of the multimodal dialogue language model is trained to obtain the trained language structure.
[0044] In one embodiment, the initial diagnosis result includes at least one of a plurality of preset indicators, wherein the plurality of preset indicators include relevant information describing the disease.
[0045] In one embodiment, extracting features from the image to be detected based on the trained visual structure to obtain image features corresponding to the image to be detected includes: inputting the image to be detected corresponding to the target object and the task type corresponding to the prompt word information into the trained multimodal dialogue language model; and extracting features from the image to be detected through a forward propagation process based on the trained visual structure to obtain image features corresponding to the image to be detected.
[0046] The method further comprises: representing the image to be detected as a feature vector containing the image features.
[0047] In one embodiment, constructing the object analysis result corresponding to the target object based on the initial diagnosis result and the medical order text includes: associating the initial diagnosis result with the medical order text; constructing the initial diagnosis result and the medical order text following the initial diagnosis result into a coherent description based on the trained language structure to obtain the object analysis result.
[0048] In one embodiment, the positive sample pairs are created by applying data augmentation to original images in the image sample dataset, and the negative sample pairs are created by randomly selecting images from the image sample dataset.
[0049] In a second aspect, the present application also provides an object analysis device, comprising: a first acquisition module, a feature extraction module, an initial diagnosis module, and a diagnosis analysis module.
[0050] The first acquisition module is used to obtain an image to be detected corresponding to the target object and a trained multimodal dialogue language model, wherein the trained multimodal dialogue language model includes a trained visual structure, a trained attention structure, and a trained language structure.
[0051] The feature extraction module is used to extract features of the image to be detected based on the trained visual structure to obtain image features corresponding to the image to be detected.
[0052] An initial diagnosis module is used to perform data processing on the image features based on the trained attention structure to obtain an initial diagnosis result corresponding to the target object.
[0053] The diagnosis analysis module is used to perform data processing on the initial diagnosis result based on the trained language structure to obtain an object analysis result corresponding to the target object.
[0054] In one embodiment, the diagnosis analysis module is specifically used to perform data processing on the initial diagnosis result according to the language structure to obtain a medical order text;
[0055] An object analysis result corresponding to the target object is constructed based on the initial diagnosis result and the medical order text.
[0056] In one embodiment, the apparatus further includes: a second acquisition module, a first training module, a second training module, a third training module, and a fourth training module.
[0057] The second acquisition module is used to acquire an image sample data set, wherein the image sample data set includes image sample data, reference description text corresponding to each image sample data, and reference analysis results corresponding to the type of the image sample data.
[0058] The first training module is used to train the visual model based on the image sample dataset and the momentum contrast learning framework to obtain a pre-trained visual model.
[0059] The second training module is used to train the visual structure of the multimodal dialogue language model based on the output of the pre-trained visual model according to the knowledge distillation method to obtain a trained visual structure.
[0060] The third training module is used to construct an attention structure with a multi-branch architecture according to preset indicators, and train the attention structure based on image sample data and reference description text corresponding to the image sample data to obtain a trained attention structure.
[0061] The fourth training module is used to train the language structure of the multimodal dialogue language model based on the reference description text corresponding to the image sample data and the reference analysis results corresponding to the type of the image sample data to obtain a trained language structure.
[0062] In one embodiment, the second acquisition module is specifically configured to acquire initial image data; and filter the initial image data according to the clarity of the initial image data to obtain image sample data.
[0063] The first training module is specifically used to generate positive sample pairs and negative sample pairs based on the momentum contrast learning framework and each image sample data; train the visual model according to the first loss function, the positive sample pairs and the negative sample pairs to obtain a pre-trained visual model.
[0064] In one embodiment, the first loss function is a normalized mutual information loss function, and the pre-trained visual model includes a query encoder and a target encoder.
[0065] The first training module is specifically used to process the positive sample pairs according to the query encoder and the target encoder respectively to obtain the first feature vector group corresponding to the query encoder and the second feature vector group corresponding to the target encoder; process the negative sample pairs according to the query encoder to obtain the third feature vector group corresponding to the negative sample pairs; train the visual model based on the normalized mutual information loss function, the first feature vector group, the second feature vector group and the third feature vector group to obtain a pre-trained visual model.
[0066] In one embodiment, the visual structure includes a backbone network and a first bypass.
[0067] The second training module is specifically used to train the visual structure of the multimodal dialogue language model based on a preset training method. During the training process, the backbone network and the first bypass are used as student models, and the pre-trained visual model is used as the teacher model; the first output result of the teacher model and the second output result of the student model are measured according to the target measurement function to obtain the similarity between the teacher model and the student model; the backbone network and the first bypass corresponding to the student model are trained through the second loss function and similarity to obtain the trained visual structure.
[0068] In one embodiment, the attention structure includes a plurality of second bypass paths.
[0069] The third training module is specifically used to classify image sample data based on preset indicators to obtain multiple types of target image sample sets, reference description text corresponding to each target image sample in the target image sample set, and analysis results corresponding to each target sample in each type of target image sample set; for each second bypass, determine the target type corresponding to each second bypass, and perform data processing on the target image sample set and the reference description text of the target type according to the second bypass to obtain a visual vector and a first semantic code, respectively; train the attention structure based on the contrast loss of the visual vector and the first semantic code to obtain a trained attention structure.
[0070] In one embodiment, the fourth training module is specifically used to perform data processing on the target image samples according to the trained attention structure to obtain an initial diagnosis result; for each type of target image sample set, the initial diagnosis result corresponding to the target image sample is data processed according to the language structure to obtain a second semantic code; the language structure of the multimodal dialogue language model is trained based on the contrast loss of the second semantic code and the third semantic code corresponding to the reference description text to obtain the language structure of the trained multimodal dialogue language model.
[0071] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0072] Obtaining an image to be detected corresponding to the target object and a trained multimodal dialogue language model, wherein the trained multimodal dialogue language model includes a trained visual structure, a trained attention structure, and a trained language structure;
[0073] Extracting features of the image to be detected based on the trained visual structure to obtain image features corresponding to the image to be detected;
[0074] Performing data processing on the image features based on the trained attention structure to obtain an initial diagnosis result corresponding to the target object;
[0075] The initial diagnosis result is subjected to data processing based on the trained language structure to obtain an object analysis result corresponding to the target object.
[0076] In a fourth aspect, the present application further provides a non-volatile computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0077] Obtaining an image to be detected corresponding to the target object and a trained multimodal dialogue language model, wherein the trained multimodal dialogue language model includes a trained visual structure, a trained attention structure, and a trained language structure;
[0078] Performing feature extraction on the image to be detected according to the trained visual structure to obtain image features corresponding to the image to be detected;
[0079] Performing data processing on the image features based on the trained attention structure to obtain an initial diagnosis result corresponding to the target object;
[0080] The initial diagnosis result is subjected to data processing based on the trained language structure to obtain an object analysis result corresponding to the target object.
[0081] In a fifth aspect, the present application further provides a computer program product, comprising computer-executable instructions, which, when executed by a processor, implement the following steps:
[0082] Obtaining an image to be detected corresponding to the target object and a trained multimodal dialogue language model, wherein the trained multimodal dialogue language model includes a trained visual structure, a trained attention structure, and a trained language structure;
[0083] Extracting features of the image to be detected based on the trained visual structure to obtain image features corresponding to the image to be detected;
[0084] Performing data processing on the image features based on the trained attention structure to obtain an initial diagnosis result corresponding to the target object;
[0085] The initial diagnosis result is subjected to data processing based on the trained language structure to obtain an object analysis result corresponding to the target object.
[0086] The object analysis method, apparatus, computer device, storage medium, and computer program product described above, in response to a prompt word input by a user, obtain an image to be detected and a trained multimodal dialogue language model corresponding to the target object, wherein the trained multimodal dialogue language model includes a trained visual structure, a trained attention structure, and a trained language structure; perform feature extraction on the image to be detected based on the trained visual structure to obtain image features corresponding to the image to be detected; perform data processing on the image features based on the trained attention structure to obtain an initial diagnosis result corresponding to the target object; and perform data processing on the initial diagnosis result based on the trained language structure to obtain an object analysis result corresponding to the target object. Using this method, by performing feature extraction and data processing based on the trained visual structure and attention structure, an initial diagnosis result for the target object can be obtained based on the image to be detected. The initial diagnosis result can reflect indicator information contained in the image to be detected in a variety of preset indicators. The initial diagnosis result is analyzed through the language structure to output a treatment plan corresponding to the disease information, thereby improving the efficiency of providing coronary heart disease diagnosis and treatment results. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0088] FIG1 is a diagram illustrating an application environment of an object analysis method according to an embodiment of the present application;
[0089] FIG2 is a flow chart of an object analysis method according to an embodiment of the present application;
[0090] FIG3 is a flow chart of an example of an object analysis method in one embodiment of the present application;
[0091] FIG4 is a schematic diagram of a process for constructing object analysis results in one embodiment of the present application;
[0092] FIG5 is a flowchart illustrating a training process of a multimodal conversational language model in one embodiment of the present application;
[0093] FIG6 is a schematic diagram of an example of a training process of a multimodal conversational language model according to an embodiment of the present application;
[0094] FIG7 is a flow chart of a visual structure training process according to an embodiment of the present application;
[0095] FIG8 is a schematic diagram of an example of a visual structure training process in one embodiment of the present application;
[0096] FIG9 is a schematic diagram of a process for performing loss training on visual structure in one embodiment of the present application;
[0097] FIG10 is a schematic diagram of a process for training visual structure in one embodiment of the present application;
[0098] FIG11 is a schematic diagram showing the principle of LoRA fine-tuning training in one embodiment of the present application;
[0099] FIG12 is a flow chart of the training process of the attention structure in one embodiment of the present application;
[0100] FIG13 is a flow chart of a language structure training process according to an embodiment of the present application;
[0101] FIG14 is a flow chart illustrating an example of a training process for a language structure according to an embodiment of the present application;
[0102] FIG15 is a structural block diagram of an object analysis device according to an embodiment of the present application;
[0103] FIG16 is a diagram showing the internal structure of a computer device in one embodiment of the present application. DETAILED DESCRIPTION
[0104] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0105] The object analysis method provided in the embodiments of the present application can be applied in the application environment shown in FIG1 . Terminal 102 communicates with server 104 via a communication network. A data storage system can store data that server 104 needs to process. The data storage system can be integrated with server 104, or located on a cloud or other network server. Terminal 102 sends an image to be detected corresponding to a target object to server 104. Server 104 obtains the image to be detected corresponding to the target object and a trained multimodal conversational language model. The multimodal conversational language model includes a visual structure, an attention structure, and a language structure. Server 104 extracts features from the image to be detected based on the trained visual structure to obtain image features corresponding to the image to be detected. Server 104 processes the image features based on the attention structure to obtain an initial diagnosis result corresponding to the target object. The initial diagnosis result includes at least one of a plurality of preset indicators. Server 104 processes the initial diagnosis result based on the language structure to obtain an object analysis result corresponding to the target object. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, and tablet computers. Server 104 can be implemented as a standalone server or a server cluster consisting of multiple servers.
[0106] In an exemplary embodiment, as shown in FIG2 , the present application provides an object analysis method, which is described by taking the application of the method to the server 104 in FIG1 as an example, and includes the following steps 202 to 208 .
[0107] Step 202: Obtain the image to be detected corresponding to the target object and the trained multimodal dialogue language model.
[0108] Among them, the trained multimodal dialogue language model includes a trained visual structure, a trained attention structure and a trained language structure.
[0109] In an embodiment of the present application, the target object may be a patient, and the image to be detected may be a medical image, such as an ultrasound image, a nuclear magnetic resonance imaging (MRI), etc. The multimodal conversational language model is used to determine the lesion information or symptom information contained in the patient's image to be detected, i.e., the medical image, based on the visual structure and attention structure, and to determine a targeted medical advice for the patient based on the lesion information or symptom information. This embodiment of the present application and the following embodiments are described as follows, using the image to be detected as a fundus image of a patient and the multimodal conversational language model as a large language model for data analysis of coronary heart disease.
[0110] The user inputs the patient's image to be tested into the trained multimodal dialogue language model, and at the same time enters prompt information or command information, for example, "Hello, please give the possible vascular calcification situation of this patient based on this fundus image." The server responds to the upload operation of the user terminal and receives the image to be tested uploaded by the user.
[0111] Step 204 : extract features of the image to be detected based on the trained visual structure to obtain image features corresponding to the image to be detected.
[0112] In an embodiment of the present application, after the server receives the image to be detected and the prompt word information of the target object, the task type corresponding to the image to be detected and the prompt word information is input into the trained multimodal dialogue language model. According to the trained visual structure in the trained multimodal dialogue language model, through the forward propagation process, the feature extraction of the image to be detected is performed, and the image to be detected is represented as a feature vector containing image features, which is used to analyze the image to be detected and determine the possible lesion information in the image to be detected. For example, the image information contains the edge, texture, morphology, etc. of the lesion, which can be used as a basis for judging whether there is a lesion in the image to be detected and the condition of the lesion.
[0113] Step 206: Perform data processing on the image features based on the trained attention structure to obtain an initial diagnosis result corresponding to the target object.
[0114] The initial diagnosis result includes at least one of a plurality of preset indicators.
[0115] In an embodiment of the present application, the preset indicators include relevant information describing the disease, which can be relevant information reflecting a certain type of disease in the field of medical diagnosis. Taking the coronary angiography diagnosis corresponding to coronary heart disease examination as an example, the preset indicators include coronary artery stenosis, coronary artery calcification score, degree of coronary atherosclerosis, and whether myocardial infarction exists. Among them, each indicator is a prediction object of the fundus, and the attention structure is a multi-task learning structure with a multi-branch architecture. The server processes the feature vector containing the image features through the attention structure. According to the relationship between the visual vector of the image learned during the training process and the semantic encoding of the descriptive text, the feature vector of the image features is extracted by each branch of the multi-task learning structure, and the targeted features corresponding to each branch are obtained. The initial diagnosis result of the target object is determined based on the targeted features.
[0116] Step 208 : performing data processing on the initial diagnosis result based on the trained language structure to obtain an object analysis result corresponding to the target object.
[0117] In an embodiment of the present application, the server passes the initial diagnosis result backward and passes the initial diagnosis result of the target object to the language structure. The language structure is a large language model, which is obtained by fine-tuning and training through coronary angiography diagnosis conclusions and corresponding medical orders to obtain corresponding diagnosis and treatment measures based on the predicted initial diagnosis result.
[0118] In this embodiment, based on the image to be detected, feature extraction is performed through the trained visual structure, and data processing of the image features is performed through the trained attention structure, so that the initial diagnosis result of the target object can be obtained. The initial diagnosis result can reflect the indicator information contained in the image to be detected in a variety of preset indicators, and the initial diagnosis result is processed and analyzed through the language structure, and the treatment plan corresponding to the disease information is output, so that the efficiency of providing coronary heart disease diagnosis and treatment results is improved.
[0119] In an exemplary embodiment, as shown in FIG3 , the server obtains a fundus image of the target subject and extracts features from the fundus image using the trained visual structure to obtain a feature representation of the key information contained in the fundus image. Then, based on the trained attention structure, the server receives prompt word information from the user, such as “analyzing coronary heart disease based on the fundus.” Using the pre-trained weights in the trained attention structure, the feature representation of the fundus image is further expressed using model parameters with different attentional targets to obtain fundus features. These fundus features can reflect the initial diagnostic results of the target subject, i.e., the conditions present in the target subject. Finally, the fundus features are input into the language structure, and a diagnostic analysis is generated based on the initial diagnostic results represented by the fundus features.
[0120] In an exemplary embodiment, as shown in FIG4 , step 208 performs data processing on the initial diagnosis result based on the trained language structure to obtain the object analysis result corresponding to the target object, including steps 402 to 404 .
[0121] Step 402: Process the initial diagnosis result according to the trained language structure to obtain a medical order text.
[0122] In an embodiment of the present application, the server analyzes the descriptive text of each preset indicator contained in the initial diagnosis result according to the language structure, and generates a medical order text corresponding to the descriptive text of each preset indicator, and the medical order text contains a diagnosis and treatment plan for each indicator.
[0123] Step 404: construct an object analysis result corresponding to the target object based on the initial diagnosis result and the medical order text.
[0124] In this embodiment of the present application, the server associates the initial diagnosis result with the medical order text. Based on the trained language structure, the server constructs the initial diagnosis result and the medical order text that follows the initial diagnosis result into a coherent description, thereby obtaining the final object analysis result. For example, the language structure simultaneously outputs descriptions of multiple lesions such as coronary artery stenosis, calcification, atherosclerosis, and myocardial infarction, and associates the description of each lesion with a diagnosis and treatment recommendation, outputting an object analysis result consisting of the lesion description and the diagnosis and treatment recommendation.
[0125] In this embodiment, by constructing the object analysis results with the initial diagnosis results and the medical order text, the user can clearly understand the extent of the disease and output targeted diagnosis and treatment recommendations based on the current extent of the disease, thereby improving the effectiveness of the output object analysis results.
[0126] In an exemplary embodiment, as shown in FIG5 , before step 202 of acquiring the image to be detected corresponding to the target object and the trained multimodal dialogue language model, the object analysis method of the present application further includes steps 502 to 510 .
[0127] Step 502 : Acquire an image sample data set, wherein the image sample data set includes image sample data, reference description text corresponding to each image sample data, and reference analysis results corresponding to the type of the image sample data.
[0128] In an embodiment of the present application, the image sample data set includes image sample data such as fundus images, reference description texts corresponding to fundus images of different degrees of lesions, and the type of image sample data can be symptom types under different preset indicators. For example, the image sample data set contains fundus images of multiple symptom types, and the reference description text indicates whether the symptoms corresponding to the preset indicators are included. The reference analysis results are diagnosis and treatment recommendations corresponding to the disease types contained in the fundus images. In one embodiment of the present application, the image sample data is fundus images that have undergone quality control to eliminate poor imaging quality, and the fundus images are pre-processed. The fundus images are cropped to a size of 224×224, and the cropped fundus images are enhanced by changing brightness, flipping, random cropping, displacement, and adding Gaussian noise, and then normalized to construct a multimodal data set of fundus image-coronary angiography diagnostic report pairs.
[0129] Step 504 : Train the visual model based on the image sample dataset and the momentum contrast learning framework to obtain a pre-trained visual model.
[0130] In an embodiment of the present application, the server uses a self-supervised momentum contrast (MoCo, Momentum Contrust) learning framework to train the visual model to construct a pre-trained visual model. The pre-trained visual model learns a common visual representation of the fundus during training. These common feature representations can be transferred to the coronary heart disease detection task, and good performance can be achieved using only relatively small labeled data. Positive samples and negative samples are constructed through image sample data sets, and positive sample pairs and negative sample pairs are respectively composed. The similarity between the positive sample pairs is maximized and the similarity between the negative sample pairs is minimized, and the pre-trained visual model is trained.
[0131] In step 506 , the visual structure of the multimodal dialogue language model is trained based on the output of the pre-trained visual model in accordance with the knowledge distillation method to obtain a trained visual structure.
[0132] In an embodiment of the present application, knowledge distillation is a model compression technology. The server uses the pre-trained visual model as a teacher model to guide the visual structure learning of the multimodal dialogue language model to extract fundus features. That is, the model parameters of the pre-trained visual model are indirectly embedded into the multimodal dialogue language model, and the output of the pre-trained visual model is used as the target of the language structure to complete the training of the visual structure.
[0133] Step 508: construct an attention structure with a multi-branch architecture according to preset indicators, train the attention structure based on the image sample data and the reference description text corresponding to the image sample data, and obtain a trained attention structure.
[0134] The server can build a multi-branch attention structure based on a variety of preset indicators. The multi-branch architecture mainly sets multiple bypasses around the visual structure of the multimodal dialogue language model. In different types of image sample data, it is fine-tuned based on different types of image sample data and the reference description text corresponding to the image sample data to obtain a trained attention structure.
[0135] Step 510 : Based on the reference description text corresponding to the image sample data and the reference analysis result corresponding to the type of the image sample data, the language structure of the multimodal dialogue language model is trained to obtain a trained language structure.
[0136] In the embodiment of the present application, the reference description text corresponding to the image sample data is the initial diagnosis result, and the reference analysis result corresponding to the type of image sample data can be text data combining the medical order text and the initial diagnosis result. The server uses the low-rank adaptation of large language models (LoRA) fine-tuning method to learn the relationship between the coronary angiography diagnosis conclusion and the medical order through a bypass, and then generates a reference analysis result corresponding to the reference description text based on preset indicators.
[0137] In this embodiment, knowledge distillation training using a pre-trained visual model yields a trained visual structure, which can improve the efficiency and accuracy of image feature extraction by a multimodal conversational language model. By constructing an attention structure and training it based on image sample data and the corresponding reference description text, a trained attention structure is obtained. Furthermore, a language structure is trained based on the corresponding reference description text and the analysis results corresponding to the image sample data type. This allows the nonlinear relationship between the image and the reference description text to be learned, as well as the nonlinear relationship between the reference description text and the reference analysis results, resulting in a trained language structure. Furthermore, a multimodal conversational language model, which incorporates the trained visual structure, attention structure, and language structure, is used to process the image and text data to be detected, thereby improving the accuracy of object analysis results.
[0138] In an exemplary embodiment, as shown in Figure 6, the multimodal conversational language model training process includes training for visual structure, attention structure, and language structure. Using knowledge distillation technology, the pre-trained visual model serves as the teacher model, and the visual structure serves as the student model. Backward gradient propagation training is performed on the visual structure, and the parameters of the visual structure are adjusted until the output of the visual structure meets a preset similarity condition with the output of the teacher model, resulting in a trained visual structure.
[0139] For the training of the attention structure, the contrast loss is calculated between the fundus features obtained by the attention structure and the semantic codes extracted from the diagnostic report. The attention structure is fine-tuned by LoRA through the contrast loss to complete the training of the attention structure.
[0140] For the training of language structure, autoregressive calculation is performed through the diagnostic report to obtain the sub-regression loss, and the language structure is fine-tuned by LoRA through the sub-regression loss to improve the accuracy of the language structure output.
[0141] In an exemplary embodiment, as shown in FIG. 7 , step 502 of acquiring an image sample dataset includes steps 702 to 704 .
[0142] Step 702: Acquire initial image data.
[0143] In the embodiment of the present application, the server obtains initial image data, which can be used to capture a fundus image of the subject. A medical staff member uses a fundus camera to capture the fundus image of the subject, using non-contact photography technology, automatic focus, and automatic exposure to ensure that each capture is as clear and appropriately exposed as possible.
[0144] Step 704 : Screen the initial image data according to its clarity to obtain image sample data.
[0145] In an embodiment of the present application, the server first performs quality control on a large number of fundus images, that is, the entire data set is divided into three subsets according to the imaging quality and clarity of the fundus images, each representing different levels of imaging quality: the first level - clear imaging; the second level - blurred imaging, but can provide optic disc information and vascular information; the third level - poor imaging, erroneously capturing objects other than the fundus or providing almost no fundus information.
[0146] Next, the third data set is eliminated, and image enhancement is performed on the first and second fundus images to increase the contrast of blood vessels and optic discs compared to other pixels.
[0147] Step 504 trains the visual model based on the image sample dataset and the momentum contrast learning framework to obtain a pre-trained visual model, including steps 706 to 708.
[0148] Step 706 : Generate positive sample pairs and negative sample pairs based on the momentum contrast learning framework and each image sample data.
[0149] In the embodiment of the present application, the server constructs positive sample pairs and negative sample pairs based on the fundus dataset. The positive sample pairs are created by applying data enhancement to the original images in the image sample dataset, while the negative sample pairs are created by randomly selecting images from the image sample dataset.
[0150] Step 708: Train the visual model according to the first loss function, the positive sample pairs, and the negative sample pairs to obtain a pre-trained visual model.
[0151] In the embodiment of the present application, as shown in FIG8 , a first loss function represents the contrast loss corresponding to the positive sample pair and the negative sample pair, respectively. The server performs feature encoding and nonlinear mapping based on the enhanced image 1 and enhanced image 2 corresponding to the positive sample pair, calculates a similarity measure based on the momentum comparison between enhanced image 1 and enhanced image 2, and then calculates the contrast loss. Furthermore, feature encoding and nonlinear mapping are calculated based on the negative sample pair, and the similarity measure and contrast loss are calculated using the same method. The goal is to reduce the similarity between the positive sample pairs and increase the similarity between the negative sample pairs.
[0152] In this embodiment, the momentum contrast training method can be used to preliminarily obtain the feature representation of the fundus image, and then fine-tune the visual structure on a smaller labeled dataset in specific fundus-related tasks, thereby improving the accuracy of feature extraction of the multimodal dialogue language model.
[0153] In one exemplary embodiment, the first loss function is a normalized mutual information loss function, and the pre-trained visual model includes a query encoder and a target encoder. As shown in FIG9 , step 708 trains the visual model based on the first loss function, positive sample pairs, and negative sample pairs to obtain a pre-trained visual model, including steps 902 to 906.
[0154] Step 902 : Process the positive sample pairs according to the query encoder and the target encoder respectively to obtain a first feature vector group corresponding to the query encoder and a second feature vector group corresponding to the target encoder.
[0155] In the embodiment of the present application, Momentum Contrast (MoCo) uses a target encoder and a query encoder. The target encoder updates its own representation through a momentum update formula, which fuses the feature vector of the current query encoder with the feature vector of the target encoder. The purpose of updating this formula is to maintain the stability and consistency of the target encoder. The update formula is as follows: E k+1 =m×E k +(1-m)×q k (1)
[0156] Among them, E k+1 represents the updated target encoder, E k represents the current target encoder, q k represents the current query encoder and m is a momentum factor.
[0157] The server processes the positive samples according to the query encoder and the target encoder respectively to obtain two feature vector groups, namely a first feature vector group and a second feature vector group, which are used to perform momentum update on the target encoder.
[0158] Step 904 : Process the negative sample pairs according to the query encoder to obtain a third feature vector group corresponding to the negative sample pairs.
[0159] In the embodiment of the present application, for a negative sample pair, the query feature vector generated by the query encoder is denoted as q, and the target feature vector generated by the target encoder is denoted as k. + , and the target feature vector randomly selected from the batch target feature vector is denoted as k i Among them, the number of negative sample pairs K can be controlled by setting.
[0160] Step 906 : Training the visual model based on the normalized mutual information loss function, the first eigenvector group, the second eigenvector group, and the third eigenvector group to obtain a pre-trained visual model.
[0161] In the embodiment of the present application, MoCo uses normalized mutual information loss (InfoNCE Loss) to train the network. This loss function measures the similarity between the query encoder and the target encoder. It compares the similarity of the positive sample pairs with the similarity of the negative sample pairs to maximize the similarity of the positive sample pairs and minimize the similarity of the negative sample pairs, that is, the momentum factor in formula (1) is parameterized according to the normalized mutual information loss function, the first eigenvector group, the second eigenvector group, and the third eigenvector group. Normalized mutual information loss function L q The formula is as follows:
[0162] Among them, the query feature vector generated by the query encoder is recorded as q, and the target feature vector generated by the target encoder is recorded as k + , τ is the temperature coefficient. The goal of normalized mutual information loss is to maximize the similarity of positive sample pairs.
[0163] In this embodiment, the momentum contrast training method can be used to preliminarily obtain the feature representation of the fundus image, and then fine-tune the visual structure on a smaller labeled dataset in specific fundus-related tasks, thereby improving the accuracy of feature extraction of the multimodal dialogue language model.
[0164] In one exemplary embodiment, the visual structure includes a backbone network and a first bypass. As shown in FIG10 , step 506 trains the visual structure of the multimodal dialogue language model based on the output of the pre-trained visual model using a knowledge distillation approach to obtain a trained visual structure, including steps 1002 to 1006.
[0165] Step 1002: Train the visual structure of the multimodal dialogue language model based on a preset training method. During the training process, the backbone network and the first bypass are used as student models, and the pre-trained visual model is used as a teacher model.
[0166] In an embodiment of the present application, a pre-trained visual model is constructed by using the MoCo contrastive learning method. Compared with the visual structure of the multimodal dialogue language model, the model can more fully capture the complex features and patterns in the fundus image, such as blood vessels, optic discs and lesions, thereby encoding the fundus image. Next, the pre-trained visual model will be used for fine-tuning to adjust the output of the pre-trained visual model from the multimodal dialogue language model. Using the pre-trained visual model as the teacher model, LoRA fine-tuning is injected into the pre-trained visual model of the multimodal dialogue language model, that is, a bypass (bypass1) is added, and the backbone network of the visual structure and the bypass are used together as student models. As shown in Figure 11, Figure 11 shows the principle of the LoRA algorithm, which reduces the number of parameters required to fit downstream tasks by freezing the pre-trained model weights and injecting the trainable rank decomposition matrix into each layer of the transformer.
[0167] Step 1004: measure the first output result of the teacher model and the second output result of the student model according to the target measurement function to obtain the similarity between the teacher model and the student model.
[0168] In the embodiment of the present application, for the target sample data, the server introduces an objective function to measure the similarity between the teacher model and the student model, which can be expressed as follows: loss = -∑(T×log(P)) (3)
[0169] Among them, T is the fundus feature representation output by the teacher model, and P is the fundus feature representation output by the student model.
[0170] Step 1006: Train the backbone network and the first bypass corresponding to the student model using the second loss function and similarity to obtain a trained visual structure.
[0171] In an embodiment of the present application, the backbone network weights of the trained visual structure will be frozen, and the bypass weights will be randomly initialized and continuously updated with gradients.
[0172] During the distillation process, the server uses the teacher model to encode the fundus image and uses the teacher model output as the target of the student model. The student model generates a feature representation based on the input fundus image and compares it with the target of the teacher model. By minimizing the difference between the teacher model output and the student model output, the student model gradually learns to encode the fundus in the encoding method of the teacher model, that is, after knowledge distillation, the weight of bypass1 is retained for downstream tasks. Among them, according to the difference between the teacher model output and the student model output, the LoRA algorithm is used for training. Specifically, during the training process, the full parameters of the original model are represented by the symbol φ, a new bypass is created to initialize the weight and the gradient is continuously updated. The weight parameter of the bypass is recorded as θ, and |θ|<<|φ|. Each iteration is φ=φ+Δφ(θ). At this time, the training process is equivalent to optimizing the following formula:
[0173] Among them, φ represents all the parameters in the model, p φ +Δφ(θ)(y t |x, y<t) is the next token task predicted in the dialogue language model, and z represents the entire dataset, that is, the fundus image-to-be-predicted indicator data pair. Each data piece consists of input data x and output sequence y.
[0174] As shown in Figure 11, matrix W represents a fully connected layer in the trained visual architecture. Assuming the dimension of this layer is d×d, during fine-tuning, all weights in this layer are frozen and no gradient updates are performed. On the right is the newly created network bypass, consisting of two parts: fully connected layer A and fully connected layer B. The dimensions of fully connected layer A are d×r, and the dimensions of fully connected layer B are r×d. Therefore, the total number of parameters in this bypass is 2×r×d. Since r << d, (2×r×d) << (d×d). During fine-tuning, only the weights in this bypass are gradient updated.
[0175] During forward propagation, data x enters both the backbone network and the bypass network. The output dimensions of the matrix W on the left and the two matrices A and B on the right are the same, both d. The final output is obtained by directly adding the outputs on the left and right sides, expressed as: h = Wx + BAx (5)
[0176] In this embodiment, during the fine-tuning process, the pre-trained visual model based on the fundus is indirectly embedded into the trained visual structure through knowledge distillation, so that the trained visual structure can better extract fundus image features, thereby improving the accuracy of feature extraction of the trained visual structure.
[0177] In one exemplary embodiment, the attention structure includes multiple second bypasses. As shown in FIG12 , step 508 constructs a multi-branch attention structure based on preset indicators, trains the attention structure based on image sample data and reference description text corresponding to the image sample data, and obtains a trained attention structure, including steps 1202 to 1206.
[0178] Step 1202 , classify the image sample data based on preset indicators to obtain multiple types of target image sample sets, reference description text corresponding to each target image sample in the target image sample set, and analysis results corresponding to each target sample in each type of target image sample set.
[0179] In an embodiment of the present application, the preset indicators can be different symptoms in the coronary angiography diagnosis conclusion, such as the coronary artery lumen stenosis, coronary artery calcification score, the degree of coronary atherosclerosis, and whether there is myocardial infarction. The server exports the coronary angiography diagnosis report of the subject. A number of key indicators in the coronary angiography examination conclusion are extracted by regular matching, and are combined with the fundus images respectively to construct different data sets. The data set is composed of reference description texts corresponding to the image sample data, and the subject's medical advice is obtained, combined with the coronary angiography diagnosis conclusion, and organized into data pairs. Among them, the data set is shown in Table 1 below:
[0180] Table 1
[0181] Step 1204 : for each second bypass, determine the target type corresponding to each second bypass, perform data processing on the target image sample set and the reference description text of the target type according to the second bypass, and obtain a visual vector and a first semantic code respectively.
[0182] In this embodiment of the present application, the multi-branch architecture is primarily implemented by setting up multiple bypasses around the attention structure of the multimodal conversational language model, each fine-tuned based on different sub-datasets. In this embodiment, four additional bypasses are added for different prediction tasks, namely, for the four prediction objects of coronary artery stenosis, coronary artery calcification score, coronary artery atherosclerosis degree, and the presence of myocardial infarction. The server processes the target image sample set and reference description text of the target type based on the second bypass, respectively, to obtain the visual vector of the target image sample set and the first semantic encoding of the reference description text.
[0183] Step 1206: Train the attention structure based on the contrast loss of the visual vector and the first semantic code to obtain a trained attention structure.
[0184] In the embodiment of the present application, the idea of LoRA fine-tuning is also followed. Under the condition of freezing the weights of the backbone network, the gradients of the four bypasses are continuously updated to adapt to different tasks. The network structure of this bypass is the same as the bypass passby1 idea. The fine-tuning process is to obtain the visual vector of the fundus image and the semantic encoding of the corresponding descriptive text under a specific task, that is, the first semantic encoding and the visual vector and calculate the contrast loss, and then minimize the back propagation gradient by the contrast loss to achieve feature alignment of the visual vector and the semantic encoding, and obtain the trained attention structure.
[0185] In this embodiment, by adopting a multi-task learning approach and setting multiple bypasses for fine-tuning, some key indicators can be better predicted, so that the initial diagnosis results can reflect the indicator information contained in the image to be detected in a variety of preset indicators, thereby improving the accuracy of generating diagnostic analysis.
[0186] In an exemplary embodiment, as shown in FIG13 , step 510 trains the language structure of the multimodal conversational language model based on the reference description text corresponding to the image sample data and the reference analysis results corresponding to the type of the image sample data to obtain a trained language structure, including steps 1302 to 1306 . In particular:
[0187] Step 1302: Process the target image sample according to the trained attention structure to obtain an initial diagnosis result.
[0188] In an embodiment of the present application, in order to enable the multimodal dialogue language model to determine accurate diagnosis and treatment information based on the initial diagnosis result, the server obtains the initial diagnosis result through the trained attention structure, so as to train the language structure through the initial diagnosis result and the reference analysis result.
[0189] Step 1304 : For each type of target image sample set, data processing is performed on the initial diagnosis results corresponding to the target image samples according to the language structure to obtain a second semantic code.
[0190] In an embodiment of the present application, the server performs data processing on each type of target image sample set through the language structure to be trained to obtain an object analysis result corresponding to the initial diagnosis result of each type of target image sample set. The object analysis result is in the form of encoding, namely, the second semantic encoding.
[0191] Step 1306 : Based on the contrast loss between the second semantic encoding and the third semantic encoding corresponding to the reference description text, the language structure of the multimodal dialogue language model is trained to obtain a trained language structure.
[0192] In an embodiment of the present application, the language structure of the multimodal dialogue language model is fine-tuned by LoRA based on the subject's coronary angiography diagnosis conclusion and the corresponding doctor's order, and the relationship between the coronary angiography diagnosis conclusion and the doctor's order is learned through a bypass. As shown in Figure 14, the server first encodes the reference analysis result to obtain a third semantic code, and calculates the autoregressive loss based on the third semantic code and the second semantic code corresponding to the diagnosis and treatment recommendation (initial diagnosis result) output by the language structure, and trains based on the autoregressive loss to obtain a trained language structure. The trained visual structure, trained attention structure, and trained language structure constitute a trained multimodal dialogue language model.
[0193] In this embodiment, the language structure can not only identify whether the patient has coronary heart disease, but also generate descriptive text to describe the disease condition. Through multiple rounds of dialogue, personalized diagnosis and treatment suggestions can be automatically output, thereby improving the efficiency of determining the object analysis results.
[0194] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0195] Based on the same inventive concept, embodiments of the present application also provide an object analysis device for implementing the object analysis method described above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more object analysis device embodiments provided below can be found in the above-mentioned limitations of the object analysis method and will not be repeated here.
[0196] In an exemplary embodiment, as shown in FIG15 , an object analysis device 1500 is provided, comprising: a first acquisition module 1501 , a feature extraction module 1502 , an initial diagnosis module 1503 and a diagnosis analysis module 1504 .
[0197] The first acquisition module 1501 is used to acquire the image to be detected corresponding to the target object and the trained multimodal dialogue language model. The trained multimodal dialogue language model includes a trained visual structure, a trained attention structure, and a trained language structure.
[0198] The feature extraction module 1502 is used to extract features of the image to be detected based on the trained visual structure to obtain image features corresponding to the image to be detected.
[0199] The initial diagnosis module 1503 is configured to process image features based on the trained attention structure to obtain an initial diagnosis result corresponding to the target object. The initial diagnosis result includes at least one of a plurality of preset indicators.
[0200] The diagnosis analysis module 1504 is used to perform data processing on the initial diagnosis result based on the trained language structure to obtain an object analysis result corresponding to the target object.
[0201] In one embodiment, the diagnosis analysis module 1504 is specifically used to perform data processing on the initial diagnosis result according to the trained language structure to obtain a medical order text; and construct an object analysis result corresponding to the target object according to the initial diagnosis result and the medical order text.
[0202] In one embodiment, the apparatus 1500 further includes: a second acquisition module, a first training module, a second training module, a third training module, and a fourth training module.
[0203] The second acquisition module is used to acquire an image sample data set, wherein the image sample data set includes image sample data, reference description text corresponding to each image sample data, and reference analysis results corresponding to the type of the image sample data.
[0204] The first training module is used to train the visual model based on the image sample dataset and the momentum contrast learning framework to obtain a pre-trained visual model.
[0205] The second training module is used to train the visual structure of the multimodal dialogue language model based on the output of the pre-trained visual model according to the knowledge distillation method to obtain a trained visual structure.
[0206] The third training module is used to construct an attention structure with a multi-branch architecture according to preset indicators, and train the attention structure based on image sample data and reference description text corresponding to the image sample data to obtain a trained attention structure.
[0207] The fourth training module is used to train the language structure of the multimodal dialogue language model based on the reference description text corresponding to the image sample data and the reference analysis results corresponding to the type of the image sample data to obtain a trained language structure.
[0208] In one embodiment, the second acquisition module is specifically configured to acquire initial image data; and filter the initial image data according to the clarity of the initial image data to obtain image sample data.
[0209] The first training module is specifically used to generate positive sample pairs and negative sample pairs based on the momentum contrast learning framework and each image sample data; train the visual model according to the first loss function, the positive sample pairs and the negative sample pairs to obtain a pre-trained visual model.
[0210] In one embodiment, the first loss function is a normalized mutual information loss function, and the pre-trained visual model includes a query encoder and a target encoder.
[0211] The first training module is specifically used to process the positive sample pairs according to the query encoder and the target encoder respectively to obtain the first feature vector group corresponding to the query encoder and the second feature vector group corresponding to the target encoder; process the negative sample pairs according to the query encoder to obtain the third feature vector group corresponding to the negative sample pairs; train the visual model based on the normalized mutual information loss function, the first feature vector group, the second feature vector group and the third feature vector group to obtain a pre-trained visual model.
[0212] In one embodiment, the visual structure includes a backbone network and a first bypass.
[0213] The second training module is specifically used to train the visual structure of the multimodal dialogue language model based on a preset training method. During the training process, the backbone network and the first bypass are used as student models, and the pre-trained visual model is used as the teacher model; the first output result of the teacher model and the second output result of the student model are measured according to the target measurement function to obtain the similarity between the teacher model and the student model; the backbone network and the first bypass corresponding to the student model are trained through the second loss function and similarity to obtain the trained visual structure.
[0214] In one embodiment, the attention structure includes a plurality of second bypass paths.
[0215] The third training module is specifically used to classify image sample data based on preset indicators to obtain multiple types of target image sample sets, reference description text corresponding to each target image sample in the target image sample set, and analysis results corresponding to each target sample in each type of target image sample set; for each second bypass, determine the target type corresponding to each second bypass, and perform data processing on the target image sample set and the reference description text of the target type according to the second bypass to obtain a visual vector and a first semantic code, respectively; train the attention structure based on the contrast loss of the visual vector and the first semantic code to obtain a trained attention structure.
[0216] In one embodiment, the fourth training module is specifically used to perform data processing on the target image samples according to the trained attention structure to obtain an initial diagnosis result; for each type of target image sample set, the initial diagnosis result corresponding to the target image sample is data processed according to the language structure to obtain a second semantic code; the language structure of the multimodal dialogue language model is trained based on the contrast loss of the second semantic code and the third semantic code corresponding to the reference description text to obtain the language structure of the trained multimodal dialogue language model.
[0217] Each module in the object analysis device described above may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0218] In an exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be shown in Figure 16. The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, memory, and I / O interface are connected via a system bus, and the communication interface is connected to the system bus via the I / O interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store images to be detected and trained multimodal dialogue language models. The I / O interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements an object analysis method.
[0219] Those skilled in the art will understand that the structure shown in Figure 16 is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0220] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0221] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0222] In one embodiment, a computer program product is provided, comprising computer-executable instructions, which, when executed by a processor, implement the steps in the above-mentioned method embodiments.
[0223] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0224] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.
[0225] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0226] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. An object analysis method, comprising: Obtaining an image to be detected corresponding to the target object and a trained multimodal dialogue language model, wherein the trained multimodal dialogue language model includes a trained visual structure, a trained attention structure, and a trained language structure; Extracting features of the image to be detected based on the trained visual structure to obtain image features corresponding to the image to be detected; Performing data processing on the image features based on the trained attention structure to obtain an initial diagnosis result corresponding to the target object; The initial diagnosis result is subjected to data processing based on the trained language structure to obtain an object analysis result corresponding to the target object.
2. The method according to claim 1, characterized in that The performing data processing on the initial diagnosis result based on the trained language structure to obtain an object analysis result corresponding to the target object includes: Performing data processing on the initial diagnosis result according to the trained language structure to obtain a medical order text; An object analysis result corresponding to the target object is constructed based on the initial diagnosis result and the medical order text.
3. The method according to claim 1 or 2, characterized in that Before obtaining the image to be detected corresponding to the target object and the trained multimodal dialogue language model, the method further includes: Acquire an image sample data set, wherein the image sample data set includes image sample data, reference description text corresponding to the image sample data, and reference analysis results corresponding to the type of the image sample data; Training a visual model based on the image sample dataset and a momentum contrast learning framework to obtain a pre-trained visual model; Training the visual structure of the multimodal dialogue language model based on the output of the pre-trained visual model in a knowledge distillation manner to obtain the trained visual structure; Constructing a multi-branch attention structure according to preset indicators, and training the attention structure based on the image sample data and the reference description text corresponding to the image sample data to obtain the trained attention structure; Based on the reference description text corresponding to the image sample data and the reference analysis result corresponding to the type of the image sample data, the language structure of the multimodal dialogue language model is trained to obtain the trained language structure.
4. The method according to claim 3, characterized in that The acquiring of the image sample data set includes: Obtaining initial image data; The initial image data is screened according to the clarity of the initial image data to obtain image sample data.
5. The method according to claim 3 or 4, characterized in that The step of training the visual model based on the image sample dataset and the momentum contrast learning framework to obtain the pre-trained visual model includes: Based on the momentum contrastive learning framework and the image sample data, generating positive sample pairs and negative sample pairs; The visual model is trained according to the first loss function, the positive sample pairs, and the negative sample pairs to obtain the pre-trained visual model.
6. The method according to claim 5, characterized in that: The first loss function is a normalized mutual information loss function, and the pre-trained visual model includes a query encoder and a target encoder; The pre-trained visual model is trained according to the first loss function, the positive sample pairs, and the negative sample pairs to obtain the pre-trained visual model, comprising: Performing data processing on the positive sample pairs according to the query encoder and the target encoder respectively to obtain a first feature vector group corresponding to the query encoder and a second feature vector group corresponding to the target encoder; performing data processing on the negative sample pair according to the query encoder to obtain a third feature vector group corresponding to the negative sample pair; The visual model is trained based on the normalized mutual information loss function, the first eigenvector group, the second eigenvector group, and the third eigenvector group to obtain the pre-trained visual model.
7. The method according to any one of claims 3 to 5, characterized in that: The visual structure includes a backbone network and a first bypass; The visual structure of the multimodal dialogue language model is modified based on the output of the pre-trained visual model in the way of knowledge distillation. Perform training to obtain the trained visual structure, including: Training the visual structure of the multimodal dialogue language model based on a preset training method, and during the training process, using the backbone network and the first bypass as student models and the pre-trained visual model as a teacher model; Measuring the first output result of the teacher model and the second output result of the student model according to the target measurement function to obtain the similarity between the teacher model and the student model; The backbone network and the first bypass corresponding to the student model are trained using a second loss function and the similarity to obtain a trained visual structure.
8. The method according to any one of claims 3 to 7, characterized in that: The attention structure includes a plurality of second bypass paths; The method of constructing a multi-branch attention structure according to preset indicators, and training the attention structure based on the image sample data and the reference description text corresponding to the image sample data to obtain the trained attention structure includes: Classifying the image sample data based on preset indicators to obtain multiple types of target image sample sets, reference description text corresponding to each target image sample in the target image sample sets, and analysis results corresponding to each target sample in each type of the target image sample sets; For each of the second bypass paths, determining a target type corresponding to each of the second bypass paths, and performing data processing on the target image sample set and the reference description text of the target type according to the second bypass paths to obtain a visual vector and a first semantic code, respectively; The attention structure is trained based on the contrast loss of the visual vector and the first semantic code to obtain a trained attention structure.
9. The method according to any one of claims 3 to 8, characterized in that: The training of the language structure of the multimodal dialogue language model based on the reference description text corresponding to the image sample data and the reference analysis result corresponding to the type of the image sample data to obtain the trained language structure includes: Performing data processing on the target image sample according to the trained attention structure to obtain an initial diagnosis result; For each type of the target image sample set, performing data processing on the initial diagnosis results corresponding to the target image samples according to the language structure to obtain a second semantic code; Based on the contrast loss between the second semantic encoding and the third semantic encoding corresponding to the reference description text, the language structure of the multimodal dialogue language model is trained to obtain the trained language structure.
10. The method according to any one of claims 1 to 9, characterized in that The initial diagnosis result includes at least one of a plurality of preset indicators; the plurality of preset indicators include relevant information describing the disease.
11. The method according to any one of claims 1 to 10, characterized in that: The step of extracting features from the image to be detected based on the trained visual structure to obtain image features corresponding to the image to be detected includes: Inputting the image to be detected corresponding to the target object and the task type corresponding to the prompt word information into the trained multimodal dialogue language model; According to the trained visual structure, feature extraction is performed on the image to be detected through a forward propagation process to obtain image features corresponding to the image to be detected; The method further comprises: representing the image to be detected as a feature vector containing the image features.
12. The method according to claim 2, wherein: The constructing the object analysis result corresponding to the target object according to the initial diagnosis result and the medical order text includes: Associating the initial diagnosis result with the medical order text; According to the trained language structure, the initial diagnosis result and the medical order text following the initial diagnosis result are constructed into a coherent description to obtain the object analysis result.
13. The method according to claim 5, wherein: The positive sample pairs are created by applying data augmentation to original images in the image sample dataset, and the negative sample pairs are created by randomly selecting images from the image sample dataset.
14. An object analysis device, characterized in that: The device comprises: A first acquisition module is configured to acquire an image to be detected corresponding to a target object and a trained multimodal dialogue language model, wherein the trained multimodal dialogue language model includes a trained visual structure, a trained attention structure, and a trained language structure; A feature extraction module is used to extract features of the image to be detected based on the trained visual structure to obtain image features corresponding to the image to be detected; An initial diagnosis module, configured to perform data processing on the image features based on the trained attention structure to obtain an initial diagnosis result corresponding to the target object; The diagnosis analysis module is used to perform data processing on the initial diagnosis result based on the trained language structure to obtain an object analysis result corresponding to the target object.
15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 13 are implemented.
16. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
Citation Information
Patent Citations
Image prediction model generation method and device, computer equipment and storage medium
CN115359005A
Medical image interpretability analysis system and analysis method
CN116485777A
Visual positioning method, device and equipment based on double knowledge distillation and memory
CN116778140A
Ultrasonic diagnosis intelligent interaction system based on liver attribute analysis
CN117333462A
GPT-based knee joint lesion diagnosis intelligent self-generation method, device and equipment
CN117352120A
Cited By
Method and device for predicting peak value of video topic
CN120804361A
Strategy optimization method and device for multi-modal action model, equipment and medium
CN120877387A
Language guidance feature decoupling infrared target detection method
CN121330249A