Subject analysis method and apparatus, and computer device and storage medium

Through the multi-head self-attention structure and multi-layer perceptron feature extractor combined with a large-scale language analysis model, the problem of ignoring fundus image characteristics in traditional methods is solved, and a high-accuracy diagnosis of diabetic retinopathy is achieved. It is suitable for terminal and server image analysis systems.

WO2025138569A1PCT designated stage expired Publication Date: 2025-07-03TSINGHUA UNIVERSITY +1

Patent Information

Application Number
PCT/CN2024/095777
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-05-28
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing diagnostic methods for diabetic retinopathy rely on visual examinations by professional ophthalmologists, making it difficult to achieve timely and accurate diagnosis in resource-poor and primary medical environments. The traditional convolutional neural network model ignores some characteristics of fundus images, resulting in poor accuracy of diagnosis and treatment results.

Method used

The feature extractor of multi-head self-attention structure and multi-layer perceptron is used to extract the image sequence feature, combine pre-trained classification sub-models and large-scale language analysis models, and generate target image sequences through image segmentation and convolution operations, and combine clinical data for object analysis to improve feature extraction and diagnostic accuracy.

Benefits of technology

By conducting more comprehensive feature extraction and multi-factor analysis of the details of the detected image, the analysis accuracy of the image processing model is improved, the ability to identify lesions is enhanced, and more accurate diagnostic results are provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024095777_03072025_PF_FP_ABST
    Figure CN2024095777_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A subject analysis method and apparatus, and a computer device, a storage medium and a computer program product. The subject analysis method comprises: acquiring a target image sequence, clinical data and an image processing model, wherein the target image sequence is obtained by means of performing image sequence conversion on an image to be subjected to detection, and the image processing model includes a pre-trained feature extractor and a trained classification sub-model; performing feature extraction on the target image sequence by means of the feature extractor, and performing analysis processing on a feature extraction result by means of the classification sub-model, so as to obtain a target analysis result; and on the basis of a large-scale language analysis model, performing subject analysis on the target analysis result and the clinical data, so as to obtain a subject analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

Object analysis method, device, computer equipment and storage medium

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 202311846783.0, filed on December 29, 2023, entitled “Object Analysis Method, Apparatus, Computer Equipment and Storage Medium,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to the field of artificial intelligence technology, and in particular to an object analysis method, apparatus, computer equipment, storage medium, and computer program product. Background Art

[0004] Diabetic retinopathy (DR) is one of the leading causes of visual impairment and blindness worldwide. Current DR diagnostic methods rely primarily on visual inspection and assessment by professional ophthalmologists. However, timely and accurate diagnosis and treatment of DR remain a challenge in resource-poor and primary care settings.

[0005] In traditional technology, by applying neural network and machine learning techniques, the DR diagnostic model is trained based on a large number of retinal fundus image datasets, and the retinal fundus images provided by the patient are analyzed based on the Convolutional Neural Network (CNN) model. The fundus images are convolved through the CNN model to output the current patient's diagnosis and treatment results.

[0006] However, in traditional technology, global convolution is performed on fundus images for feature extraction, which ignores some features and only analyzes the symptoms contained in the fundus images. The analysis factors are relatively simple, resulting in poor accuracy of diagnosis and treatment results.

[0007] Summary of the Invention

[0008] Based on this, it is necessary to provide an object analysis method, apparatus, computer device, computer-readable storage medium and computer program product to address the above technical issues.

[0009] In a first aspect, the present application provides an object analysis method, including: acquiring a target image sequence, clinical data and an image processing model; the target image sequence is obtained by performing image sequence conversion on the image to be detected; the image processing model includes a pre-trained feature extractor and a trained classification sub-model; feature extraction is performed on the target image sequence by the feature extractor, and the result of the feature extraction is analyzed and processed by the classification sub-model to obtain a target analysis result; object analysis is performed on the target analysis result and the clinical data based on a large-scale language analysis model to obtain an object analysis result.

[0010] In one embodiment, acquiring the target image sequence includes: performing image segmentation on the image to be detected according to a preset segmentation method to obtain a plurality of image blocks; and performing a convolution operation on each of the image blocks to obtain the target image sequence.

[0011] In one embodiment, the feature extractor includes a multi-head self-attention structure and a multi-layer perceptron.

[0012] The feature extractor is used to extract features from the target image sequence, and the classification sub-model is used to analyze and process the feature extraction results to obtain target analysis results, including: calculating the triple features of the target image sequence; transforming the triple features based on a multi-head self-attention structure and a multi-layer perceptron to obtain a target feature vector; and classifying the target feature vector according to the classification sub-model to obtain a target analysis result.

[0013] In one embodiment, the triple features are transformed based on the multi-head self-attention structure and the multi-layer perceptron to obtain the target feature vector, including: mapping the triple features according to each self-attention layer in the multi-head self-attention mechanism to obtain the mapping result of each self-attention layer, and obtaining a joint mapping result based on each mapping result; transforming the joint mapping result through the multi-layer perceptron to obtain the target feature vector.

[0014] In one embodiment, the classification sub-model includes a quality assessment sub-model, a first symptom grading sub-model, a second symptom judgment sub-model and a lesion segmentation sub-model.

[0015] The target feature vector is classified and processed according to the classification sub-model to obtain a target analysis result, including: the target feature vector is classified and processed according to the image quality assessment sub-model, the first symptom grading sub-model, the second symptom judgment sub-model and the lesion segmentation sub-model respectively to obtain an image quality assessment result, a first symptom level, a second symptom diagnosis result and a lesion segmentation result; and the image quality assessment result, the first symptom level, the second symptom diagnosis result and the lesion segmentation result are used as the target analysis result.

[0016] In one embodiment, before acquiring the target image sequence, clinical data and image processing model, the method further includes: acquiring a sample image set, and performing data enhancement on each sample image in the sample image set to obtain a first reconstructed image block and a second reconstructed image block; performing feature extraction on the first reconstructed image block and the second reconstructed image block respectively to obtain query features and key feature sequences, and determining a contrast loss value based on the query features and the key feature sequence; reconstructing the reconstructed image block according to a preset decoder to obtain a reconstructed image, and determining a reconstruction loss value based on the reconstructed image and the sample image corresponding to the reconstructed image; training a feature extractor included in the image processing model based on the contrast loss value and the reconstruction loss value to obtain the pre-trained feature extractor.

[0017] In one embodiment, the feature extraction is performed on the first reconstructed image block and the second reconstructed image block respectively to obtain query features and key feature sequences, and the contrast loss value is determined based on the query features and key feature sequences, including: for each of the first reconstructed image block and the second reconstructed image block in the sample image set, feature extraction is performed on the first reconstructed image block according to a first encoder to obtain query features, and feature extraction is performed on the second reconstructed image block according to a second encoder to obtain key features; each key feature in the key feature sequence is matched with the query feature, and a positive sample key feature in the key feature sequence is determined based on the matching result, and a contrast loss value is determined by comparing the query feature with the positive sample key features in the key feature sequence and other key features in the key feature sequence.

[0018] In one embodiment, the target image sequence is a one-dimensional embedded sequence.

[0019] In one embodiment, the large-scale language analysis model includes an adapter and a LoRA (Low-Rank Adaptation) module. The adapter is embedded in the Transformer layer, located after the feedforward network layer and before the residual connection. The adapter is a multilayer perceptron (MLP) that is used to reduce and increase the dimensionality of the feature representation of the Transformer layer. The LoRA module is set to bypass the LLaMA model.

[0020] In one embodiment, the LoRA module includes a dimensionality reduction matrix A and a dimensionality increase matrix B.

[0021] Before performing object analysis on the target analysis result and the clinical data based on the large-scale language analysis model to obtain the object analysis result, the method includes: training the large-scale language analysis model.

[0022] The training of the large-scale language analysis model includes: initializing the dimension reduction matrix A with a random Gaussian distribution and initializing the dimension increase matrix B with a zero matrix; fine-tuning the large-scale language analysis model, and updating the large-scale language analysis model to: h = W0x + ΔWx = W0x + BAx

[0023] W0 is the initialization parameter of the large-scale language analysis model, which is fixed. ΔW is the parameter to be updated. x is the input of the large-scale language analysis model. h is the dimension of the output of the large-scale language analysis model. The dimensionality reduction matrix A and the dimensionality increase matrix B contain training parameters.

[0024] In one embodiment, the plurality of image blocks are vector image blocks, and performing a convolution operation on each of the image blocks to obtain a target image sequence includes: performing a two-dimensional convolution operation on each of the vector image blocks to generate a one-dimensional embedding sequence of the image to be detected.

[0025] In a second aspect, the present application also provides an object analysis device, including: an acquisition module, a first analysis module and a second analysis module.

[0026] An acquisition module is used to acquire a target image sequence, clinical data, and an image processing model; the target image sequence is obtained by performing image sequence conversion on the image to be detected; and the image processing model includes a pre-trained feature extractor and a trained classification sub-model.

[0027] The first analysis module is used to extract features from the target image sequence through the feature extractor, and analyze and process the feature extraction results through the classification sub-model to obtain target analysis results.

[0028] The second analysis module is used to perform object analysis on the target analysis results and the clinical data based on a large-scale language analysis model to obtain object analysis results.

[0029] In one embodiment, the acquisition module is specifically configured to perform image segmentation on the image to be detected according to a preset segmentation method to obtain a plurality of image blocks; and perform a convolution operation on each of the image blocks to obtain a target image sequence.

[0030] In one embodiment, the feature extractor includes a multi-head self-attention structure and a multi-layer perceptron.

[0031] The first analysis module is specifically used to calculate the triple features of the target image sequence; transform the triple features based on the multi-head self-attention structure and the multi-layer perceptron to obtain the target feature vector; and classify the target feature vector according to the classification sub-model to obtain the target analysis result.

[0032] In one embodiment, the first analysis module is specifically used to map the triplet features according to each self-attention layer in the multi-head self-attention mechanism, obtain the mapping results of each self-attention layer, and obtain a joint mapping result based on each mapping result; and transform the joint mapping result through a multi-layer perceptron to obtain a target feature vector.

[0033] In one embodiment, the classification sub-model includes a quality assessment sub-model, a first symptom grading sub-model, a second symptom judgment sub-model and a lesion segmentation sub-model.

[0034] The first analysis module is specifically used to classify the target feature vector according to the image quality assessment sub-model, the first symptom grading sub-model, the second symptom judgment sub-model and the lesion segmentation sub-model, respectively, to obtain image quality assessment results, first symptom level, second symptom diagnosis results and lesion segmentation results; and use the image quality assessment results, the first symptom level, the second symptom diagnosis results and the lesion segmentation results as target analysis results.

[0035] In one embodiment, the device further includes: a data enhancement module, a feature extraction module, a reconstruction module and a training module.

[0036] The data enhancement module is used to obtain a sample image set and perform data enhancement on each sample image in the sample image set to obtain a first reconstructed image block and a second reconstructed image block.

[0037] A feature extraction module is used to perform feature extraction on the first reconstructed image block and the second reconstructed image block respectively to obtain query features and a key feature sequence, and determine a contrast loss value based on the query features and the key feature sequence.

[0038] The reconstruction module is configured to reconstruct the reconstructed image block according to a preset decoder to obtain a reconstructed image, and determine a reconstruction loss value based on the reconstructed image and the sample image corresponding to the reconstructed image.

[0039] A training module is used to train the feature extractor included in the image processing model based on the contrast loss value and the reconstruction loss value to obtain the pre-trained feature extractor.

[0040] In one embodiment, the feature extraction module is specifically used to perform feature extraction on the first reconstructed image block according to the first encoder to obtain a query feature, and perform feature extraction on the second reconstructed image block according to the second encoder to obtain a key feature for each of the first reconstructed image block and the second reconstructed image block in the sample image set; match each key feature in the key feature sequence with the query feature, and determine the positive sample key feature in the key feature sequence based on the matching result, and determine the contrast loss value through the query feature and the positive sample key feature in the key feature sequence, as well as other key features in the key feature sequence.

[0041] In a third aspect, the present application also provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the following steps when executing the computer program: obtaining a target image sequence, clinical data and an image processing model; the target image sequence is obtained by performing image sequence conversion on the image to be detected; the image processing model includes a pre-trained feature extractor and a trained classification sub-model; feature extraction is performed on the target image sequence by the feature extractor, and the result of the feature extraction is analyzed and processed by the classification sub-model to obtain a target analysis result; object analysis is performed on the target analysis result and the clinical data based on a large-scale language analysis model to obtain an object analysis result.

[0042] In a fourth aspect, the present application also provides a non-volatile computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps: acquiring a target image sequence, clinical data, and an image processing model; the target image sequence is obtained by performing image sequence conversion on the image to be detected; the image processing model includes a pre-trained feature extractor and a trained classification sub-model; feature extraction is performed on the target image sequence by the feature extractor, and the result of the feature extraction is analyzed and processed by the classification sub-model to obtain a target analysis result; object analysis is performed on the target analysis result and the clinical data based on a large-scale language analysis model to obtain an object analysis result.

[0043] In a fifth aspect, the present application also provides a computer program product, comprising executable instructions, which, when executed by a processor, implement the following steps: obtaining a target image sequence, clinical data, and an image processing model; the target image sequence is obtained by performing image sequence conversion on the image to be detected; the image processing model includes a pre-trained feature extractor and a trained classification sub-model; feature extraction is performed on the target image sequence by the feature extractor, and the result of the feature extraction is analyzed and processed by the classification sub-model to obtain a target analysis result; object analysis is performed on the target analysis result and the clinical data based on a large-scale language analysis model to obtain an object analysis result.

[0044] The above-mentioned object analysis method, apparatus, computer equipment, storage medium and computer program product obtain a target image sequence, clinical data and an image processing model; the target image sequence is obtained by performing image sequence conversion on the image to be detected; the image processing model includes a pre-trained feature extractor and a trained classification sub-model; the feature extractor performs feature extraction on the target image sequence, and the classification sub-model analyzes and processes the feature extraction results to obtain a target analysis result; the target analysis result and the clinical data are subjected to object analysis based on a large-scale language analysis model to obtain an object analysis result. Using this method, by performing sequence conversion on the image to be detected, the image processing model's feature extraction capability for the detailed parts of the image to be detected can be improved, the pre-trained feature extractor and the trained classification sub-model can be used to improve the analysis accuracy of the image processing model for the image to be detected, and the large-scale language analysis model can be used to perform object analysis on the target prediction results combined with clinical data, which can be analyzed based on multiple factors to improve the accuracy of the object analysis results. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0046] FIG1 is a diagram illustrating an application environment of an object analysis method according to an embodiment of the present application;

[0047] FIG2 is a schematic diagram of the training principle of a large-scale language analysis model in one embodiment of the present application;

[0048] FIG3 is a schematic diagram of a process for converting an image to be detected into an image sequence in one embodiment of the present application;

[0049] FIG4 is a schematic diagram of a process for analyzing and processing a target image sequence by an image processing model in one embodiment of the present application;

[0050] FIG5 is a schematic diagram of the structure of a feature extractor in one embodiment of the present application;

[0051] FIG6 is a schematic diagram of a process for determining a target feature vector by a feature extractor in one embodiment of the present application;

[0052] FIG7 is a schematic diagram of the model structure of each classification sub-model in one embodiment of the present application;

[0053] FIG8 is a flow chart of the training steps of a feature extractor in one embodiment of the present application;

[0054] FIG9 is a schematic diagram of a pre-trained network of an image processing model in one embodiment of the present application;

[0055] FIG10 is a schematic diagram of a process for determining contrast loss in another embodiment of the present application;

[0056] FIG11 is a flow chart of an example of an object analysis method according to an embodiment of the present application;

[0057] FIG12 is a structural block diagram of an object analysis device according to an embodiment of the present application;

[0058] FIG13 is a diagram showing the internal structure of a computer device in one embodiment of the present application. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0060] In one embodiment, as shown in FIG1 , an object analysis method is provided. This embodiment uses the method applied to a terminal as an example. It is understood that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes steps 102 through 106.

[0061] Step 102: Obtain target image sequence, clinical data, and image processing model.

[0062] The target image sequence is obtained by performing image sequence conversion on the image to be detected.

[0063] Among them, the image processing model includes a pre-trained feature extractor and a trained classification sub-model.

[0064] In an embodiment of the present application, a terminal receives user-input images to be tested and clinical data, and then obtains a trained image processing model. The images to be tested may be medical images. This embodiment of the present application uses fundus images used to detect diabetic retinopathy as an example. The terminal performs image sequence conversion on the images to be tested, converting them into a one-dimensional embedding sequence to obtain a target image sequence corresponding to the images to be tested.

[0065] Clinical data can be examination data of the subject (patient) to be tested, such as gender, age, height, weight, blood pressure, and other data obtained through examination. Multiple data types are preset in the clinical data. There may be missing data based on the clinical data provided by different subjects to be tested. In this case, the terminal can set the data type corresponding to the missing data to the default state in the preset data type, that is, all data with examination records can be included and saved in the form of a dictionary. The image processing model is a trained model used to perform image analysis on the image to be tested and determine the disease analysis results contained in the image to be tested.

[0066] Step 104 , extracting features from the target image sequence using a pre-trained feature extractor, and analyzing and processing the feature extraction results using the trained classification sub-model to obtain target analysis results.

[0067] In an embodiment of the present application, the terminal can extract key features from a one-dimensional target image sequence through a feature extractor of an image processing model. The feature extractor can be a method based on deep learning, for example, an attention mechanism model (Transformer), a convolutional neural network (CNN), or a long short-term memory network (LSTM). Then, the terminal inputs the result obtained by feature extraction into a classification sub-model, classifies the result of feature extraction through the classification sub-model, and uses the classification result as the target analysis result. Optionally, the classification sub-model serves as a downstream task of the image processing model. There can be multiple classification sub-models, each of which has different points of interest during training. Image analysis can be performed for different task types, and the analysis results for different dimensions are summarized to obtain the target analysis result.

[0068] Step 106: Perform object analysis on the target analysis results and clinical data based on the large-scale language analysis model to obtain object analysis results.

[0069] In an embodiment of the present application, the large-scale language analysis model can be based on the LLaMA large language model framework. The terminal fuses the target analysis results in text form and clinical data, and inputs the fused text data into the large-scale language analysis model. The large-scale language analysis model is used to process the data to obtain the object analysis results of the object to be detected.

[0070] The training process of the large-scale language analysis model is shown in Figure 2. The large-scale language analysis model includes additional network layers, including an adapter network embedded in the Transformer layer and a LoRA (Low-Rank Adaptation) module set in the LLaMA model bypass. The adapter is located after the feedforward network layer and before the residual connection. The adapter is a multilayer perceptron (MLP), which is responsible for reducing and increasing the dimensionality of the feature expression of the Transformer layer. The adapter method is based on effective parameter learning for the pre-trained large-scale language analysis model. Only a small number of trainable parameters need to be added to achieve the effect of full parameter training, which significantly saves time during the training process and improves time efficiency. Therefore, in practical applications, for specific tasks, the terminal does not need to re-fine-tune all the parameters of the pre-trained large-scale language analysis model. It only needs to save the adapter, while the other parameters of the pre-trained large-scale language analysis model remain in the original pre-trained state, thereby optimizing space efficiency.

[0071] In the process of training a pre-trained large-scale language analysis model based on LoRA, the dimension reduction matrix A and dimension increase matrix B in the LoRA module in the bypass are replaced. In the downstream task processing, only the dimension reduction matrix A and dimension increase matrix B are updated. One dimension reduction and one dimension increase operation are performed to simulate the intrinsic rank, ensuring that the input and output dimensions of the pre-trained large-scale language analysis model remain unchanged. The outputs of the dimension reduction matrix A and dimension increase matrix B are superimposed with the parameters of LLaMA. Specifically, the terminal uses a random Gaussian distribution to initialize the dimension reduction matrix A and the dimension increase matrix B with a zero matrix, so as to ensure that the matrix of this bypass is still a zero matrix at the beginning of training. When fine-tuning a pre-trained language model in the downstream task, the pre-trained large-scale language analysis model is updated, as shown below: h=W0x+ΔWx=W0x+BAx (1)

[0072] Where W0 is the parameter initialized by the pre-trained model, ΔW is the parameter to be updated, x is the input of the large-scale language analysis model, and h is the dimension of the output of the large-scale language analysis model. It can be seen that fine-tuning all the parameters of a large language model is very challenging with limited computing resources. Early research found that pre-trained language models have low "internal dimensionality," meaning that even random projections onto smaller subspaces can effectively learn during task adaptation. Therefore, the LoRA module can add a small parameter module to learn the variation ΔW. During training, W0 is fixed; only the dimension reduction matrix A and the dimension increase matrix B contain training parameters and can be varied. During inference, this variation simply needs to be re-embedded into the pre-trained large-scale language analysis model, without any delay. Therefore, in this embodiment, the terminal uses doctor-patient conversations as training samples to analyze the diagnosis and treatment plan for diabetic retinopathy using the pre-trained large-scale language analysis model. The doctor-patient conversations contain 500,000 paired data items of doctor diagnoses and patient information, including fundus image symptom descriptions and patient clinical data. Specifically, the dimensionality reduction matrix A and the dimensionality increase matrix B are removed from the model, and then the trained dimensionality increase matrix B′ and the dimensionality reduction matrix A′ obtained by training with the doctor-patient dialogue information as training samples are added. This can complete the personalized training for the output diagnosis and treatment plan for diabetic retinopathy, and improve the accuracy of the generated diagnosis and treatment plan.

[0073] In the above-mentioned object analysis method, by performing sequence conversion on the image to be detected, the image processing model's ability to extract features of the details in the image to be detected can be improved. Through the pre-trained feature extractor and the trained classification sub-model, the image processing model's analysis accuracy of the image to be detected can be improved. The target prediction results are combined with clinical data for object analysis through a large-scale language analysis model. It can perform analysis based on multiple factors and improve the accuracy of the object analysis results.

[0074] In an exemplary embodiment of the present application, before performing feature extraction on the image to be detected, in order to improve the effect of feature extraction, it is necessary to pre-process the image to be detected, as shown in FIG3 , step 102 includes steps 302 to 304. Among them:

[0075] Step 302 : segment the image to be detected according to a preset segmentation method to obtain a plurality of image blocks.

[0076] In the embodiment of the present application, the input sequence length of the Transformer layer is fixed, and it receives the feature embedding Z∈R as a sequence. L×C As input, where L is the length of the sequence and C is the hidden channel size, the terminal needs to obtain the image to be detected x∈R H×W×3 Convert to Z, where H is the height of the image to be detected and W is the width of the image to be detected. Specifically, the input sequence length L of the Transformer is:

[0077] Therefore, the terminal divides the image to be detected into image blocks.

[0078] Step 304: Perform a convolution operation on each image block to obtain a target image sequence.

[0079] In the embodiment of the present application, after the terminal segments the image to be detected, a two-dimensional convolution operation f:p→e∈R is performed on each image block. C , where p represents a vector image block and C is the dimension. The terminal further maps each vector image block p to a potential representation space of dimension C, i.e., generates a one-dimensional embedding sequence of the image to be detected. The convolution kernel size of the two-dimensional convolution operation can be 16×16, and the step size can be 16. Optionally, based on the different sizes of the images to be detected, the parameters of the two-dimensional convolution operation can be adaptively adjusted. Optionally, in order to encode the spatial information of each image block, the terminal can learn each position i (i=0, 1, 2…L) as a specific embedding p by using the truncated normal distribution method. i , and embed the i Add to e i , e iThe feature vector representing the image block at each position i contains the visual features of the image block, and p i Represents the spatial embedding vector of position i, which is used to capture the spatial information of position i. The final target image sequence can be expressed in the form of E = {e1+p1,e2+p2,…,e L +p L}.

[0080] In this embodiment, by performing image segmentation on the image to be detected, more comprehensive detail features are obtained for each image block when the image processing model performs feature extraction. At the same time, each image block is convolutionally mapped, and the final sequence obtained contains not only the visual features of the image block, but also the spatial information of each position of the image block. The image processing model processes the positional relationship between the image blocks to provide more accurate feature extraction results.

[0081] In an exemplary embodiment, as shown in FIG4 , the feature extractor includes a multi-head self-attention structure and a multi-layer perceptron, and step 104 includes steps 402 to 406 .

[0082] Step 402: Calculate triplet features of the target image sequence.

[0083] In the embodiment of the present application, as shown in FIG5 , the feature extractor of the image processing model includes 24 stacked Transformer layers. Each Transformer layer has a global receptive field and is composed of a multi-headed self-attention module (MSA) and a multi-layer perceptron (MLP). The input of each Transformer layer is from Z l-1 ∈R L×C The triplet (query vector Q, key vector K, value vector V) calculated in , that is: Q = Z l-1 W Q ,K=Z l-1 W K ,V=Z l-1 W V (3)

[0084] Among them, W Q 、W K 、W V ∈R C×d are the learnable weight parameters in the three linear mapping layers, d is the feature dimension of the triplet, and Z l-1The l in the figure represents the number of Transformer layers. Since the Transformer layer in this embodiment is 24, l = 1, 2, ... 24. Therefore, the terminal first calculates the triplet feature corresponding to the image to be detected, namely the query vector Q, the key vector K, and the value vector V.

[0085] Step 404: transform the triplet features based on the multi-head self-attention structure and the multi-layer perceptron to obtain the target feature vector.

[0086] In an embodiment of the present application, the terminal maps the triplet features to different self-attention layers through a multi-head self-attention structure, where each self-attention layer learns a specific focus point. In each attention layer, the correlation of the input features is calculated according to the self-attention mechanism to obtain a more discriminative feature representation. The terminal applies a multi-layer perceptron to transform the fused feature vector. MLP consists of multiple fully connected layers, each layer contains a set of weights and activation functions. By stacking multiple fully connected layers, MLP can nonlinearly change the representation of the feature vector, and can obtain a more compact and expressive target feature vector by operations such as reducing the dimension.

[0087] Step 406: classify the target feature vector according to the classification sub-model to obtain a target analysis result.

[0088] In an embodiment of the present application, the terminal transmits the feature vector extracted by the feature extractor to the molecular class model of the downstream task, classifies the target feature vector through the classification sub-model, and obtains the type of lesions that may exist in the image to be detected as the target analysis result.

[0089] In this embodiment, a multi-head self-attention mechanism and a multi-layer perceptron are used to perform conversion operations on the triple features of the image to be detected to obtain a target feature vector, which can improve the performance and accuracy of the image processing model and enhance the perception and recognition ability of the image processing model to analyze the lesions contained in the image to be detected.

[0090] In an exemplary embodiment, as shown in FIG6 , step 404 includes steps 602 to 604 .

[0091] Step 602: Map the triplet features according to each self-attention layer in the multi-head self-attention mechanism to obtain the mapping result of each self-attention layer, and obtain a joint mapping result based on each mapping result.

[0092] In the embodiment of the present application, each self-attention mechanism (SA) of the multi-head self-attention structure can be expressed as:

[0093] The terminal maps the joint output of the self-attention mechanism operation of each self-attention layer according to the multi-head self-attention structure, that is: MSA(Z l-1 )=[SA1(Z l-1 );SA2(Z l-1 );…;SA m (Z l-1 )]W O (5)

[0094] Among them, W O ∈R md×C , d is usually set to C / m, where C represents the number of channels and m represents the number of attention heads, representing the number of parallel heads used in the multi-head self-attention structure.

[0095] Step 604: transform the joint mapping result through a multi-layer perceptron to obtain a target feature vector.

[0096] In the embodiment of the present application, the terminal transforms the output of the multi-head self-attention structure through a multi-layer perceptron and uses the residual connection as the output of the layer: Z l =MSA(Z l-1 )+MLP(MSA(Z l-1 ))∈R L×C (6)

[0097] In addition, as shown in Figure 5, layer normalization is applied before MSA and MLP, and {Z 1 ,Z 2 ,…,Z l-1} as the feature of the Transformer layer, that is, the target feature vector.

[0098] In this embodiment, a multi-head self-attention mechanism and a multi-layer perceptron are used to perform conversion operations on the triple features of the image to be detected to obtain a target feature vector, which can improve the performance and accuracy of feature extraction of the image processing model.

[0099] In an exemplary embodiment, the classification sub-model includes a quality assessment sub-model, a first symptom grading sub-model, a second symptom judgment sub-model and a lesion segmentation sub-model, and step 406 includes steps 4061 to 4062.

[0100] Step 4061, classify the target feature vector according to the image quality assessment sub-model, the first symptom grading sub-model, the second symptom judgment sub-model and the lesion segmentation sub-model to obtain the image quality assessment result, the first symptom level, the second symptom diagnosis result and the lesion segmentation result.

[0101] In the embodiment of the present application, the classification and discrimination process is explained by taking the image to be detected as a fundus image as an example. The model structure of the quality assessment submodel, the first symptom grading submodel, the second symptom judgment submodel and the lesion segmentation submodel is shown in Figure 7. After the terminal obtains the target feature vector, the target feature vector is transmitted to the different classification submodels of the downstream task respectively. The fundus image is quality assessed by the quality assessment submodel to determine the image quality assessment result of the fundus image, wherein the quality assessment result of the fundus image can be used to reflect whether the current fundus image to be detected is usable. The first symptom grading submodel can be a DR grading submodel of the fundus image. The fundus image is analyzed by the DR grading submodel to obtain the patient's DR grading result (grade 0 to 5). The second symptom judgment submodel can be a diabetic macular edema (DME) classification model to obtain an analysis result of whether the fundus image has diabetic macular edema. The lesion segmentation submodel segments the lesions that may exist in the fundus image and obtains a lesion segmentation result, for example, whether the fundus image contains lesions such as microaneurysms, hemorrhages, soft infiltrations and hard infiltrations. After the classification sub-model of each downstream task outputs the analysis results of the fundus image, the terminal constructs each analysis result into a dictionary format, for example<key,value> The final classification result of the fundus image is formed in the form of.

[0102] Specifically, the image quality assessment submodel involves a binary classification task, using labeled sample images of different clarity as training samples for the image quality assessment submodel. After the input fundus image passes through the encoder, it is connected to a mean pooling layer and a linear layer, and then the probability of the predicted category is output. During the training of this image quality assessment submodel, the supervised loss of the input labeled sample image can be calculated using binary cross-entropy loss, thereby completing the training of the image quality assessment submodel. The first symptom grading submodel involves a five-category task, using labeled sample images of different DR levels as training samples for the first symptom grading submodel. After the input fundus image passes through the encoder, it is connected to a mean pooling layer and a linear layer, and then the probability of the predicted category is output. The training of the first symptom grading submodel can use cross-entropy loss to fine-tune the labeled sample images, resulting in a trained first symptom grading submodel. The second symptom assessment submodel involves a binary classification task, trained on labeled sample images containing and not containing DME. After the input image passes through the encoder, it is connected to a mean pooling layer and a linear layer to output the predicted category probability. Binary cross-entropy loss is used to fine-tune the labeled images, resulting in the trained second symptom assessment submodel. The lesion segmentation submodel involves a five-category segmentation task, trained on labeled sample images containing different lesions. After the input image passes through the encoder-decoder, the predicted mask image is output, which is fine-tuned using cross-entropy loss to obtain the trained lesion segmentation submodel.

[0103] Step 4062: The image quality assessment result, the first symptom level, the second symptom diagnosis result, and the lesion segmentation result are used as target analysis results.

[0104] In an embodiment of the present application, the terminal performs summary processing based on the image quality assessment result, the first symptom level, the second symptom diagnosis result and the lesion segmentation result to obtain the target analysis result.

[0105] In this embodiment, the image to be detected is analyzed by the downstream classification submodel of the quality assessment submodel, the first symptom grading submodel, the second symptom judgment submodel and the lesion segmentation submodel to obtain analysis results in multiple dimensions. Through the analysis results of the image to be detected in multiple dimensions, the data processing accuracy of the large-scale language analysis model can be improved, thereby improving the accuracy of the object analysis results.

[0106] In an exemplary embodiment, before using the image processing model, the image processing model needs to be trained. As shown in FIG8 , the process before step 102 includes steps 802 to 808 .

[0107] Step 802: Acquire a sample image set, and perform data enhancement on each sample image in the sample image set to obtain a first reconstructed image block and a second reconstructed image block.

[0108] In an embodiment of the present application, the pre-trained network of the image processing model adopts a dual self-supervised network (DSN, Dual-Supervised Nets) method to maximize the learning of unlabeled fundus image features. Its structure is shown in Figure 9. By reconstructing the missing image blocks in the local view and measuring the similarity of positive and negative samples in the global view, it includes an image block generation step, two encoders and a decoder. Among them, the image block generation step is to divide the input unlabeled sample fundus image into regular non-overlapping image blocks, and then the terminal samples a part of the image blocks and masks the remaining image blocks (see the gray image blocks in Figure 9), and sets the masking ratio to 0.75. The encoders are E q and E k , both encoders are based on the Transformer structure, let θ k and θ q Respectively represent E k and E q The parameters of θ are calculated by adopting the following momentum update strategy k : θ k ←mθ k +(1-m)θ q (7)

[0109] Among them, m∈[0,1) is the momentum coefficient, which can be 0.999. It should be noted that only the parameter θ q Updated by back propagation. From the momentum update strategy in the above equation, we can see that θ k The change ratio θ q Smoother.

[0110] In addition, the decoder uses Transformer as the decoder structure. Since Transformer has fewer layers, it can reduce training time.

[0111] The terminal performs two data enhancements on each sample image, as shown in Figure 9, to obtain the first reconstructed image block x q and the second reconstructed image block x k .

[0112] Step 804 : performing feature extraction on the first reconstructed image block and the second reconstructed image block respectively to obtain query features and a key feature sequence, and determining a contrast loss value based on the query features and the key feature sequence.

[0113] In an embodiment of the present application, the terminal performs feature extraction on the first reconstructed image block and the second reconstructed image block according to different encoders, obtains the query feature corresponding to the first reconstructed image block, that is, the query feature q in Figure 9, and obtains the key feature sequence corresponding to the second reconstructed image block, that is, the key feature sequence composed of the key feature k in Figure 9, and calculates the contrast loss value based on the query feature and each key feature in the key feature sequence.

[0114] Step 806 : reconstruct the reconstructed image block according to a preset decoder to obtain a reconstructed image, and determine a reconstruction loss value based on the reconstructed image and a sample image corresponding to the reconstructed image.

[0115] In the embodiment of the present application, as shown in FIG9 , the terminal decodes and reconstructs the reconstructed image through the encoder and decoder respectively to obtain a reconstructed image. Then, the terminal calculates the reconstruction loss value based on the difference between the original sample image and the reconstructed image. In the encoder-decoder branch for image reconstruction, the terminal uses the mean square error (MSE) between the reconstructed image and the original fundus image as the reconstruction loss. Finally, for an unlabeled fundus image, the unsupervised loss The calculation method is:

[0116] Among them, w1 and w2 are two learnable parameters used to weight and Their initial values ​​are empirically set to 0.4 and 0.6.

[0117] Step 808: Train the feature extractor included in the image processing model based on the contrast loss value and the reconstruction loss value to obtain a pre-trained feature extractor.

[0118] In an embodiment of the present application, the terminal forward-propagates a sample image through a pre-trained network to obtain an output result from the trained network, calculates a comparative loss value based on the output result, and obtains a reconstruction loss value based on the output result of the pre-trained network and the original sample image. The terminal then obtains a total loss based on the comparative loss value and the reconstruction loss value, and uses the total loss to perform backpropagation and update the network parameters of the pre-trained network until the pre-trained network meets preset training conditions, thereby obtaining a pre-trained feature extractor. The preset training conditions may be satisfying a preset number of iterations, or the total loss being lower than a preset loss threshold.

[0119] In this embodiment, by training the feature extractor based on the contrast loss value and the reconstruction loss value, the image processing model's ability to extract image features can be improved, the feature extractor's generalization ability to extract features for images to be detected of different qualities can be improved, the accuracy of the target image sequence can be improved, and the accuracy of the object analysis results can be improved.

[0120] In an exemplary embodiment, as shown in FIG. 10 , step 804 includes steps 1002 to 1006 .

[0121] Step 1002: For each first reconstructed image block and second reconstructed image block in the sample image set, feature extraction is performed on the first reconstructed image block according to the first encoder to obtain query features, and feature extraction is performed on the second reconstructed image block according to the second encoder to obtain key features.

[0122] In the embodiment of the present application, as shown in FIG9 , the terminal reconstructs the first image block x q Pass in an encoder E q , through a pooling layer and a normalization (BN) layer, the query feature q is generated. At the same time, the terminal passes the second reconstructed image block into the momentum encoder E k , generate a key feature f k , and then adds it to an initially defined dictionary queue. As iterations proceed, this queue generates a set of encoder key features {k0, k1, k2, ...}. The terminal then uses a global average pooling layer and a normalization layer to fix the dimensions of the query feature q and the key feature k.

[0123] Step 1004, match each key feature in the key feature sequence with the query feature, and determine the positive sample key feature in the key feature sequence based on the matching result, and determine the comparison loss value through the query feature and the positive sample key feature in the key feature sequence, as well as other key features in the key feature sequence.

[0124] In this embodiment of the present application, there is a single key feature (denoted as k+) in the dictionary that is used to match the query feature q. The terminal then uses the query feature q as the query feature, k+ as the positive sample, and all other key features in the dictionary as negative samples. Therefore, the contrast loss of the terminal for the unlabeled image x is The calculation of is achieved by making the query feature q similar to the key feature k+ as the positive sample and dissimilar to other key features, that is:

[0125] Where τ is a temperature hyperparameter, and its value is empirically set to 0.07. The length of the queue K can be set to 16,384. The summation in the above equation is for one positive sample and K negative samples. Intuitively, is a logarithmic loss based on a (K+1)-class softmax classifier, trying to classify q as k+.

[0126] In this embodiment, the sample images in the sample image set that serve as positive samples are determined through unsupervised contrast loss, without the need to label the sample images. This allows the feature extractor to be learned for a large number of sample images, thereby improving the learning efficiency of the feature extractor.

[0127] In an exemplary embodiment, as shown in FIG11 , in the first stage, the terminal obtains the patient's left and right eye fundus images, and predicts the left and right eye fundus images respectively using the TransDR model to obtain fundus image prediction results. In one example, the prediction result may be: "{Eye examination: [{left or right eye (0-left; 1-right): 1, DR grade (0-1-2-3-4-5): 2, DME grade (0-1)): 0}, {left or right eye (0-left; 1-right): 0, DR grade (0-1-2-3-4): 2, DME grade (0-1): 0]}".

[0128] In the second stage, the terminal obtains the patient's clinical data. In an example, the patient's clinical data may be: "{Gender: Male, Date of Birth: 1977 / 8 / 27 0:00, Examination Date: 2018 / 11 / 20, Age: 41, Systolic Blood Pressure: 157, Diastolic Blood Pressure: 100, Heart Rate: 112, Smoking: No smoking, Drinking: No drinking, Blood Pressure History: High blood pressure on physical examination, normal blood pressure on self-measurement, Circulatory System History: No chest tightness, chest pain or other discomfort, Family History: Father has hypertension and diabetes, Dietary habits: Balanced meat and vegetables, Exercise type: Physical exercise frequency: Walking 15,000 times a day First, history of endocrine system diseases: diagnosed with type 2 diabetes at around 36 years old, currently on diet control, methotrexate, and insulin treatment. Regular monitoring of fasting blood sugar is around 9mmol / l, blood sugar 2 hours after meal is around 11mmol / l, serum triglycerides: 2.8, total cholesterol: 6.56, high-density lipoprotein: 1.3, low-density lipoprotein: 4.3, eGFR: 110.6508, fasting serum glucose: 11.91, glycosylated hemoglobin: 9.9, aspartate aminotransferase: 28.1, alanine aminotransferase: 39.6, gamma-glutamyl transpeptidase (GGT): 65}".

[0129] In the third stage, the terminal integrates the fundus image prediction results with the patient's clinical data to form the consultation content, which serves as input information for the large-scale language analysis model.

[0130] In the fourth stage, the terminal inputs the consultation content and language model prompts into the large-scale language analysis model, which analyzes and processes the consultation content through the large-scale language analysis model to obtain a personalized diagnosis and treatment plan for the current patient. Among them, the large-scale language analysis model is a large language model that has been fine-tuned for model parameters based on a large amount of consultation data. The language model prompt represents the task identifier of the large-scale language analysis model, which is used to indicate the task type of data processing by the large-scale language analysis model. In one example, the personalized diagnosis and treatment plan output by the large-scale language analysis model can be: "For this patient, according to the Chinese Diabetes Prevention and Treatment Guidelines, I recommend: 1. In terms of blood sugar control, it is recommended to continue active treatment, strengthen diet control and exercise, and adjust the insulin dosage according to blood sugar levels. The recommended medication regimen is: maintain the original dosage of insulin Humulin in the morning and evening, and consider adding oral medications such as insulin-like or α-glucosidase inhibitors to lower blood sugar levels as much as possible. It is recommended to monitor blood sugar regularly to achieve the target range for reasonable control of diabetes. 2. In terms of blood lipid management, it is recommended to increase the intake of foods rich in dietary fiber and w-3 fatty acids, and actively engage in aerobic exercise, control weight, stop smoking, limit drinking and other bad habits. . The recommended drug treatment plan is: oral statins, such as atorvastatin, to control total cholesterol and low-density lipoprotein levels. 3. In terms of blood pressure management, it is recommended to continue taking lisinopril, levamlodipine and other drugs to control blood pressure within the target range. It is recommended to monitor blood pressure regularly and adjust medication. 4. Pay attention to eye examinations, perform regular retinal examinations, and deal with diabetic retinopathy in a timely manner. 5. Maintain a positive mental state, avoid mood swings, maintain a stable living habits, and at the same time follow a healthy diet and lifestyle, increase exercise, and control weight appropriately. I hope these suggestions are helpful to you. If you have any questions or need more details, please feel free to consult a doctor. This patient advocates annual fundus follow-up and ophthalmological follow-up when eye discomfort occurs."

[0131] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0132] Based on the same inventive concept, embodiments of the present application also provide an object analysis device for implementing the object analysis method described above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more object analysis device embodiments provided below can be found in the above-mentioned limitations of the object analysis method and will not be repeated here.

[0133] In an exemplary embodiment, as shown in FIG12 , an object analysis device 1200 is provided, including: an acquisition module 1201 , a first analysis module 1202 , and a second analysis module 1203 .

[0134] The acquisition module 1201 is used to obtain the target image sequence, clinical data and image processing model; the target image sequence is obtained by converting the image to be detected into an image sequence; the image processing model includes a pre-trained feature extractor and a trained classification sub-model.

[0135] The first analysis module 1202 is configured to extract features from a target image sequence using a feature extractor, and analyze and process the feature extraction results using a classification sub-model to obtain target analysis results.

[0136] The second analysis module 1203 is used to perform object analysis on the target analysis results and clinical data based on the large-scale language analysis model to obtain object analysis results.

[0137] In one embodiment, the acquisition module 1201 is specifically configured to perform image segmentation on the image to be detected according to a preset segmentation method to obtain a plurality of image blocks; and perform a convolution operation on each image block to obtain a target image sequence.

[0138] In one embodiment, the feature extractor includes a multi-head self-attention structure and a multi-layer perceptron.

[0139] The first analysis module 1202 is specifically used to calculate the triple features of the target image sequence; transform the triple features based on the multi-head self-attention structure and the multi-layer perceptron to obtain the target feature vector; classify the target feature vector according to the classification sub-model to obtain the target analysis result.

[0140] In one embodiment, the first analysis module 1202 is specifically used to map the triple features according to each self-attention layer in the multi-head self-attention mechanism, obtain the mapping result of each self-attention layer, and obtain a joint mapping result based on each mapping result; the joint mapping result is transformed through a multi-layer perceptron to obtain a target feature vector.

[0141] In one embodiment, the classification sub-model includes a quality assessment sub-model, a first symptom grading sub-model, a second symptom judgment sub-model and a lesion segmentation sub-model.

[0142] The first analysis module 1202 is specifically used to classify the target feature vector according to the image quality assessment sub-model, the first symptom grading sub-model, the second symptom judgment sub-model and the lesion segmentation sub-model, and obtain the image quality assessment results, the first symptom level, the second symptom diagnosis results and the lesion segmentation results; and use the image quality assessment results, the first symptom level, the second symptom diagnosis results and the lesion segmentation results as the target analysis results.

[0143] In one embodiment, the apparatus 1200 further includes: a data enhancement module, a feature extraction module, a reconstruction module and a training module.

[0144] The data enhancement module is used to obtain a sample image set and perform data enhancement on each sample image in the sample image set to obtain a first reconstructed image block and a second reconstructed image block.

[0145] The feature extraction module is used to extract features from the first reconstructed image block and the second reconstructed image block respectively to obtain query features and key feature sequences, and determine the contrast loss value based on the query features and the key feature sequences.

[0146] The reconstruction module is used to reconstruct the reconstructed image block according to the preset decoder to obtain a reconstructed image, and determine the reconstruction loss value based on the reconstructed image and the sample image corresponding to the reconstructed image.

[0147] The training module is used to train the feature extractor included in the image processing model based on the contrast loss value and the reconstruction loss value to obtain a pre-trained feature extractor.

[0148] In one embodiment, the feature extraction module is specifically used to perform feature extraction on the first reconstructed image block according to the first encoder to obtain a query feature, and to perform feature extraction on the second reconstructed image block according to the second encoder to obtain a key feature for each first reconstructed image block and the second reconstructed image block in the sample image set; match each key feature in the key feature sequence with the query feature, and determine the positive sample key feature in the key feature sequence based on the matching result, and determine the contrast loss value through the query feature and the positive sample key feature in the key feature sequence, as well as other key features in the key feature sequence.

[0149] Each module in the object analysis device described above may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0150] In an exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be shown in Figure 13. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store target image sequences and clinical data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an object analysis method is implemented.

[0151] Those skilled in the art will understand that the structure shown in FIG13 is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.

[0152] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0153] In one embodiment, a non-volatile computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0154] In one embodiment, a computer program product is provided, comprising executable instructions, which implement the steps of the above method embodiments when executed by a processor.

[0155] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0156] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0157] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0158] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. An object analysis method, characterized in that, The method includes: Obtaining a target image sequence, clinical data, and an image processing model; the target image sequence is obtained by transforming a to-be-detected image into an image sequence; the image processing model includes a pre-trained feature extractor and a trained classification sub-model; Performing feature extraction on the target image sequence through the feature extractor, and performing analysis and processing on the result of the feature extraction through the classification sub-model to obtain a target analysis result; Performing object analysis on the target analysis result and the clinical data based on a large-scale language analysis model to obtain an object analysis result.

2. The method according to claim 1, characterized in that, The obtaining of the target image sequence includes: Performing image segmentation on the to-be-detected image according to a preset segmentation method to obtain a plurality of image blocks; Performing a convolution operation on each of the image blocks to obtain a target image sequence.

3. The method according to claim 1 or 2, characterized in that The feature extractor includes a multi-head self-attention structure and a multi-layer perceptron; The performing of feature extraction on the target image sequence through the feature extractor, and the performing of analysis and processing on the result of the feature extraction through the classification sub-model to obtain the target analysis result includes: Calculating the triple features of the target image sequence; Transforming the triple features based on the multi-head self-attention structure and the multi-layer perceptron to obtain a target feature vector; Performing classification processing on the target feature vector according to the classification sub-model to obtain a target analysis result.

4. The method according to claim 3, wherein The transforming of the triple features based on the multi-head self-attention structure and the multi-layer perceptron to obtain a target feature vector includes: Mapping the triple features according to each self-attention layer in the multi-head self-attention mechanism to obtain the mapping result of each self-attention layer, and obtaining a joint mapping result based on the mapping results; Transforming the joint mapping result through the multi-layer perceptron to obtain the target feature vector.

5. The method according to claim 3 or 4, characterized in that, The classification sub-model includes a quality assessment sub-model, a first symptom grading sub-model, a second symptom judgment sub-model, and a lesion segmentation sub-model; The performing of classification processing on the target feature vector according to the classification sub-model to obtain a target analysis result includes: Performing classification processing on the target feature vector according to the image quality assessment sub-model, the first symptom grading sub-model, the second symptom judgment sub-model, and the lesion segmentation sub-model respectively to obtain an image quality assessment result, a first symptom grade, a second symptom diagnosis result, and a lesion segmentation result; Taking the image quality assessment result, the first symptom grade, the second symptom diagnosis result, and the lesion segmentation result as the target analysis result.

6. The method according to claim 1, characterized in that Before obtaining the target image sequence, clinical data, and image processing model, the method further includes: Obtaining a sample image set, and performing data augmentation on each sample image in the sample image set to obtain a first reconstructed image block and a second reconstructed image block; Performing feature extraction on the first reconstructed image block and the second reconstructed image block respectively to obtain a query feature and a key feature sequence, and determining a contrast loss value based on the query feature and the key feature sequence; Reconstruct the reconstructed image block according to a preset decoder to obtain a reconstructed image, and determine a reconstruction loss value based on the reconstructed image and the corresponding sample image of the reconstructed image; Determine a reconstruction loss value corresponding to the sample image based on the reconstructed image; Train a feature extractor included in the image processing model based on the contrast loss value and the reconstruction loss value to obtain the pre-trained feature extractor.

7. The method according to claim 6, wherein The extracting features from the first reconstructed image block and the second reconstructed image block respectively to obtain a query feature and a key feature sequence, and determining a contrast loss value based on the query feature and the key feature sequence includes: For each of the first reconstructed image block and the second reconstructed image block in the sample image set, extract features from the first reconstructed image block according to a first encoder to obtain a query feature, and extract features from the second reconstructed image block according to a second encoder to obtain a key feature; Match each key feature in the key feature sequence with the query feature, determine a positive sample key feature in the key feature sequence based on the matching result, and determine a contrast loss value through the query feature, the positive sample key feature in the key feature sequence, and other key features in the key feature sequence.

8. The method according to claim 1, characterized in that, The target image sequence is a one-dimensional embedding sequence.

9. The method according to claim 1, wherein The large-scale language analysis model includes: An adapter, which is embedded in a Transformer layer, after a feed-forward network layer and before a residual connection. The adapter is a multi-layer perceptron (MLP) for performing dimensionality reduction and dimensionality increase operations on the feature representation of the Transformer layer; A LoRA (Low-Rank Adaptation) module, which is set on the bypass of the LLaMA model.

10. The method according to claim 9, wherein: The LoRA module includes a dimensionality reduction matrix A and a dimensionality increase matrix B; Before performing object analysis on the target analysis result and the clinical data based on the large-scale language analysis model to obtain an object analysis result, the method includes: training the large-scale language analysis model; Wherein, the training the large-scale language analysis model includes: Initialize the dimensionality reduction matrix A with a random Gaussian distribution and initialize the dimensionality increase matrix B with a zero matrix; Fine-tune the large-scale language analysis model, and update the large-scale language analysis model to: h = W0x + ΔWx = W0x + BAx Wherein, W0 is the initialized parameter of the large-scale language analysis model, which is fixed and unchanged, ΔW is the parameter to be updated, x is the input of the large-scale language analysis model, h is the dimension of the output of the large-scale language analysis model, and the dimensionality reduction matrix A and the dimensionality increase matrix B include training parameters.

11. The method according to claim 2, wherein: The multiple image blocks are vector image blocks; The performing a convolution operation on each of the image blocks to obtain a target image sequence includes: Performing a two-dimensional convolution operation on each of the vector image blocks to generate a one-dimensional embedding sequence of the image to be detected.

12. An object analysis device, characterized in that, The device includes: An acquisition module, configured to acquire a target image sequence, clinical data, and an image processing model; the target image sequence is obtained by performing image sequence conversion on the image to be detected; the image processing model includes a pre-trained feature extractor and a trained classification sub-model; A first analysis module, configured to perform feature extraction on the target image sequence through the feature extractor, and perform analysis and processing on the result of the feature extraction through the classification sub-model to obtain a target analysis result; A second analysis module, configured to perform object analysis on the target analysis result and the clinical data based on a large-scale language analysis model to obtain an object analysis result.

13. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.

15. A computer program product comprising executable instructions, characterized in that, When the executable instruction is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Image processing method and device, computer equipment and storage medium

    CN114387270A

  • Training method and device of power grid time sequence data feature extraction model

    CN116028785A

  • Video target segmentation method and device, computer equipment and storage medium

    CN116385947A

  • Target detection method and device, computer equipment and storage medium

    CN116912791A

  • Training method, application method and system of grape embryo auxiliary inspection model

    CN117012373A

Cited By

  • Regional ASF image data processing method and system based on multi-modal data fusion

    CN120894663A

  • Helicobacter pylori fluorescence image fine tuning detection method and device and readable storage medium thereof

    CN121147226A

  • Intestinal mucosa tissue pathology image classification method and system based on artificial intelligence

    CN121259416A

  • Method and device for predicting early Alzheimer's disease based on dynamic and static SFC characteristics

    CN121460182A

  • A method and device for predicting early alzheimer's disease based on dynamic and static sfc features

    CN121460182B