Clinical data analysis method and device, computer equipment and storage medium

Through the method of feature classification and feature vector processing on clinical data, the problems of incomplete data information and interpretation error in the prior art are solved, and the accuracy and efficiency of clinical data analysis are improved.

CN119993353APending Publication Date: 2025-05-13PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411304271.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the existing case analysis methods, incomplete data source information or interpretation errors lead to insufficient utilization of analytical data, low analysis accuracy, and high manual analysis cost.

Method used

By obtaining the medical clinical data of the current disease type, the feature classification processing of text and picture data is performed, the data is processed with the preset pathological feature vector processing model, and feature integration is performed to determine the classification results of the disease type.

Benefits of technology

It improves the meticulousness of medical clinical data, enhances the model's ability to judge the characteristics of the disease, improves the accuracy and efficiency of clinical data analysis, and reduces the cost of manual analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993353A_ABST
    Figure CN119993353A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the field of artificial intelligence and image recognition, and relates to a clinical data analysis method, which comprises the following steps: acquiring medical clinical data corresponding to a current disease; performing feature classification processing on the text clinical data and the picture clinical data to obtain text clinical sub-data and picture clinical sub-data corresponding to feature classification; respectively carrying out feature vector processing on the text clinical sub-data and the picture clinical sub-data of the corresponding feature classification through a preset pathological feature vector processing model to obtain a text feature vector and a picture feature vector corresponding to the current disease category; performing feature integration processing on the picture feature vector and the text feature vector to obtain a medical feature vector of the current disease; and determining at least one classification result of the current disease based on the medical feature vector. By deepening the hierarchical structure of the medical clinical data, the fine degree of the medical clinical data is improved, and the accuracy of clinical data analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence and image recognition technology, and in particular to clinical data analysis methods, devices, computer equipment and storage media. Background Art

[0002] In modern medical practice, accurately identifying diseases and developing effective treatment plans rely on a detailed analysis of patients’ clinical data. These data usually include clinical records and image data in text form, and doctors are required to integrate this information for disease analysis, a process that relies on doctors’ expertise and experience.

[0003] However, these methods have some limitations. For example, they cannot automatically and accurately fuse clinical text and image data. Therefore, existing case analysis methods may have incomplete information or interpretation errors on the data source, which may lead to insufficient utilization of analysis data, low analysis accuracy, and high manual analysis costs. Summary of the invention

[0004] The purpose of the embodiments of the present application is to propose a clinical data analysis method, apparatus, computer equipment and storage medium to solve the problems in existing case analysis methods, such as incomplete information or interpretation errors that may appear on the data source, resulting in insufficient utilization of analysis data, low analysis accuracy, and high manual analysis costs.

[0005] In order to solve the above technical problems, the present application embodiment provides a clinical data analysis method, which adopts the following technical solution:

[0006] Obtaining medical clinical data corresponding to the current disease, wherein the medical data includes text clinical data and image clinical data;

[0007] Performing feature classification processing on the text clinical data and the image clinical data to obtain text clinical sub-data and image clinical sub-data corresponding to the feature classification;

[0008] By using a preset pathological feature vector processing model, feature vector processing is performed on the clinical sub-data of the image corresponding to the feature classification to obtain the image feature vector corresponding to the current disease type;

[0009] Performing feature vector processing on the text clinical sub-data of the corresponding feature classification by using a preset pathological feature vector processing model to obtain a text feature vector corresponding to the current disease type;

[0010] Performing feature integration processing on the image feature vector and the text feature vector to obtain a medical feature vector of the current disease;

[0011] Based on the medical feature vector, at least one classification result of the current disease is determined.

[0012] Furthermore, the text clinical sub-data includes title-level text clinical sub-data, paragraph-level text clinical sub-data and sentence-level text clinical sub-data, and the step of performing feature classification processing on the text clinical data to obtain text clinical sub-data corresponding to the feature classification includes:

[0013] Performing text parsing on the text clinical data to determine the text content corresponding to the title level, paragraph level, and sentence level in the text clinical data;

[0014] According to the hierarchical coding method, semantic features are extracted from the text contents corresponding to the title level, paragraph level and sentence level to obtain the clinical sub-data of the title level text, the clinical sub-data of the paragraph level text and the clinical sub-data of the sentence level text.

[0015] Furthermore, the step of performing feature classification processing on the image clinical data to obtain the image clinical sub-data corresponding to the feature classification also includes:

[0016] Dividing the image clinical data by a preset segmentation size to obtain a plurality of image data of the same size;

[0017] Performing image feature clustering on the plurality of image data of the same size by using a clustering algorithm to obtain a plurality of clustered image data;

[0018] The plurality of clustered image data are used as the image clinical sub-data.

[0019] Furthermore, the step of performing feature vector processing on the text clinical sub-data of the corresponding feature classification by using a preset pathological feature vector processing model to obtain a text feature vector corresponding to the current disease type specifically includes:

[0020] Based on the number of characters in each sentence in the sentence-level text clinical sub-data, performing text feature vector processing on the plurality of sentence-level text clinical sub-data to obtain a plurality of first text feature vectors;

[0021] Aggregating multiple first text feature vectors at the current paragraph level by a weighted average pooling algorithm to obtain a second text feature vector, and repeating the steps of the weighted average pooling algorithm to obtain multiple second text feature vectors;

[0022] Based on the number of paragraph-level text clinical sub-data, performing text feature vector processing on the title-level text clinical sub-data to obtain a plurality of third text feature vectors;

[0023] Determine the weight values ​​corresponding to the first text feature vector, the second text feature vector, and the third text feature vector;

[0024] Based on the corresponding weight values, feature vector connection is performed on the first text feature vector, the second text feature vector and the third text feature vector to obtain a text feature vector corresponding to the current disease type.

[0025] Furthermore, the step of performing feature vector processing on the clinical sub-data of the picture corresponding to the feature classification by using a preset pathological feature vector processing model to obtain the picture feature vector corresponding to the current disease type specifically includes:

[0026] Extracting lesion features corresponding to a plurality of clustered image data from the clinical sub-data of the image;

[0027] Performing vector processing on the lesion features to obtain a plurality of corresponding lesion feature vectors;

[0028] Based on the multiple corresponding lesion feature vectors, determine the image feature vector corresponding to the current disease type.

[0029] Furthermore, the step of performing feature integration processing on the image feature vector and the text feature vector to determine at least one classification result of the current disease type includes:

[0030] Through the trained image-text joint model, feature integration processing is performed on the image feature vector and the text feature vector to determine the medical feature vector of the current disease;

[0031] Based on the medical feature vector, at least one classification result of the current disease is determined.

[0032] Furthermore, before the step of performing feature integration processing on the image feature vector and the text feature vector by the trained image-text joint model to determine the medical feature vector of the current disease, the method further includes:

[0033] Obtain the training text clinical data, training image clinical data and the image-text joint model to be trained for the corresponding disease;

[0034] Marking the training text clinical data and the training image clinical data with feature vector labels to obtain a first training data set;

[0035] Performing feature vector fusion processing on the training text clinical data and the training image clinical data marked by the feature vector label in the first training data set, so that the medical feature vectors fused from the training text clinical data and the training image clinical data marked by the feature vector label are cross-referenced to obtain a second training data set;

[0036] The image-text joint model to be trained is iteratively trained based on the second training data set, and the trained image-text joint model is obtained after the iterative training is completed.

[0037] In order to solve the above technical problems, the present application also provides a clinical data analysis device, which adopts the following technical solution:

[0038] A first acquisition module is used to acquire medical clinical data corresponding to the current disease, wherein the medical data includes text clinical data and image clinical data;

[0039] A second acquisition module is used to perform feature classification processing on the text clinical data and the image clinical data to obtain text clinical sub-data and image clinical sub-data corresponding to the feature classification;

[0040] The first processing module is used to perform feature vector processing on the clinical sub-data of the picture corresponding to the feature classification through a preset pathological feature vector processing model to obtain the picture feature vector corresponding to the current disease type;

[0041] The second processing module is used to perform feature vector processing on the text clinical sub-data of the corresponding feature classification through a preset pathological feature vector processing model to obtain a text feature vector corresponding to the current disease type;

[0042] An integration processing module is used to perform feature integration processing on the image feature vector and the text feature vector to obtain a medical feature vector of the current disease;

[0043] The classification module is used to determine at least one classification result of the current disease based on the medical feature vector.

[0044] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:

[0045] A computer device comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the clinical data analysis method when executing the computer-readable instructions.

[0046] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0047] A computer-readable storage medium, characterized in that computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the clinical data analysis method are implemented.

[0048] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0049] The embodiment of the present application obtains the medical clinical data corresponding to the current disease; performs feature classification processing on the text clinical data and the image clinical data to obtain the text clinical sub-data and the image clinical sub-data of the corresponding feature classification; performs feature vector processing on the text clinical sub-data and the image clinical sub-data of the corresponding feature classification respectively through a preset pathological feature vector processing model to obtain the text feature vector and the image feature vector corresponding to the current disease; performs feature integration processing on the image feature vector and the text feature vector to determine at least one classification result of the current disease. The present application improves the refinement of the medical clinical data by deepening the hierarchical structure of the medical clinical data, so that the model can judge the characteristics of the current disease based on more refined medical clinical data, thereby improving the accuracy of clinical data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the scheme in the present application, a brief introduction is given below to the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0051] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;

[0052] Figure 2 A flowchart according to an embodiment of a clinical data analysis method of the present application;

[0053] Figure 3 is a schematic structural diagram of an embodiment of a clinical data analysis device according to the present application;

[0054] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by technicians in the technical field of the present application; the terms used in the specification of the application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application; the terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of the present application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0056] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0057] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0058] like Figure 1 As shown, the system architecture 100 may include a terminal device 101, a network 102 and a server 103. The terminal device 101 may be a laptop 1011, a tablet computer 1012 or a mobile phone 1013. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links or optical fiber cables.

[0059] The user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0060] The terminal device 101 can be any electronic device with a display screen and supporting web browsing. In addition to a laptop computer 1011, a tablet computer 1012 or a mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV), a laptop computer, a desktop computer, etc.

[0061] The server 103 may be a server that provides various services, such as a background server that provides support for a web page displayed on the terminal device 101 .

[0062] It should be noted that the clinical data analysis method provided in the embodiment of the present application is generally executed by a server / terminal device, and accordingly, the clinical data analysis device is generally set in the server / terminal device.

[0063] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0064] Continue to refer Figure 2 , shows a flow chart of an embodiment of a method for clinical data analysis according to the present application. The clinical data analysis method comprises the following steps:

[0065] Step S201, obtaining medical clinical data corresponding to the current disease.

[0066] In this embodiment, the clinical data analysis method can be deployed in a clinical data analysis system. The clinical data analysis system can be constructed by a server or a server cluster. The server or server cluster can be any electronic device with functions such as (text, image recognition, text, image processing, data transmission, data storage), and the clinical data analysis method runs on the electronic device (for example Figure 1 The server / terminal device shown in the figure can obtain the above medical clinical data through a wired connection or a wireless connection. It should be noted that the above wireless connection method may include but is not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultrawideband) connection, and other wireless connection methods currently known or to be developed in the future.

[0067] Specifically, the above-mentioned medical data may include, but are not limited to, text clinical data and image clinical data. The above-mentioned text clinical data may include detailed descriptive content, such as the patient's existing medical history and previous medical records, basic demographic information (such as age, gender), and laboratory test results. It should be noted that demographic data and laboratory results are generally regarded as structured chief complaint content. For unstructured chief complaint content, a maximum length limit can be set, such as 40 characters. In an operation example, if the patient's description exceeds this length limit, the system only retains the content up to the limit; if it does not exceed, the system uses zero padding to reach the predetermined length. In addition, laboratory test results can also be processed by standardizing the test data of different laboratories for the same disease to calculate the average test result of the disease and use it as the standard laboratory result of the disease.

[0068] The above-mentioned image clinical data may include image scans of characteristic features associated with a specific disease. These images may be analyzed and processed by an image recognition model equipped with pathological image processing, feature extraction, and feature tracking functions to scan and process the collected disease-related images.

[0069] In a possible embodiment, the clinical data analysis system processes and analyzes the text and images input by the user or patient, converts them into target text clinical data and target image data, creates files based on the user or patient, and stores them in a database corresponding to the file.

[0070] Step S202, performing feature classification processing on the text clinical data and the image clinical data to obtain text clinical sub-data and image clinical sub-data corresponding to the feature classification.

[0071] In this embodiment, the classification process may refer to a process of obtaining more refined text clinical sub-data by performing hierarchical coding on the text clinical data. For example, the text clinical data may be divided into three levels: title, paragraph, and sentence, to obtain three corresponding hierarchical structures to achieve multi-level text representation.

[0072] Alternatively, the pathological features of the segmented images may be acquired by segmenting the clinical data of the images and using a local attention mechanism and a local feature extraction method on the segmented images.

[0073] The above-mentioned text clinical sub-data may be data information with specific classification characteristics extracted or derived from the text clinical data after the above-mentioned feature classification processing. The data information with specific classification characteristics has the function of focusing on specific information points. Specifically, the text clinical sub-data after feature classification processing through the title may have the ability to summarize the title, for example, summarizing the characteristic content of the current disease based on text analysis of the title.

[0074] The above-mentioned image clinical sub-data may be data information with specific region classification characteristics extracted or derived from the image clinical data after the above-mentioned feature classification processing, and the data information with specific region classification characteristics may be used to describe the pathological characteristics in the current region.

[0075] In a possible embodiment, the case analysis system performs feature classification processing on the text clinical data and image clinical data input by the user or patient to obtain text clinical sub-data and image clinical sub-data of corresponding feature classification, and stores them in the database corresponding to the archive based on the existing archive.

[0076] Step S203, performing feature vector processing on the clinical sub-data of the image corresponding to the feature classification by using a preset pathological feature vector processing model to obtain the image feature vector corresponding to the current disease type;

[0077] Step S204: perform feature vector processing on the text clinical sub-data of the corresponding feature classification through a preset pathological feature vector processing model to obtain a text feature vector corresponding to the current disease type.

[0078] In this embodiment, the above-mentioned preset pathological feature vector processing model can be used to convert the above-mentioned text clinical sub-data and image clinical sub-data into corresponding feature vectors through the mapping relationship between vectors and data. It should be noted that the above-mentioned feature vectors can be processed by splicing or other arbitrary vectors to calculate and correct the eigenvalues ​​of the feature vectors themselves. The above-mentioned preset pathological feature vector processing model may include but is not limited to BERT (Bidirectional Encoder Representations from Transformers), CLIP (Contrastive Language–Image Pre-training) and MMBT (Multimodal Bitransformer) and other deep learning models that can perform feature vector processing on the above-mentioned text clinical sub-data and image clinical sub-data.

[0079] The above mapping relationship may refer to the intermediate conditions for converting the original medical clinical data into feature vectors, which can then be used for various prediction models, classification tasks or in-depth data analysis. Specifically, the above text clinical sub-data and image clinical sub-data may be processed and analyzed through NLP technology and image processing technology to obtain corresponding feature data, and the corresponding feature data may be vectorized. More specifically, the above feature data may be converted into numerical form, and a certain point in the vector space of the corresponding feature data may be taken as the vector value corresponding to the word or sentence of the text content, or the image data may be vector encoded through an encoding method to obtain the image feature vector corresponding to the image clinical sub-data.

[0080] Step S205, performing feature integration processing on the image feature vector and the text feature vector to obtain a medical feature vector of the current disease;

[0081] Step S206: Determine at least one classification result of the current disease based on the medical feature vector.

[0082] In this embodiment, the above-mentioned feature integration processing can perform feature integration processing on the above-mentioned image feature vector and text feature vector through a trained image-text joint model, and output a splicing vector of text features and image features. The dimension of the splicing vector can be a matrix of (N+3,D), and probability distribution statistics is performed on the splicing vector, and at least one classification result of the current disease is determined according to the statistical results.

[0083] In a possible embodiment, the case analysis system performs feature integration processing on the image feature vector and the text feature vector through a trained image-text joint model to obtain a spliced ​​vector, and performs probability distribution statistics on the spliced ​​vector, and determines at least one classification result of the current disease based on the statistical results.

[0084] The present application obtains the medical clinical data corresponding to the current disease; performs feature classification processing on the text clinical data and the image clinical data to obtain the text clinical sub-data and the image clinical sub-data of the corresponding feature classification; performs feature vector processing on the text clinical sub-data and the image clinical sub-data of the corresponding feature classification through a preset pathological feature vector processing model to obtain the text feature vector and the image feature vector corresponding to the current disease; performs feature integration processing on the image feature vector and the text feature vector to determine at least one classification result of the current disease. The present application improves the refinement of the medical clinical data by deepening the hierarchical structure of the medical clinical data, so that the model can judge the characteristics of the current disease based on more refined medical clinical data, thereby improving the accuracy of clinical data analysis.

[0085] In some optional implementations, step S202 includes the following steps:

[0086] Perform text analysis on the text clinical data to determine the text content corresponding to the title level, paragraph level and sentence level in the text clinical data;

[0087] According to the hierarchical coding method, semantic features are extracted from the text contents corresponding to the title level, paragraph level and sentence level to obtain the clinical sub-data of the title level text, the clinical sub-data of the paragraph level text and the clinical sub-data of the sentence level text.

[0088] In the present embodiment, the above-mentioned text parsing can be a process of content splitting the above-mentioned text clinical data according to the text features of titles, sentences and paragraphs. Specifically, it can be obtained by feature parsing the text content of the above-mentioned text clinical data through a convolutional neural network model or a Transformer model. By extracting and analyzing the text features of titles, sentences and paragraphs in the text content of the above-mentioned text clinical data, the title text, sentence text and paragraph text are determined respectively. More specifically, the title text can be extracted by the size and boldness of the text, the sentence text can be extracted by the punctuation features, and the paragraph text can be extracted by the features of multiple paragraph blanks and the features corresponding to the sentence text.

[0089] The above hierarchical encoding method can divide the above text clinical data into titles, sentences and paragraphs to achieve multi-level text representation. By constructing different levels of text feature representation, the understanding and processing capabilities of subsequent models can be improved, thereby more effectively processing and analyzing text clinical data.

[0090] In a possible embodiment, the above-mentioned clinical data analysis system processes and analyzes text clinical data through a hierarchical encoding method to extract fine text clinical sub-data from structured text clinical data, wherein text clinical data is parsed to identify and separate title-level, paragraph-level, and sentence-level content in the document. Each level of text data can be processed by a specific deep learning model, such as using a pre-trained language model such as BERT to extract semantic features at each level. This hierarchical processing can analyze the text content in detail and extract text clinical sub-data at the title level, paragraph level, and sentence level. Data extraction at each level focuses on capturing relevant clinical information, such as disease descriptions, treatment methods, or patient responses, which can then be used for further clinical research or to support medical decisions.

[0091] This application uses hierarchical encoding and deep learning models to more accurately understand and utilize complex information in clinical texts, which helps reduce the burden on doctors when manually processing large amounts of documents. It can also provide more accurate data support such as disease diagnosis and patient management, thereby improving the quality of medical services and patient treatment outcomes.

[0092] In some optional implementations, step S202 further includes the following steps:

[0093] By presetting the segmentation size, the image clinical data is divided to obtain multiple image data of the same size;

[0094] Performing image feature clustering on multiple image data of the same size by using a clustering algorithm to obtain multiple clustered image data;

[0095] The multiple clustered image data are used as image clinical sub-data.

[0096] In this embodiment, the preset segmentation size can be determined based on the size of the clinical data of the image and image features such as resolution. Generally speaking, the lower the resolution of the image, the larger the preset segmentation size can be set to ensure the clarity of the image. Specifically, the input image can be divided into multiple tiles through the ViT model, and features can be extracted for each tile, wherein the tiles can be divided into grids of equal size. Assuming that the width of the input image is W and the height is H, the ViT model can divide the image into a grid of W0 rows and H0 columns, wherein each grid has a width of P and a height of Q. Wherein, Indicates the number of rows and columns of the grid.

[0097] The above clustering algorithm may include but is not limited to any deep learning model such as K-means clustering and hierarchical clustering that can cluster the image features of the above multiple image data of the same size. Specifically, by identifying and analyzing the pathological features contained in the above multiple image data of the same size, the image data with the same pathological features are aggregated. For example, when there are pathological features corresponding to the rash in picture A, and there are pathological features corresponding to the rash in picture B, then picture A and picture B are aggregated to form clustered image data, or there are pathological features corresponding to the rash in picture A, and there are pathological features in picture B that are inferred from the pathological features of picture A. For example, when the pathological texture in picture A is derived through the above clustering algorithm and spreads to picture B, and there may also be corresponding pathological textures in picture B, it means that there are the same features in picture A and picture B, that is, picture A and picture B are also aggregated to form clustered image data.

[0098] In a possible embodiment, the clinical data analysis platform uses preset size parameters to evenly divide a large amount of image clinical data into multiple image blocks of the same size, and applies a clustering algorithm, such as K-means or DBSCAN, to cluster these image blocks according to their visual features to obtain image clinical sub-data.

[0099] This application uses fine-grained processing and clustering of image data to quickly locate and identify key pathological features and accelerate the diagnosis process. Through cluster analysis, vague or unclear pathological features can be matched with known disease patterns. Automated image processing reduces labor costs and improves the ability to process large-scale data sets, thereby improving the efficiency of clinical data analysis.

[0100] In some optional implementations, step S203 includes the following steps:

[0101] Based on the number of characters in each sentence in the sentence-level text clinical sub-data, performing text feature vector processing on the plurality of sentence-level text clinical sub-data to obtain a plurality of first text feature vectors;

[0102] Aggregating multiple first text feature vectors at the current paragraph level by a weighted average pooling algorithm to obtain a second text feature vector, and repeating the steps of the weighted average pooling algorithm to obtain multiple second text feature vectors;

[0103] Based on the number of paragraph-level text clinical sub-data, performing text feature vector processing on the title-level text clinical sub-data to obtain a plurality of third text feature vectors;

[0104] Determine weight values ​​corresponding to the first text feature vector, the second text feature vector, and the third text feature vector;

[0105] Based on the corresponding weight values, feature vector connection is performed on the first text feature vector, the second text feature vector and the third text feature vector to obtain a text feature vector corresponding to the current disease type.

[0106] In this embodiment, the sentence-level text clinical sub-data can be processed into text feature vectors by the following formula:

[0107] S i =BERT(Xi);

[0108] Among them, S i is the first text feature vector, where represents the th sentence, i is the output corresponding to the cls character of BERT, and is a 1×D vector. D is the vector length, which is generally 768 bytes.

[0109] The text feature vector processing of the paragraph-level text clinical sub-data can be obtained by aggregating the first text feature vector through the weighted average pooling algorithm. Specifically, the weight vector W of the sentences in the paragraph can be defined first. p , with a size of N, where N represents the number of sentences in the paragraph. The importance of the first text feature vector can be normalized using the softmax function through the following formula to obtain a weight vector:

[0110]

[0111] Among them, Wi is the weight vector corresponding to the i-th sentence, l i Represented as the length of the i-th sentence, represents the length of the sentence, and N represents the total number of sentences. The current weight vector W p For {W1, W2, W3...W i} is the weight vector of .

[0112] The weight vector W p The dimension expansion is performed through the following formula:

[0113] W p =W p T ×ones

[0114] Among them, W p is a 1×N vector, ones is a 1xD vector of all 1s, ^T represents the transpose operation, W p The dimension of is (N,D). It should be noted that W pThe weights can be determined based on the length of the sentences. For example, longer sentences may contain more information and thus be given higher weights, while shorter sentences may be given lower weights.

[0115] Finally, the second text feature vector P corresponding to each paragraph is obtained by multiplying the first text feature vector with the corresponding weight and performing weighted average pooling processing:

[0116] P = sum(S i *W p [i])

[0117] Where sum represents the sum operation, W p [i] represents the weight of the i-th sentence. The vector dimension of P is 1×D.

[0118] For the title-level feature vector, the number of the above-mentioned paragraph-level text clinical sub-data can be calculated to obtain the number of corresponding title-level text clinical sub-data, and then the above-mentioned text feature vector processing process for the paragraph-level text clinical sub-data can be repeated to obtain the title-level feature vector, that is, the third text feature vector T, and the dimension of T can also be 1×D.

[0119] The above eigenvector connection can be calculated by the following formula:

[0120]

[0121] H_T=T;

[0122] H_Text = [H_S; H_P; H_T];

[0123] Among them, H_Text is the text feature vector corresponding to the current disease type, H_S represents the average dimension of all sentence vectors, which is (1, D), H_P represents the average of all paragraphs, which is (1, D), and H_T represents the title vector, which is (1, D). H_Text is the concatenation vector of the above three vectors, and the vector dimension is (3, D).

[0124] In a possible embodiment, the above-mentioned case analysis system adopts a hierarchical text feature extraction and aggregation method to analyze and process text clinical data, and obtains the corresponding first text feature vector, second text feature vector and third text feature vector by extracting features from sentence-level text clinical sub-data, paragraph-level text clinical sub-data and title-level text clinical sub-data. Then, the above-mentioned first text feature vector, second text feature vector and third text feature vector are vector-connected according to predefined weight values ​​to obtain the text feature vector corresponding to the current disease type.

[0125] Through fine-grained disease text feature extraction and intelligent aggregation, this application can more deeply understand the intrinsic structure and semantic information of text clinical data, thereby improving the accuracy of pathological diagnosis.

[0126] In some optional implementations, step S203 further includes the following steps:

[0127] Extracting lesion features corresponding to multiple clustered image data in the clinical sub-data of the image;

[0128] Performing vector processing on the lesion features to obtain a plurality of corresponding lesion feature vectors;

[0129] Based on multiple corresponding lesion feature vectors, determine the image feature vector corresponding to the current disease type.

[0130] In this embodiment, the above pathological features can be the characteristics of the current disease, and can be described by image data. For example, the characteristics can be white hair of albinism, red skin with red spots of erythematosus, etc. The above pathological features can also be graded according to the degree of damage to the body, and the classification can be determined according to the specific implementation plan. Generally speaking, the relative level of symptoms that appear on the surface, such as white hair and white lips, can be set lower.

[0131] Specifically, the following formula can be used to perform vector processing on multiple clustered image data in the divided image clinical sub-data:

[0132] H_Image = N × D;

[0133] Where H_Image is the lesion feature vector, and the feature vector N representing the image represents the number of blocks. The number of N can be calculated by N=W0×H0.

[0134] In a possible embodiment, the case analysis system processes the clinical sub-data of the image through image processing technology and clustering algorithm, so as to identify and classify the image areas containing lesions, and then automatically extracts key lesion features from these areas through models such as convolutional neural networks (CNN), and converts these features into feature vectors in numerical form. It should be noted that each of the above feature vectors can capture key visual information about the lesion, such as shape, size, texture and edge characteristics, and the lesion feature vectors are combined through a weighted average algorithm to obtain the image feature vector corresponding to the current disease type.

[0135] This application uses image processing technology and pathological feature extraction to extract the lesion features of the clinical sub-data of the image, and then obtains the corresponding lesion feature vector by converting the lesion features into a vector in numerical form. Finally, a weighted average algorithm is used to calculate multiple lesion feature vectors to obtain the image feature vector corresponding to the current disease, which increases the accuracy of lesion feature recognition for clinical data in the form of images and improves the efficiency of clinical data analysis based on pathological images.

[0136] In some optional implementations, step S204 includes the following steps:

[0137] Through the trained image-text joint model, the image feature vector and the text feature vector are integrated to determine the medical feature vector of the current disease;

[0138] Based on the medical feature vector, at least one classification result of the current disease is determined.

[0139] In this embodiment, the above-mentioned trained joint graph-text model can be used to fuse and integrate different types of data sources, such as text data and image data. Specifically, the above-mentioned text clinical data and image clinical data can be fused through the above-mentioned trained joint graph-text model to obtain the classification results of the above-mentioned diseases. The above-mentioned trained joint graph-text model may include but is not limited to BERT (Bidirectional Encoder Representations from Transformers), CLIP (Contrastive Language–Image Pre-training) and MMBT (Multimodal Bitransformer) and other deep learning models that can fuse the above-mentioned text clinical data and image clinical data.

[0140] The above medical feature vectors can be integrated through the above trained image-text joint model in the following way:

[0141] H_merge = [H_Text; H_Image]

[0142] Among them, H_merge represents the concatenated vector of the image and text vectors, that is, the above-mentioned medical feature vector. The vector dimension can be (N+3,D), which can be obtained by multiplying the above-mentioned text feature vector and the image feature vector.

[0143] Then, the fusion vector model is processed by a multi-layer perceptron to process the above medical feature vectors. The specific method is as follows:

[0144] The above medical feature vector is used as a sample of the multi-layer perceptron (Multi-Layer Perceptron) processing fusion vector model for input, and then the dimension of the above medical feature vector is converted from (N+3,D) to (1,D) according to the global average pooling operation in the global pooling layer, and then the output is nonlinearly transformed through the activation function of the pooling layer, such as the ReLU function, to obtain a numerical value, and then the numerical value is output through the fully connected layer to facilitate subsequent probability distribution calculations.

[0145] After calculating the probability distribution of this value and combining it with the incidence probability of the corresponding disease, the classification result of the current disease is calculated.

[0146] The above classification results can be the corresponding development stage determined by the current disease according to different shape manifestations, or the mutated disease type of the current disease under different shape characteristics. In a possible embodiment, the above clinical data analysis platform fuses the text clinical data and the image clinical data through the above trained graphic joint model, at least obtaining the corresponding development stage determined by the disease under different shape manifestations, or the mutated disease type of the current disease under different shape characteristics, and stores the classification results of the current disease in the database of the corresponding archive.

[0147] This application processes and calculates the image feature vector and text feature vector through a trained joint image-text model to obtain the classification result of the current disease, effectively integrates the key information in the image and text data, and significantly improves the accuracy and efficiency of pathological diagnosis.

[0148] In some optional implementations, the following steps are further included before step S204:

[0149] Obtain the training text clinical data, training image clinical data and the image-text joint model to be trained for the corresponding disease;

[0150] Marking the training text clinical data and the training image clinical data with feature vector labels to obtain a first training data set;

[0151] Performing feature vector fusion processing on the training text clinical data and the training image clinical data marked by the feature vector label in the first training data set, so that the medical feature vectors fused from the training text clinical data and the training image clinical data marked by the feature vector label are cross-referenced to obtain a second training data set;

[0152] The image-text joint model to be trained is iteratively trained based on the second training data set, and a trained image-text joint model is obtained after the iterative training is completed.

[0153] In this embodiment, the training text clinical data and training image clinical data can be obtained through any channels or institutions that can provide text descriptions corresponding to the characteristics of known diseases, such as text records of clinical studies, public medical databases, and medical academics. It should be noted that the acquisition of the training text clinical data and training image clinical data of the corresponding diseases by the above channels or institutions needs to comply with relevant privacy and legal regulations, and be legally obtained under relevant requirements.

[0154] The above-mentioned image-text joint model to be trained can be any deep learning model that has not been trained, such as BERT (Bidirectional Encoder Representations from Transformers), CLIP (Contrastive Language–Image Pre-training), and MMBT (Multimodal Bitransformer), which can fuse the above-mentioned image feature vectors and text feature vectors but cannot accurately output the corresponding medical feature vectors.

[0155] Generally speaking, inputting the above-mentioned image feature vectors and text feature vectors corresponding to the disease types into the above-mentioned image-text joint model to be trained may not necessarily output the corresponding medical feature vectors corresponding to the disease types.

[0156] The above-mentioned feature vector label marking can be a process of marking descriptions of the same shape content features between the training text clinical data and the training image clinical data. Generally speaking, the above-mentioned feature vector label marking can be regarded as a process of connecting the training text clinical data and the training image clinical data, that is, the same or similar trait features between the training text clinical data and the training image clinical data are expressed with the same vector. For example, when the text content of the training text clinical data describes blisters with convex and large-scale distribution trait features, the pictures corresponding to the trait features with the same convex and large-scale distribution in the training text clinical data can be marked with feature vector labels, that is, the above-mentioned multimodal model to be trained can match the corresponding picture data of blisters with convex and large-scale distribution trait features through the text corresponding to the above-mentioned blisters with convex and large-scale distribution trait features.

[0157] In a possible embodiment, the clinical data analysis platform can vectorize the training text clinical data, cross-reference the corresponding training image clinical data and the training text clinical data in the same feature space, and construct the training image clinical data and the training text clinical data according to the vector feature label to obtain a classifier based on text content, which can be used as the input of the image-text joint model to be trained. It should be noted that multiple classifiers can be used as the second training data set.

[0158] In a possible embodiment, the second training data set can be divided into a training set, a validation set and a test set. The graphic-text joint model to be trained is trained according to the second training data set, and the parameters of the model are adjusted according to the output results. The performance of the model is optimized through the validation set. Finally, the graphic-text joint model to be trained is evaluated according to the test set. When the graphic-text joint model to be trained processes the input training set and the validation set, and the output result is close to the output threshold, it means that the training of the graphic-text joint model to be trained is completed, and a trained graphic-text joint model is obtained.

[0159] It should be noted that during the training phase, a loss function (such as cross entropy loss) can be used to measure the difference between the model's prediction results and the true label, and the model parameters can be updated through optimization algorithms such as back propagation and gradient descent.

[0160] This application enhances the model's ability to understand multimodal data by comprehensively utilizing key information in text and image data, thereby improving the accuracy and reliability of disease diagnosis.

[0161] It is understandable that in the specific implementation of this application, medical clinical data and other related data are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0162] It should be emphasized that in order to further ensure the privacy and security of the above-mentioned medical clinical data, the above-mentioned medical clinical data can also be stored in a blockchain node.

[0163] The blockchain referred to in this application is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.

[0164] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0165] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0166] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through computer-readable instructions, and the computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0167] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0168] Further references Figure 3 , as a response to the above Figure 2 In order to realize the method shown in the figure, the present application provides an embodiment of a clinical data analysis device, and the device embodiment is Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0169] like Figure 3As shown, the clinical data analysis device 300 described in this embodiment includes: a first acquisition module 301, a second acquisition module 302, a first processing module 303, a second processing module 304, an integration processing module 305, and a classification module 306. Among them:

[0170] The first acquisition module 301 is used to acquire medical clinical data corresponding to the current disease, wherein the medical data includes text clinical data and image clinical data;

[0171] The second acquisition module 302 is used to perform feature classification processing on the text clinical data and the image clinical data to obtain text clinical sub-data and image clinical sub-data corresponding to the feature classification;

[0172] The first processing module 303 is used to perform feature vector processing on the clinical sub-data of the picture corresponding to the feature classification through a preset pathological feature vector processing model to obtain the picture feature vector corresponding to the current disease type;

[0173] The second processing module 304 is used to perform feature vector processing on the text clinical sub-data of the corresponding feature classification through a preset pathological feature vector processing model to obtain a text feature vector corresponding to the current disease type;

[0174] An integration processing module 305 is used to perform feature integration processing on the image feature vector and the text feature vector to obtain a medical feature vector of the current disease;

[0175] The classification module 306 is used to determine at least one classification result of the current disease based on the medical feature vector.

[0176] In a possible embodiment, the second acquisition module 302 includes:

[0177] The first determination submodule is used to perform text analysis on the text clinical data to determine the text content corresponding to the title level, paragraph level and sentence level in the text clinical data;

[0178] The first acquisition submodule is used to extract semantic features of text content corresponding to the title level, paragraph level and sentence level according to the hierarchical coding method to obtain clinical sub-data of title level text, clinical sub-data of paragraph level text and clinical sub-data of sentence level text.

[0179] In a possible embodiment, the second obtaining module 302 further includes:

[0180] The second acquisition submodule is used to divide the image clinical data by a preset segmentation size to obtain multiple image data of the same size;

[0181] A first clustering submodule is used to perform image feature clustering on a plurality of image data of the same size by using a clustering algorithm to obtain a plurality of clustered image data;

[0182] The third acquisition submodule is used to use the multiple clustered image data as image clinical sub-data.

[0183] In a possible embodiment, the first processing module 303 includes:

[0184] A fourth acquisition submodule is used to perform text feature vector processing on a plurality of sentence-level text clinical sub-data based on the number of characters of each sentence in the sentence-level text clinical sub-data to obtain a plurality of first text feature vectors;

[0185] A fifth acquisition submodule is used to aggregate multiple first text feature vectors at the current paragraph level through a weighted average pooling algorithm to obtain a second text feature vector, and repeat the steps of the weighted average pooling algorithm to obtain multiple second text feature vectors;

[0186] a sixth acquisition submodule, configured to perform text feature vector processing on the title-level text clinical sub-data based on the number of paragraph-level text clinical sub-data, to obtain a plurality of third text feature vectors;

[0187] A second determination submodule is used to determine weight values ​​corresponding to the first text feature vector, the second text feature vector, and the third text feature vector;

[0188] The seventh acquisition submodule is used to perform feature vector connection on the first text feature vector, the second text feature vector and the third text feature vector based on the corresponding weight values ​​to obtain the text feature vector corresponding to the current disease type.

[0189] In a possible embodiment, the first processing module 303 further includes:

[0190] An extraction submodule is used to extract lesion features corresponding to multiple clustered image data in the image clinical sub-data;

[0191] A vector processing submodule is used to perform vector processing on lesion features to obtain a plurality of corresponding lesion feature vectors;

[0192] The third determination submodule is used to determine the image feature vector corresponding to the current disease type based on multiple corresponding lesion feature vectors.

[0193] In a possible embodiment, the second processing module 304 includes:

[0194] The fourth determination submodule is used to perform feature integration processing on the image feature vector and the text feature vector through the trained image-text joint model to determine the medical feature vector of the current disease;

[0195] The fifth determination submodule is used to determine at least one classification result of the current disease based on the medical feature vector.

[0196] In a possible embodiment, the device further includes:

[0197] The second acquisition module is used to acquire the training text clinical data, training image clinical data and the image-text joint model to be trained for the corresponding disease;

[0198] A first labeling module is used to label the training text clinical data and the training image clinical data with feature vector labels to obtain a first training data set;

[0199] A second labeling module is used to perform feature vector fusion processing on the training text clinical data and the training image clinical data labeled with feature vector labels in the first training data set, so that the medical feature vectors fused from the training text clinical data and the training image clinical data labeled with feature vector labels are cross-referenced to obtain a second training data set;

[0200] The training module is used to iteratively train the image-text joint model to be trained based on the second training data set, and obtain the trained image-text joint model after the iterative training is completed.

[0201] In this embodiment, the medical clinical data corresponding to the current disease is obtained; the text clinical data and the image clinical data are subjected to feature classification processing to obtain text clinical sub-data and image clinical sub-data of the corresponding feature classification; the text clinical sub-data and the image clinical sub-data of the corresponding feature classification are subjected to feature vector processing by a preset pathological feature vector processing model to obtain the text feature vector and the image feature vector corresponding to the current disease; the image feature vector and the text feature vector are subjected to feature integration processing to determine at least one classification result of the current disease. This application improves the refinement of the medical clinical data by deepening the hierarchical structure of the medical clinical data, so that the model can judge the characteristics of the current disease based on more refined medical clinical data, thereby improving the accuracy of clinical data analysis.

[0202] In this embodiment, the operations performed by the above-mentioned units or modules respectively correspond one-to-one to the steps of the clinical data analysis method of the above-mentioned implementation mode, and will not be repeated here.

[0203] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0204] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with components 41-43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (Application Specific Integrated Circuit, ASIC), programmable gate arrays (Field-Programmable Gate Array, FPGA), digital processors (Digital Signal Processor, DSP), embedded devices, etc.

[0205] The computer device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may interact with a user through a keyboard, a mouse, a remote controller, a touch pad, or a voice control device.

[0206] The memory 41 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as a hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. equipped on the computer device 4. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of the clinical data analysis method, etc. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.

[0207] The processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run computer-readable instructions stored in the memory 41 or process data, such as computer-readable instructions for running the clinical data analysis method.

[0208] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0209] This embodiment provides a computer device, which obtains medical clinical data corresponding to the current disease; performs feature classification processing on text clinical data and image clinical data to obtain text clinical sub-data and image clinical sub-data of corresponding feature classification; performs feature vector processing on text clinical sub-data and image clinical sub-data of corresponding feature classification through a preset pathological feature vector processing model to obtain text feature vectors and image feature vectors corresponding to the current disease; performs feature integration processing on the image feature vector and the text feature vector to determine at least one classification result of the current disease. This application improves the refinement of medical clinical data by deepening the hierarchical structure of medical clinical data, so that the model can judge the characteristics of the current disease based on more refined medical clinical data, thereby improving the accuracy of clinical data analysis.

[0210] The present application also provides another embodiment, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the clinical data analysis method as described above.

[0211] This embodiment provides a computer-readable storage medium, which obtains medical clinical data corresponding to the current disease; performs feature classification processing on text clinical data and image clinical data to obtain text clinical sub-data and image clinical sub-data of corresponding feature classification; performs feature vector processing on text clinical sub-data and image clinical sub-data of corresponding feature classification through a preset pathological feature vector processing model to obtain text feature vectors and image feature vectors corresponding to the current disease; performs feature integration processing on the image feature vector and the text feature vector to determine at least one classification result of the current disease. This application improves the refinement of medical clinical data by deepening the hierarchical structure of medical clinical data, so that the model can judge the characteristics of the current disease based on more refined medical clinical data, thereby improving the accuracy of clinical data analysis.

[0212] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0213] Obviously, the embodiments described above are only some embodiments of the present application, rather than all embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application is described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific implementation methods, or to perform equivalent replacement of some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of this application, directly or indirectly used in other related technical fields, is similarly within the scope of patent protection of this application.

Claims

1. A clinical data analysis method, characterized in that: The steps include: Obtaining medical clinical data corresponding to the current disease, wherein the medical data includes text clinical data and image clinical data; Performing feature classification processing on the text clinical data and the image clinical data to obtain text clinical sub-data and image clinical sub-data corresponding to the feature classification; By using a preset pathological feature vector processing model, feature vector processing is performed on the clinical sub-data of the image corresponding to the feature classification to obtain the image feature vector corresponding to the current disease type; Performing feature vector processing on the text clinical sub-data of the corresponding feature classification by using a preset pathological feature vector processing model to obtain a text feature vector corresponding to the current disease type; Performing feature integration processing on the image feature vector and the text feature vector to obtain a medical feature vector of the current disease; Based on the medical feature vector, at least one classification result of the current disease is determined.

2. The clinical data analysis method according to claim 1, characterized in that: The text clinical sub-data includes title-level text clinical sub-data, paragraph-level text clinical sub-data and sentence-level text clinical sub-data. The step of performing feature classification processing on the text clinical data to obtain text clinical sub-data corresponding to the feature classification includes: Performing text parsing on the text clinical data to determine the text content corresponding to the title level, paragraph level, and sentence level in the text clinical data; According to the hierarchical coding method, semantic features are extracted from the text contents corresponding to the title level, paragraph level and sentence level to obtain the clinical sub-data of the title level text, the clinical sub-data of the paragraph level text and the clinical sub-data of the sentence level text.

3. The clinical data analysis method according to claim 1, characterized in that: The step of performing feature classification processing on the image clinical data to obtain image clinical sub-data corresponding to the feature classification also includes: Dividing the image clinical data by a preset segmentation size to obtain a plurality of image data of the same size; Performing image feature clustering on the plurality of image data of the same size by using a clustering algorithm to obtain a plurality of clustered image data; The plurality of clustered image data are used as the image clinical sub-data.

4. The clinical data analysis method according to claim 2, characterized in that: The step of performing feature vector processing on the text clinical sub-data of the corresponding feature classification by using a preset pathological feature vector processing model to obtain a text feature vector corresponding to the current disease type specifically includes: Based on the number of characters in each sentence in the sentence-level text clinical sub-data, performing text feature vector processing on the plurality of sentence-level text clinical sub-data to obtain a plurality of first text feature vectors; Aggregating multiple first text feature vectors at the current paragraph level by a weighted average pooling algorithm to obtain a second text feature vector, and repeating the steps of the weighted average pooling algorithm to obtain multiple second text feature vectors; Based on the number of paragraph-level text clinical sub-data, performing text feature vector processing on the title-level text clinical sub-data to obtain a plurality of third text feature vectors; Determine the weight values ​​corresponding to the first text feature vector, the second text feature vector, and the third text feature vector; Based on the corresponding weight values, feature vector connection is performed on the first text feature vector, the second text feature vector and the third text feature vector to obtain a text feature vector corresponding to the current disease type.

5. The clinical data analysis method according to claim 1, characterized in that: The step of performing feature vector processing on the clinical sub-data of the picture corresponding to the feature classification by using a preset pathological feature vector processing model to obtain the picture feature vector corresponding to the current disease type specifically includes: Extracting lesion features corresponding to a plurality of clustered image data from the clinical sub-data of the image; Performing vector processing on the lesion features to obtain a plurality of corresponding lesion feature vectors; Based on the multiple corresponding lesion feature vectors, determine the image feature vector corresponding to the current disease type.

6. The clinical data analysis method according to claim 1, characterized in that: The step of determining at least one classification result of the current disease based on the medical feature vector comprises: The medical feature vector of the current disease is classified by using the trained image-text joint model to obtain at least one classification result of the current disease.

7. The clinical data analysis method according to claim 6, characterized in that: Before the step of performing feature integration processing on the image feature vector and the text feature vector through the trained image-text joint model to determine the medical feature vector of the current disease, the method further includes: Obtain the training text clinical data, training image clinical data and the image-text joint model to be trained for the corresponding disease; Marking the training text clinical data and the training image clinical data with feature vector labels to obtain a first training data set; Performing feature vector fusion processing on the training text clinical data and the training image clinical data marked by the feature vector label in the first training data set, so that the medical feature vectors fused from the training text clinical data and the training image clinical data marked by the feature vector label are cross-referenced to obtain a second training data set; The image-text joint model to be trained is iteratively trained based on the second training data set, and the trained image-text joint model is obtained after the iterative training is completed.

8. A clinical data analysis device, characterized in that: include: A first acquisition module is used to acquire medical clinical data corresponding to the current disease, wherein the medical data includes text clinical data and image clinical data; A second acquisition module is used to perform feature classification processing on the text clinical data and the image clinical data to obtain text clinical sub-data and image clinical sub-data corresponding to the feature classification; The first processing module is used to perform feature vector processing on the clinical sub-data of the picture corresponding to the feature classification through a preset pathological feature vector processing model to obtain the picture feature vector corresponding to the current disease type; The second processing module is used to perform feature vector processing on the text clinical sub-data of the corresponding feature classification through a preset pathological feature vector processing model to obtain a text feature vector corresponding to the current disease type; An integration processing module is used to perform feature integration processing on the image feature vector and the text feature vector to obtain a medical feature vector of the current disease; The classification module is used to determine at least one classification result of the current disease based on the medical feature vector.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the clinical data analysis method according to any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the clinical data analysis method according to any one of claims 1 to 7 are implemented.