Artificial Intelligence-Based Auxiliary Interpretation Method, Device, Terminal, and Storage Medium
Through the combination of multi-level neural networks, including convolutional neural networks and GRU recurrent neural networks, extracting and fusing the features of medical images, the problem of limited auxiliary interpretation effects in the prior art is solved, and more efficient and accurate medical image processing and interpretation are achieved.
Patent Information
- Application Number
- CN201910324859.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-04-22
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2039-04-22
AI Technical Summary
The existing deep learning technology based on artificial intelligence has limited auxiliary interpretation effects in the field of medical interpretation, making it difficult to effectively identify and process high-complex medical images.
Multi-level neural network combinations are adopted, including initializing convolutional neural networks, extracting Laplace and Hesser matrix features, and fusing features through gated recurrent units (GRU) recurrent neural networks to form a combined feature vector for processing.
It improves the classification and processing efficiency of medical images, enhances the accuracy and reliability of auxiliary interpretation, and improves the efficiency of deep learning by continuously optimizing interpretation parameters.
Smart Images

Figure CN111833991B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical auxiliary interpretation, and particularly to an auxiliary interpretation method, device, terminal and storage medium based on artificial intelligence. Background Art
[0002] Some medical tools for visualizing internal human organs can fully display the condition of human organs from the inside, and at the same time, there is no pain and discomfort of conventional endoscopes, which is very easy to implement. However, the quantity and complexity of the video images generated by these visual medical tools are very challenging for recognition. For example, by using a capsule endoscope to record the normal small intestine video images of a patient, the recording can be as long as 8 hours. If a recording form of 2 frames per second is used, a video can contain up to 57,600 images (the number of images per second at 24 frames per second will reach 691,200); secondly, the visual medical tools may move freely within the human organs, and the imaging conditions of the captured images vary greatly, and sometimes they are even completely or partially affected by other substances, such as bile or food being digested. Therefore, the complexity of the video images is extremely high.
[0003] The current auxiliary interpretation devices have the problems of few recognized disease types and non - general methods for recognizing diseases. For different diseases, specific feature extraction methods and determination methods need to be designed. The emergence of deep learning neural networks has changed the traditional way of image recognition. Deep models have powerful learning capabilities and efficient feature expression capabilities, extracting information layer by layer from pixel - level raw data to abstract semantic concepts. This makes it have outstanding advantages in extracting global features and context information of images.
[0004] There have been some attempts and progress in applying deep learning to intelligent auxiliary interpretation. However, due to the characteristics of the medical interpretation field, the auxiliary interpretation based on deep learning technology of artificial intelligence often has limited effects. Summary of the Invention
[0005] Based on this, in view of the fact that the auxiliary interpretation based on deep learning technology of artificial intelligence often has limited effects, it is necessary to provide an auxiliary interpretation method, device, terminal and storage medium based on artificial intelligence.
[0006] An auxiliary interpretation method based on artificial intelligence includes:
[0007] Initializing the parameters of the first convolutional neural network, the parameters of the second convolutional neural network, and the parameters of the gated recurrent unit (GRU) recurrent neural network;
[0008] Obtaining the image information to be interpreted and performing pre - processing;
[0009] Use the first convolutional neural network to perform preliminary processing on the preprocessed image information to be interpreted, obtaining a preliminary feature vector of the image information to be interpreted and a classified sequence of images to be interpreted;
[0010] Extract the grayscale features of the classified sequence of images to be interpreted, respectively extract the Laplacian features and Hessian matrix features of the grayscale features, and use the second convolutional neural network to process the Laplacian features and the Hessian matrix features, respectively obtaining a Laplacian feature vector and a Hessian matrix feature vector;
[0011] Fuse the Laplacian feature vector, the Hessian matrix feature vector and the preliminary feature vector to form a combined feature vector, and use the GRU recurrent neural network to process the combined feature vector; and
[0012] Display the processing results for assisting in interpretation and store the interpretation results.
[0013] In one embodiment, the initialization includes:
[0014] Provide an initial training database, where the initial training database includes labeled image information with interpretation annotations;
[0015] Obtain the labeled image information and perform preprocessing;
[0016] Use a training convolutional neural network to process the preprocessed labeled image information, obtaining the parameters of the first convolutional neural network, a labeled feature vector and a classified sequence of labeled images;
[0017] Extract the labeled grayscale features of the classified sequence of labeled images, respectively extract the labeled Laplacian features and labeled Hessian matrix features of the labeled grayscale features, and use the training convolutional neural network to process the labeled Laplacian features and the labeled Hessian matrix features, obtaining the parameters of the second convolutional neural network, a labeled Laplacian feature vector and a labeled Hessian matrix feature vector; and
[0018] Fuse the labeled Laplacian feature vector, the labeled Hessian matrix feature vector and the labeled feature vector to form a feature matrix, quantize the feature matrix and use the GRU recurrent neural network to process the feature matrix, obtaining the parameters of the GRU recurrent neural network.
[0019] In one embodiment, the providing of the initial training database includes:
[0020] When the stored interpretation results reach a preset condition, update the initial training database according to the interpretation results.
[0021] In one embodiment, the preset conditions include at least one of a preset quantity and a preset time.
[0022] In one embodiment, using the trained convolutional neural network to process the preprocessed labeled image information to obtain the parameters of the first convolutional neural network includes:
[0023] Using mini-batch gradient descent to process the preprocessed labeled image information to obtain the parameters of the first convolutional neural network.
[0024] In one embodiment, using the trained convolutional neural network to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network includes:
[0025] Using stochastic gradient descent to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network.
[0026] In one embodiment, the preprocessing includes performing one or more of the following processes on the image information to be interpreted: segmentation, format conversion, deformation, flipping, distortion, brightness adjustment, data augmentation, cropping, illumination, and regularization.
[0027] An artificial intelligence-based auxiliary interpretation device includes:
[0028] A preprocessing module for obtaining image information to be interpreted and performing preprocessing;
[0029] A preliminary processing module including a first convolutional neural network for preliminarily processing the preprocessed image information to be interpreted to obtain a preliminary feature vector and a classified sequence of images to be interpreted;
[0030] A fine processing module including a second convolutional neural network for extracting the gray-scale features of the classified sequence of images to be interpreted, respectively extracting the Laplacian features and the Hessian matrix features of the gray-scale features, and using the second convolutional neural network to process the Laplacian features and the Hessian matrix features to respectively obtain a Laplacian feature vector and a Hessian matrix feature vector;
[0031] A fusion processing module including a gated recurrent unit (GRU) recurrent neural network for fusing the Laplacian feature vector, the Hessian matrix feature vector, and the preliminary feature vector to form a combined feature vector, and using the GRU recurrent neural network to process the combined feature vector;
[0032] An initialization module for initializing the parameters of the first convolutional neural network, the parameters of the second convolutional neural network, and the parameters of the GRU recurrent neural network; and
[0033] An interaction module, which is used to display the processing results for assisting in interpretation and store the interpretation results.
[0034] In one embodiment, the initialization module further includes:
[0035] An initial training database, which includes labeled image information with interpretation annotations;
[0036] Train a convolutional neural network; and
[0037] Train a GRU recurrent neural network;
[0038] The preprocessing module is further used to obtain the labeled image information and perform preprocessing;
[0039] The initialization module is further used to process the preprocessed labeled image information using the trained convolutional neural network to obtain the parameters of the first convolutional neural network, the labeled feature vectors, and the classified labeled image sequence; extract the labeled gray-scale features of the classified labeled image sequence, respectively extract the labeled Laplacian features and the labeled Hessian matrix features of the labeled gray-scale features, and use the trained convolutional neural network to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network and the labeled Laplacian feature vectors and the labeled Hessian matrix feature vectors; and fuse the labeled Laplacian feature vectors, the labeled Hessian matrix feature vectors, and the labeled feature vectors to form a feature matrix, and quantize the feature matrix and use the trained GRU recurrent neural network to process the feature matrix to obtain the parameters of the GRU recurrent neural network.
[0040] In one embodiment, the initialization module is further used to:
[0041] Update the initial training database according to the interpretation results stored in the interaction module.
[0042] In one embodiment, the training module is further used to process the preprocessed labeled image information using the mini-batch gradient descent method to obtain the parameters of the first convolutional neural network.
[0043] In one embodiment, the training module is further used to process the labeled Laplacian features and the labeled Hessian matrix features using the stochastic gradient descent method to obtain the parameters of the second convolutional neural network.
[0044] In one embodiment, the preprocessing includes performing one or more of the following processes on the image information to be interpreted: segmentation, format conversion, deformation, flipping, distortion, brightness adjustment, data augmentation, cropping, illumination, and regularization.
[0045] A terminal, comprising a memory and a processor, wherein instructions are stored in the memory, and when the instructions are executed by the processor, the processor performs the following steps:
[0046] Initialize the parameters of the first convolutional neural network, the parameters of the second convolutional neural network, and the parameters of the gated recurrent unit (GRU) recurrent neural network;
[0047] Obtain the image information to be judged and perform preprocessing;
[0048] Use the first convolutional neural network to perform preliminary processing on the preprocessed image information to be judged, and obtain a preliminary feature vector and a classified image sequence to be judged;
[0049] Extract the grayscale features of the classified image sequence to be judged, respectively extract the Laplacian features and Hessian matrix features of the grayscale features, and use the second convolutional neural network to process the Laplacian features and the Hessian matrix features to obtain a Laplacian feature vector and a Hessian matrix feature vector respectively;
[0050] Fuse the Laplacian feature vector, the Hessian matrix feature vector, and the preliminary feature vector to form a combined feature vector, and use the GRU recurrent neural network to process the combined feature vector; and
[0051] Display the processing result for auxiliary judgment and store the judgment result.
[0052] In one embodiment, when the instructions are executed by the processor, the processor further performs the following steps:
[0053] Provide an initial training database, where the initial training database includes labeled image information with judgment labels;
[0054] Obtain the labeled image information and perform preprocessing;
[0055] Use a training convolutional neural network to process the preprocessed labeled image information to obtain the parameters of the first convolutional neural network, a labeled feature vector, and a classified labeled image sequence;
[0056] Extract the labeled grayscale features of the classified labeled image sequence, respectively extract the labeled Laplacian features and labeled Hessian matrix features of the labeled grayscale features, and use the training convolutional neural network to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network, and a labeled Laplacian feature vector and a labeled Hessian matrix feature vector; and
[0057] Fuse the labeled Laplacian eigenvector, the labeled Hessian matrix eigenvector, and the labeled eigenvector to form a feature matrix, quantize the feature matrix, and process the feature matrix using a GRU recurrent neural network to obtain the parameters of the GRU recurrent neural network.
[0058] In one embodiment, when the instruction is executed by the processor, the processor is further caused to perform the following steps:
[0059] When the stored interpretation result reaches a preset condition, update the initial training database according to the interpretation result.
[0060] In one embodiment, the preset condition includes at least one of a preset quantity and a preset time.
[0061] One or more non-volatile storage media storing computer-executable instructions, which when executed by one or more processors, cause the one or more processors to perform the following steps:
[0062] Initialize the parameters of the first convolutional neural network, the parameters of the second convolutional neural network, and the parameters of the gated recurrent unit (GRU) recurrent neural network;
[0063] Obtain the image information to be interpreted and perform preprocessing;
[0064] Use the first convolutional neural network to perform preliminary processing on the preprocessed image information to be interpreted, to obtain a preliminary eigenvector and a classified sequence of images to be interpreted;
[0065] Extract the gray-scale features of the classified sequence of images to be interpreted, respectively extract the Laplacian feature and the Hessian matrix feature of the gray-scale features, and use the second convolutional neural network to process the Laplacian feature and the Hessian matrix feature to respectively obtain a Laplacian eigenvector and a Hessian matrix eigenvector;
[0066] Fuse the Laplacian eigenvector, the Hessian matrix eigenvector, and the preliminary eigenvector to form a combined eigenvector, and use the GRU recurrent neural network to process the combined eigenvector; and
[0067] Display the processing result for assisting in interpretation and store the interpretation result.
[0068] In one embodiment, when the instruction is executed by the processor, the processor is further caused to perform the following steps:
[0069] Provide an initial training database, where the initial training database includes labeled image information with interpretation annotations;
[0070] Obtain the marked image information and perform preprocessing;
[0071] Use the trained convolutional neural network to process the preprocessed marked image information to obtain the parameters of the first convolutional neural network, the marked feature vector, and the classified marked image sequence;
[0072] Extract the marked gray-scale features of the classified marked image sequence, respectively extract the marked Laplacian features and marked Hessian matrix features of the marked gray-scale features, and use the trained convolutional neural network to process the marked Laplacian features and the marked Hessian matrix features to obtain the parameters of the second convolutional neural network and the marked Laplacian feature vector and the marked Hessian matrix feature vector; and
[0073] Fuse the marked Laplacian feature vector, the marked Hessian matrix feature vector, and the marked feature vector to form a feature matrix, quantize the feature matrix, and use the GRU recurrent neural network to process the feature matrix to obtain the parameters of the GRU recurrent neural network.
[0074] In one embodiment, when the instruction is executed by the processor, the processor is further caused to perform the following steps:
[0075] When the stored interpretation result reaches a preset condition, update the initial training database according to the interpretation result.
[0076] The auxiliary interpretation method, device, terminal, and storage medium based on artificial intelligence in this application fuse the preliminary feature vector obtained by initially processing the to-be-interpreted image information after preprocessing using the first convolutional neural network, and the Laplacian feature vector and the Hessian matrix feature vector respectively obtained by using the second convolutional neural network to process the Laplacian feature and the Hessian matrix feature of the to-be-interpreted image information to classify and process the to-be-interpreted image, improving the classification and processing efficiency of the to-be-interpreted image and better assisting in interpretation. And according to the final interpretation result or processing result, continuously optimize the relevant interpretation parameters, that is, optimize the neural network model for interpretation, thereby greatly improving the efficiency of deep learning. Brief Description of the Drawings
[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0078] Figure 1 It is a schematic diagram of a terminal in an embodiment of the present application.
[0079] Figure 2 The flowchart of the auxiliary interpretation method of artificial intelligence in an embodiment of the present application.
[0080] Figure 3 The flowchart of the auxiliary interpretation method of artificial intelligence in another embodiment of the present application.
[0081] Figure 4 The schematic diagram of the convolutional neural network in an embodiment of the present application.
[0082] Figure 5 The schematic diagram of the GRU recurrent neural network in an embodiment of the present application.
[0083] Figure 6 The block diagram of the auxiliary interpretation device of artificial intelligence in an embodiment of the present application. Specific implementation manners
[0084] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0085] In one embodiment, a terminal is provided, and an application program can be installed on the terminal. The terminal can be a mobile phone, a tablet computer, a notebook computer, a desktop computer, etc. As Figure 1 shown, the terminal includes a processor, an internal memory, a non-volatile storage medium, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor is used to provide computing and control capabilities to support the operation of the entire terminal. The non-volatile storage medium of the terminal stores an operating system and computer-executable instructions, and the computer-executable instructions can be executed by the processor to implement an auxiliary interpretation method based on artificial intelligence provided by the following embodiments. The internal memory in the terminal provides an environment for the operation of the operating system and computer-executable instructions in the non-volatile storage medium. The network interface is used to connect to the network for communication. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen, etc. In this embodiment, the display screen can be used as an output device of the terminal to display various interfaces. For example, an auxiliary interpretation interface can be displayed. The input device can be a touch layer covered on the display screen, or a button, a trackball or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad or mouse, etc., for the user to input various control instructions. For example, in this embodiment, the user can input an information display instruction.
[0086] Those skilled in the art can understand, Figure 1The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the terminal to which the solution of this application is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0087] As Figure 2 shown, in one embodiment, an artificial intelligence-based auxiliary interpretation method is provided. Now, taking the application of this method to Figure 1 the terminal shown as an example for illustration, it specifically includes the following steps:
[0088] Step S20: Initialize the parameters of the first convolutional neural network, the parameters of the second convolutional neural network, and the parameters of the gated recurrent unit (GRU) recurrent neural network.
[0089] As Figure 3 shown, in one embodiment, step S20 specifically includes:
[0090] Step S202: Provide an initial training database, which includes labeled image information with interpretation annotations.
[0091] For example, the initial training database includes labeled image information accurately labeled by doctors, such as the labeled image information of capsule endoscope images. For example, this labeled information is identified and reviewed by more than two professional doctors and finally confirmed by a three-doctor team, including the lesion types labeled for images with lesions, and finally forms / provides the initial training database.
[0092] In addition, the initial training database can also be updated in a timely manner. For example, when the stored interpretation results reach a preset condition, the initial training database is updated according to the interpretation results. The preset conditions include at least one of a preset quantity and a preset time. For example, when the cumulative interpretation results made by doctors or other entities reach 200, or the newly added images with labeled information reach 2000, or after 1 month, half a year, or one year, the initial training database is updated. This feature can ensure that the initial training database is updated in a timely manner according to new interpretation results, restart training / initialization, and thus update each parameter in a timely manner.
[0093] Step S204: Obtain the labeled image information and perform preprocessing.
[0094] Read in the labeled image information, such as pictures or videos labeled by doctors with lesion types. According to the type of the labeled image information (picture or video), then perform preprocessing to obtain the corresponding and required image information. The preprocessing includes: segmentation and format conversion of videos, image format conversion, deformation, flipping, distortion, brightness adjustment, data augmentation, cropping, lighting, etc.; regularization of image data; conversion of the storage format of the final result image data to facilitate subsequent training or processing.
[0095] Step S206: Use the trained convolutional neural network to process the preprocessed labeled image information to obtain the parameters of the first convolutional neural network, the labeled feature vector, and the classified labeled image sequence.
[0096] For example, the preprocessed labeled image information can be processed using the mini-batch gradient descent method to obtain the parameters of the first convolutional neural network.
[0097] For further reference Figure 4, which shows the structure of the convolutional neural network in this embodiment. The convolutional neural network can be arbitrarily selected from the Visual Geometry Group (VGG) network structure, Google network structure, and Residual Network (ResNet) structure. In this example, the GoogleInception-V4 network of the Google network structure is selected. This convolutional neural network has the characteristics of high accuracy and stable classification in this application and can well meet the requirements of the application. As shown in the figure, this convolutional neural network sequentially includes: an input layer (input layer) 402. As an example, here, an input vector of 299*299*3 (i.e., the image information to be judged) is input, indicating 299 pixels in both length and width and including three color channels of R, G, and B; an initial layer (stem layer) 404; a 5×Inception-resnet-A layer 406; a Reduction-A layer 408; a 10×Inception-resnet-B layer 410; a Reduction-B layer 412; a 5×Inception-resnet-C layer 414; an average pooling layer (average pooling layer) 416; a dropout layer (dropout layer) 418. As an example, here, 20% of the data is randomly discarded and 80% of the data is retained; and a fully connected layer (Softmax) 420. The parameters and labeled feature vectors of the first convolutional neural network obtained in this embodiment are the parameters and labeled feature vectors of the first convolutional neural network output by the fully connected layer 420 (i.e., the last layer) of the convolutional neural network. This labeled feature vector is a one-dimensional vector. In this embodiment, the convolutional neural network divides the images in the preprocessed labeled image information into different disease image sequences without changing their order.
[0098] Step S208: Extract the labeled gray-scale features of the classified and labeled image sequence, respectively extract the labeled Laplacian features and labeled Hessian matrix features of the labeled gray-scale features, and use the trained convolutional neural network to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network and the labeled Laplacian feature vector and labeled Hessian matrix feature vector.
[0099] For example, the stochastic gradient descent method can be used to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network.
[0100] Specifically, read the above-mentioned classified and labeled image sequence, and extract the 256-level labeled gray-scale features of the labeled image information; then perform two transformations of Laplacian and Hessian matrix on the labeled gray-scale features respectively to obtain two different gray-scale features, namely the labeled Laplacian features and the labeled Hessian matrix features.
[0101] The calculation formula of the Laplace operator is as follows:
[0102]
[0103] Among them, x and y represent the positions of points on the image on the x-axis and y-axis respectively, and f(x, y) represents the gray value of the point (x, y) in the image.
[0104] The calculation formula of the Hessian matrix is as follows:
[0105]
[0106] In a two-dimensional image, the Hessian matrix is a two-dimensional positive definite matrix. Through the above calculation formula, two eigenvalues λ1 and λ2 of the Hessian matrix and the corresponding two eigenvectors can be obtained. Let λ1 be the eigenvalue with a larger absolute value, that is, |λ1| > |λ2|, then the eigenvalue of the point (x, y) in the image is max(0, λ1). And:
[0107]
[0108]
[0109]
[0110] Among them, I(x, y) represents the gray value of the point (x, y) in the image.
[0111] Then, the above-mentioned Google Inception-V4 network is used again to process the above two gray-scale features respectively, obtain the parameters of the second convolutional neural network, and output two one-dimensional vectors in the fully connected layer, namely the labeled Laplace feature vector and the labeled Hessian matrix feature vector.
[0112] Step S210: Fuse the labeled Laplace feature vector, the labeled Hessian matrix feature vector and the labeled feature vector to form a feature matrix, and quantize the feature matrix and use a GRU recurrent neural network as shown in Figure 5 to process the feature matrix to obtain the parameters of the GRU recurrent neural network.
[0113] Specifically, fuse the labeled feature vector obtained in step S206 and process it using a bidirectional gated recurrent unit (GRU) recurrent neural network; in this example, three feature vectors with the same dimension (i.e., the labeled feature vector, the labeled Laplace feature vector and the labeled Hessian matrix feature vector) form a two-dimensional feature matrix, and after quantizing the feature matrix, it is input into the bidirectional GRU recurrent neural network for training / processing to obtain the parameters of the initial GRU recurrent neural network for subsequent processing.
[0114] Step S30: Obtain the image information to be judged and perform preprocessing.
[0115] Specifically, read in the image information to be judged, such as pictures or videos of lesions collected by a capsule endoscope, and then perform preprocessing according to the type of the image information to be judged (picture or video) to obtain the corresponding and required image information. The preprocessing includes: segmentation and format conversion of videos, image format conversion, deformation, flipping, distortion, brightness adjustment, data augmentation, cropping, illumination, etc.; regularization of image data; conversion of the storage format of the final result image data to facilitate the next step of training or processing.
[0116] Step S40: Use the first convolutional neural network to perform preliminary processing on the preprocessed image information to be judged, and obtain a preliminary feature vector and a classified image sequence to be processed.
[0117] The first convolutional neural network also adopts Figure 4 the convolutional neural network shown in the figure, which will not be elaborated here. In this embodiment, the convolutional neural network classifies the images in the preprocessed image information to be judged into a classified image sequence to be processed, that is, different disease image sequences, but does not change their order.
[0118] For each image, perform preliminary processing or classification, and extract the top five possible lesion probabilities, which are expressed as:
[0119] C out = C1, C2, C3, C4, C5
[0120] where C out represents the preliminary classification, and C1 to C5 represent the top five possible lesion probabilities of the preliminary classification.
[0121] Step S50: Extract the gray-scale features of the classified image sequence to be processed, respectively extract the Laplacian feature and the Hessian matrix feature of the gray-scale features, and use the second convolutional neural network to process the Laplacian feature and the Hessian matrix feature to obtain a Laplacian feature vector and a Hessian matrix feature vector respectively.
[0122] Specifically, read the above-mentioned classified image sequence to be processed, and extract the 256-level gray-scale features of the image information to be processed; then perform two transformations on the gray-scale features, namely the Laplacian and the Hessian matrix, to obtain two different gray-scale features, that is, the Laplacian feature and the Hessian matrix feature.
[0123] Then use the second convolutional neural network, that is, the aforementioned Google Inception-V4 network, to process the above two gray-scale features respectively, and output two one-dimensional vectors in the fully connected layer, that is, the Laplacian feature vector and the Hessian matrix feature vector respectively.
[0124] For each image, perform fine processing or classification, and extract the top five possible lesion probabilities, expressed as:
[0125] J out = J1, J2, J3, J4, J5
[0126] where J out represents fine classification, and J1 to J5 represent the top five possible lesion probabilities of the fine classification.
[0127] Step S60: Combine the Laplace feature vector, the Hessian matrix feature vector, and the preliminary feature vector to form a combined feature vector, and process the combined feature vector using the GRU recurrent neural network.
[0128] Specifically, combine the preliminary feature vector obtained in step S40 and process it using the GRU recurrent neural network. In this example, three feature vectors with the same dimension (i.e., the preliminary feature vector, the Laplace feature vector, and the Hessian matrix feature vector) form a two-dimensional feature matrix. After quantizing the feature matrix, it is input into a bidirectional GRU recurrent neural network for processing. The final output uses C out and J out in a weighted summation manner. The final classification probability of each image is expressed as:
[0129]
[0130] where, Z out represents the final classification. If C i and J i represent the same lesion, the classification of this lesion is added according to the above formula; if they are different, they are calculated separately; the final results are arranged in order of magnitude and saved.
[0131] Step S70: Display the processing results for auxiliary interpretation and store the interpretation results.
[0132] Specifically, display the above processing / classification results, i.e., the preliminary interpretation results, for assisting doctors or other entities in interpretation or further processing to generate the final interpretation results or processing results. The final interpretation results or processing results will be stored. For example, the marked abnormal images will be converted and stored in a certain format.
[0133] The artificial intelligence-based auxiliary interpretation method of the present application combines and uses the preliminary feature vector obtained by initially processing the preprocessed image information to be interpreted using a first convolutional neural network, and the Laplace feature vector and the Hessian matrix feature vector obtained by using a second convolutional neural network to process the Laplace feature and the Hessian matrix feature of the image information to be interpreted, respectively, to classify and process the image to be interpreted, improving the classification and processing efficiency of the image to be interpreted and better assisting in interpretation. And according to the final interpretation result or processing result, the relevant interpretation parameters are continuously optimized, that is, the neural network model used for interpretation is optimized, thereby greatly improving the efficiency of deep learning.
[0134] It should be understood that although Figure 2-3 the steps in the flowchart of Figure 2-3 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0135] As Figure 6 shown, it is a block diagram of an artificial intelligence-based auxiliary interpretation device in an embodiment of the present application.
[0136] An artificial intelligence-based auxiliary interpretation device 60 includes:
[0137] A preprocessing module 602, configured to obtain image information to be interpreted and perform preprocessing.
[0138] A preliminary processing module 604, including a first convolutional neural network 6042, configured to initially process the preprocessed image information to be interpreted to obtain a preliminary feature vector.
[0139] A fine processing module 606, including a second convolutional neural network 6062, configured to extract the gray-scale feature of the preprocessed image information to be interpreted, respectively extract the Laplace feature and the Hessian matrix feature of the gray-scale feature, and use the second convolutional neural network to process the Laplace feature and the Hessian matrix feature to obtain a Laplace feature vector and a Hessian matrix feature vector, respectively.
[0140] The fusion processing module 608 includes a gated recurrent unit (GRU) recurrent neural network 6082, which is used to fuse the Laplace feature vector, the Hessian matrix feature vector, and the preliminary feature vector to form a combined feature vector, and to process the combined feature vector using the GRU recurrent neural network.
[0141] The initialization module 610 is used to initialize the parameters of the first convolutional neural network 6042, the parameters of the second convolutional neural network 6062, and the parameters of the GRU recurrent neural network 6082.
[0142] The interaction module 612 is used to display the processing results for auxiliary interpretation and to store the interpretation results.
[0143] The initialization module 610 further includes:
[0144] An initial training database 6102, which includes labeled image information with interpretation annotations;
[0145] A trained convolutional neural network 6104; and
[0146] A trained GRU recurrent neural network 6106
[0147] In one embodiment, the preprocessing module 602 is further used to obtain the labeled image information and perform preprocessing;
[0148] As Figure 6 shown, in one embodiment, the initialization module 610 of the auxiliary interpretation device 60 is further used to process the preprocessed labeled image information using the trained convolutional neural network 6104 to obtain the parameters of the first convolutional neural network 6042 and the labeled feature vector; extract the labeled gray-scale features of the preprocessed labeled image information, respectively extract the labeled Laplace features and labeled Hessian matrix features of the labeled gray-scale features, and use the trained convolutional neural network 6104 to process the labeled Laplace features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network 6062 and the labeled Laplace feature vector and the labeled Hessian matrix feature vector; and fuse the labeled Laplace feature vector, the labeled Hessian matrix feature vector, and the labeled feature vector to form a feature matrix, and quantize the feature matrix and process the feature matrix using the trained GRU recurrent neural network 6106 to obtain the parameters of the GRU recurrent neural network 6082.
[0149] In one embodiment, the initialization module 610 is further used to:
[0150] Update the initial training database 6102 according to the interpretation results stored by the interaction module 612.
[0151] In one embodiment, the training module 614 is further configured to process the preprocessed labeled image information using mini-batch gradient descent to obtain the parameters of the first convolutional neural network 6042.
[0152] In one embodiment, the training module 614 is further configured to process the labeled Laplacian feature and the labeled Hessian matrix feature using stochastic gradient descent to obtain the parameters of the second convolutional neural network 6062.
[0153] The artificial intelligence-based auxiliary interpretation device of the present application fuses the preliminary feature vector obtained by initially processing the preprocessed image information to be interpreted using the first convolutional neural network, and the Laplacian feature vector and the Hessian matrix feature vector obtained by processing the Laplacian feature and the Hessian matrix feature of the image information to be interpreted using the second convolutional neural network respectively, to classify and process the image to be interpreted, improving the classification and processing efficiency of the image to be interpreted and better assisting in interpretation. And according to the final interpretation result or processing result, the relevant interpretation parameters are continuously optimized, that is, the neural network model used for interpretation is optimized, thereby greatly improving the efficiency of deep learning.
[0154] Each module of the above artificial intelligence-based auxiliary interpretation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the terminal in hardware form or be independent of it, or can be stored in the memory of the terminal in software form, so that the processor can call and execute the operations corresponding to the above respective modules. The processor can be a central processing unit (CPU), a microprocessor, a single-chip microcomputer, etc.
[0155] In one embodiment, a terminal is provided, including a memory and a processor. Instructions are stored in the memory, and when the instructions are executed by the processor, the processor is caused to perform the following steps:
[0156] Initialize the parameters of the first convolutional neural network, the parameters of the second convolutional neural network, and the parameters of the gated recurrent unit (GRU) recurrent neural network;
[0157] Obtain the image information to be interpreted and perform preprocessing;
[0158] Use the first convolutional neural network to initially process the preprocessed image information to be interpreted to obtain a preliminary feature vector;
[0159] Extract the grayscale feature of the preprocessed image information to be interpreted, respectively extract the Laplacian feature and the Hessian matrix feature of the grayscale feature, and use the second convolutional neural network to process the Laplacian feature and the Hessian matrix feature to obtain a Laplacian feature vector and a Hessian matrix feature vector respectively;
[0160] Fuse the Laplacian eigenvector, the Hessian matrix eigenvector, and the preliminary eigenvector to form a combined eigenvector, and process the combined eigenvector using the GRU recurrent neural network; and
[0161] Display the processing result for auxiliary interpretation, and store the interpretation result.
[0162] In one embodiment, when the instructions are executed by the processor, the processor is further caused to perform the following steps:
[0163] Provide an initial training database, the initial training database including labeled image information with interpretation annotations;
[0164] Obtain the labeled image information and perform preprocessing;
[0165] Process the preprocessed labeled image information using a trained convolutional neural network to obtain the parameters of the first convolutional neural network and the labeled eigenvector;
[0166] Extract the labeled gray-scale features of the preprocessed labeled image information, respectively extract the labeled Laplacian features and labeled Hessian matrix features of the labeled gray-scale features, and process the labeled Laplacian features and the labeled Hessian matrix features using the trained convolutional neural network to obtain the parameters of the second convolutional neural network and the labeled Laplacian eigenvector and the labeled Hessian matrix eigenvector; and
[0167] Fuse the labeled Laplacian eigenvector, the labeled Hessian matrix eigenvector, and the labeled eigenvector to form a feature matrix, quantize the feature matrix, and process the feature matrix using a GRU recurrent neural network to obtain the parameters of the GRU recurrent neural network.
[0168] In one embodiment, when the instructions are executed by the processor, the processor is further caused to perform the following steps:
[0169] When the stored interpretation result reaches a preset condition, update the initial training database according to the interpretation result.
[0170] In one embodiment, the preset condition includes at least one of a preset quantity and a preset time.
[0171] In one embodiment, one or more non-volatile storage media storing computer-executable instructions are provided, and when the computer-executable instructions are executed by one or more processors, the one or more processors are caused to perform the following steps:
[0172] Initialize the parameters of the first convolutional neural network, the parameters of the second convolutional neural network, and the parameters of the gated recurrent unit (GRU) recurrent neural network;
[0173] Obtain the image information to be judged and perform preprocessing;
[0174] Use the first convolutional neural network to perform preliminary processing on the preprocessed image information to be judged, and obtain a preliminary feature vector;
[0175] Extract the grayscale features of the preprocessed image information to be judged, respectively extract the Laplacian features and Hessian matrix features of the grayscale features, and use the second convolutional neural network to process the Laplacian features and the Hessian matrix features, respectively obtaining a Laplacian feature vector and a Hessian matrix feature vector;
[0176] Fuse the Laplacian feature vector, the Hessian matrix feature vector and the preliminary feature vector to form a combined feature vector, and use the GRU recurrent neural network to process the combined feature vector; and
[0177] Display the processing result for assisting judgment and store the judgment result.
[0178] In one embodiment, when the instruction is executed by the processor, the processor is further caused to perform the following steps:
[0179] Provide an initial training database, where the initial training database includes labeled image information with judgment annotations;
[0180] Obtain the labeled image information and perform preprocessing;
[0181] Use a training convolutional neural network to process the preprocessed labeled image information, and obtain the parameters of the first convolutional neural network and a labeled feature vector;
[0182] Extract the labeled grayscale features of the preprocessed labeled image information, respectively extract the labeled Laplacian features and labeled Hessian matrix features of the labeled grayscale features, and use the training convolutional neural network to process the labeled Laplacian features and the labeled Hessian matrix features, obtaining the parameters of the second convolutional neural network and a labeled Laplacian feature vector and a labeled Hessian matrix feature vector; and
[0183] Fuse the labeled Laplacian feature vector, the labeled Hessian matrix feature vector and the labeled feature vector to form a feature matrix, quantize the feature matrix and use the GRU recurrent neural network to process the feature matrix, and obtain the parameters of the GRU recurrent neural network.
[0184] In one embodiment, when the instruction is executed by the processor, the processor is further caused to perform the following steps:
[0185] When the stored judgment result reaches a preset condition, update the initial training database according to the judgment result.
[0186] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0187] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0188] The above embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. An artificial intelligence-based auxiliary interpretation method, comprising: Providing an initial training database, where the initial training database includes labeled image information with interpretation labels; Obtaining the labeled image information and performing preprocessing; Using a trained convolutional neural network to process the preprocessed labeled image information to obtain the parameters of the first convolutional neural network, the labeled feature vectors, and the classified labeled image sequences; Extracting the labeled gray-scale features of the classified labeled image sequences, respectively extracting the labeled Laplacian features and the labeled Hessian matrix features of the labeled gray-scale features, and using the trained convolutional neural network to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network, the labeled Laplacian feature vectors, and the labeled Hessian matrix feature vectors; Fusing the labeled Laplacian feature vectors, the labeled Hessian matrix feature vectors, and the labeled feature vectors to form a feature matrix, quantizing the feature matrix, and using a GRU recurrent neural network to process the feature matrix to obtain the parameters of the GRU recurrent neural network; Obtaining the image information to be interpreted and performing preprocessing; Using the first convolutional neural network to preliminarily process the preprocessed image information to be interpreted to obtain the preliminary feature vectors of the image information to be interpreted and the classified image sequences to be interpreted; Extracting the gray-scale features of the classified image sequences to be interpreted, respectively extracting the Laplacian features and the Hessian matrix features of the gray-scale features, and using the second convolutional neural network to process the Laplacian features and the Hessian matrix features to obtain the Laplacian feature vectors and the Hessian matrix feature vectors respectively; Fusing the Laplacian feature vectors, the Hessian matrix feature vectors, and the preliminary feature vectors to form a combined feature vector, and using the GRU recurrent neural network to process the combined feature vector; And Displaying the processing results for auxiliary interpretation and storing the interpretation results.
2. The auxiliary interpretation method according to claim 1, wherein The providing of the initial training database includes: When the stored interpretation results reach a preset condition, updating the initial training database according to the interpretation results.
3. The auxiliary interpretation method according to claim 2, wherein The preset condition includes at least one of a preset quantity and a preset time.
4. The auxiliary interpretation method according to claim 1, wherein The using of the trained convolutional neural network to process the preprocessed labeled image information to obtain the parameters of the first convolutional neural network includes: Using the mini-batch gradient descent method to process the preprocessed labeled image information to obtain the parameters of the first convolutional neural network.
5. The auxiliary interpretation method according to claim 1, wherein The using of the trained convolutional neural network to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network includes: Using the stochastic gradient descent method to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network.
6. The auxiliary interpretation method according to claim 1, wherein The preprocessing includes performing one or more of the following processes on the image information to be interpreted: segmentation, format conversion, deformation, flipping, distortion, brightness adjustment, data augmentation, cropping, illumination, and regularization.
7. An artificial intelligence-based auxiliary interpretation device, comprising: A preprocessing module, configured to obtain the image information to be judged and perform preprocessing; A preliminary processing module, including a first convolutional neural network, configured to perform preliminary processing on the preprocessed image information to be judged, so as to obtain a preliminary feature vector and a classified image sequence to be judged; A fine processing module, including a second convolutional neural network, configured to extract the gray-scale features of the classified image sequence to be judged, respectively extract the Laplacian features and Hessian matrix features of the gray-scale features, and use the second convolutional neural network to process the Laplacian features and the Hessian matrix features to respectively obtain a Laplacian feature vector and a Hessian matrix feature vector; A fusion processing module, including a gated recurrent unit (GRU) recurrent neural network, configured to fuse the Laplacian feature vector, the Hessian matrix feature vector and the preliminary feature vector to form a combined feature vector, and use the GRU recurrent neural network to process the combined feature vector; An initialization module, configured to initialize the parameters of the first convolutional neural network, the parameters of the second convolutional neural network and the parameters of the GRU recurrent neural network; and An interaction module, configured to display the processing result for assisting judgment and store the judgment result; The initialization module further includes: An initial training database, where the initial training database includes the labeled image information with judgment labels; Training the convolutional neural network; and Training the GRU recurrent neural network; The preprocessing module is further configured to obtain the labeled image information and perform preprocessing; The initialization module is further configured to use the trained convolutional neural network to process the preprocessed labeled image information to obtain the parameters of the first convolutional neural network, a labeled feature vector and a classified labeled image sequence; extract the labeled gray-scale features of the classified labeled image sequence, respectively extract the labeled Laplacian features and labeled Hessian matrix features of the labeled gray-scale features, and use the trained convolutional neural network to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network and the labeled Laplacian feature vector and the labeled Hessian matrix feature vector; and fuse the labeled Laplacian feature vector, the labeled Hessian matrix feature vector and the labeled feature vector to form a feature matrix, and quantize the feature matrix and use the trained GRU recurrent neural network to process the feature matrix to obtain the parameters of the GRU recurrent neural network.
8. The auxiliary interpretation device according to claim 7, wherein The initialization module is further configured to: Update the initial training database according to the judgment result stored in the interaction module.
9. The auxiliary interpretation device according to claim 7, wherein The training module is further configured to use the mini-batch gradient descent method to process the preprocessed labeled image information to obtain the parameters of the first convolutional neural network.
10. The auxiliary interpretation device according to claim 7, wherein The training module also is configured to use the stochastic gradient descent method to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network.
11. The auxiliary interpretation device according to claim 7, characterized in that, The preprocessing includes Perform one or more of the following operations on the image information to be judged: segmentation, format conversion, deformation, flipping, distortion, brightness adjustment, data augmentation, cropping, illumination, and regularization.
12. A terminal, comprising a memory and a processor, wherein instructions are stored in the memory, and when the instructions are executed by the processor, the processor performs the following steps: Provide an initial training database, where the initial training database includes labeled image information with judgment labels; Obtain the labeled image information and perform preprocessing; Use a trained convolutional neural network to process the preprocessed labeled image information to obtain the parameters of the first convolutional neural network, labeled feature vectors, and a classified labeled image sequence; Extract the labeled gray-scale features of the classified labeled image sequence, respectively extract the labeled Laplacian features and labeled Hessian matrix features of the labeled gray-scale features, and use the trained convolutional neural network to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network, labeled Laplacian feature vectors, and labeled Hessian matrix feature vectors; Fuse the labeled Laplacian feature vectors, the labeled Hessian matrix feature vectors, and the labeled feature vectors to form a feature matrix, quantize the feature matrix, and use a GRU recurrent neural network to process the feature matrix to obtain the parameters of the GRU recurrent neural network; Obtain the image information to be judged and perform preprocessing; Use the first convolutional neural network to perform preliminary processing on the preprocessed image information to be judged to obtain preliminary feature vectors and a classified image sequence to be judged; Extract the gray-scale features of the classified image sequence to be judged, respectively extract the Laplacian features and Hessian matrix features of the gray-scale features, and use the second convolutional neural network to process the Laplacian features and the Hessian matrix features to obtain Laplacian feature vectors and Hessian matrix feature vectors respectively; Fuse the Laplacian feature vectors, the Hessian matrix feature vectors, and the preliminary feature vectors to form a combined feature vector, and use the GRU recurrent neural network to process the combined feature vector; And Display the processing result for assisting judgment and store the judgment result.
13. The terminal according to claim 12, wherein When the instructions are executed by the processor, the processor further performs the following steps: When the stored judgment result reaches a preset condition, update the initial training database according to the judgment result.
14. The terminal according to claim 13, wherein The preset condition includes at least one of a preset quantity and a preset time.
15. One or more non-volatile storage media storing computer-executable instructions, where the computer-executable instructions, when executed by one or more processors, cause the one or more processors to perform the following steps: Provide an initial training database, where the initial training database includes labeled image information with judgment labels; Obtain the labeled image information and perform preprocessing; Use a trained convolutional neural network to process the preprocessed labeled image information to obtain the parameters of the first convolutional neural network, labeled feature vectors, and a classified labeled image sequence; Extract the labeled grayscale features of the classified and labeled image sequence, respectively extract the labeled Laplacian features and labeled Hessian matrix features of the labeled grayscale features, and use the trained convolutional neural network to process the labeled Laplacian features and the labeled Hessian matrix features to obtain the parameters of the second convolutional neural network and the labeled Laplacian feature vector and the labeled Hessian matrix feature vector; Fuse the labeled Laplacian feature vector, the labeled Hessian matrix feature vector and the labeled feature vector to form a feature matrix, quantize the feature matrix and use the GRU recurrent neural network to process the feature matrix to obtain the parameters of the GRU recurrent neural network; Obtain the image information to be interpreted and perform preprocessing; Use the first convolutional neural network to preliminarily process the preprocessed image information to be interpreted to obtain a preliminary feature vector and classify the image sequence to be interpreted; Extract the grayscale features of the classified image sequence to be interpreted, respectively extract the Laplacian features and Hessian matrix features of the grayscale features, and use the second convolutional neural network to process the Laplacian features and the Hessian matrix features to obtain a Laplacian feature vector and a Hessian matrix feature vector respectively; Fuse the Laplacian feature vector, the Hessian matrix feature vector and the preliminary feature vector to form a combined feature vector, and use the GRU recurrent neural network to process the combined feature vector; And Display the processing results for auxiliary interpretation and store the interpretation results.
16. The non-volatile storage medium according to claim 15, characterized in that, When the instruction is executed by the processor, the processor further executes the following steps: When the stored interpretation result reaches a preset condition, update the initial training database according to the interpretation result.
Citation Information
Patent Citations
License plate recognition method, device and electronic equipment
CN108229474A
Vector road determination method and device
CN109583282A