Image Classification Method, Apparatus, Device, Storage Medium and Program Product

By obtaining the video clips composed of the current frame image of the digestive tract and its adjacent frame images, using the three-dimensional and two-dimensional feature extraction network to extract feature information and integrate it, the problem of low accuracy in the classification of the digestive tract is solved, and high-precision image classification is achieved.

CN119007055BActive Publication Date: 2025-07-22CHANGZHOU UNITED IMAGING HEALTHCARE SURGICAL TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310551945.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-16
Publication Date
2025-07-22
Estimated Expiration
2043-05-16

AI Technical Summary

Technical Problem

In the prior art, the image classification accuracy is not high, especially when classifying digestive tract parts, the classification accuracy of multiple structures is low.

Method used

By obtaining the video clips composed of the current frame image and its adjacent frame images, the three-dimensional and two-dimensional feature extraction networks are used to extract feature information separately, and the fusion process is performed, and finally classification is performed through the image classification network.

Benefits of technology

The accuracy of image classification is improved, especially in the fine-grained classification of digestive tract sites, and the classification accuracy is 89.59%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119007055B_ABST
    Figure CN119007055B_ABST
Patent Text Reader

Abstract

The present application relates to an image classification method, apparatus, device, storage medium, and program product. The method includes: obtaining a current frame image of a current part; determining a video segment corresponding to the current frame image according to the current frame image; the video segment includes the current frame image and adjacent frame images corresponding to the current frame image; respectively extracting feature information in the video segment and the current frame image through a feature extraction network to obtain video feature information and image feature information, and an image classification network performs fusion processing on the video feature information and the image feature information, and classifies the result of the fusion processing to obtain a classification result of the current frame image. Using this method can improve the accuracy of image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image technology, and in particular, to an image classification method, apparatus, device, storage medium, and program product. Background Art

[0002] With the development of artificial intelligence, the process of image classification has received increasing attention. Image classification is an image processing method that differentiates target regions of different categories based on different features reflected in image information.

[0003] Taking image classification as an example, the more image categories obtained by classifying an image according to the feature information in the image, the higher the accuracy of the image classification.

[0004] However, in related technologies, there is a problem of low accuracy in image classification when classifying images. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide an image classification method, apparatus, device, storage medium, and program product that can improve the accuracy of image classification.

[0006] In a first aspect, this application provides an image classification method, which includes:

[0007] Obtain the current frame image of the current part;

[0008] Determine the video segment corresponding to the current frame image according to the current frame image; the video segment includes the current frame image and the adjacent frame images corresponding to the current frame image;

[0009] Extract the feature information in the video segment and the current frame image respectively through a feature extraction network to obtain video feature information and image feature information. The image classification network performs fusion processing on the video feature information and the image feature information, and classifies the result of the fusion processing to obtain the classification result of the current frame image.

[0010] In one embodiment, the feature extraction network includes a three-dimensional feature extraction network and a two-dimensional feature extraction network; extracting the feature information in the video segment and the current frame image respectively through the feature extraction network to obtain video feature information and image feature information includes:

[0011] Extract the feature information in the video segment through the three-dimensional feature extraction network to obtain video feature information; and extract the feature information of the current frame image through the two-dimensional feature extraction network to obtain image feature information.

[0012] In one embodiment, extracting the feature information in the video segment through the three-dimensional feature extraction network to obtain video feature information includes:

[0013] Extract local feature information in the video segment through a three-dimensional feature extraction network to obtain local video feature information; and extract global feature information in the video segment through a three-dimensional feature extraction network to obtain global video feature information;

[0014] Overlay the local video feature information and the global video feature information to obtain video feature information.

[0015] In one embodiment, extract the feature information of the current frame image through a two-dimensional feature extraction network to obtain image feature information; including:

[0016] Extract local feature information in the current frame image through a two-dimensional feature extraction network to obtain local image feature information; and extract global feature information in the current frame image through a two-dimensional feature extraction network to obtain global image feature information;

[0017] Overlay the local image feature information and the global image feature information to obtain image feature information.

[0018] In one embodiment, perform fusion processing on the video feature information and the image feature information to obtain a fusion result, including:

[0019] Extract the reference feature information corresponding to the current frame image from the video feature information;

[0020] Fuse the reference feature information and the image feature information to obtain a fusion result.

[0021] In one embodiment, fuse the reference feature information and the image feature information to obtain a fusion result, including:

[0022] Obtain the similarity between the reference feature information and the image feature information;

[0023] If the similarity is greater than a preset threshold, use the reference feature information or the image feature information as the fusion result;

[0024] If the similarity is less than or equal to the preset threshold, perform weighted summation on the reference feature information and the image feature information, and use the weighted summation result as the fusion result.

[0025] In a second aspect, the present application also provides an image classification device, and the device includes:

[0026] An acquisition module, configured to acquire the current frame image of the current part;

[0027] A determination module, configured to determine a video segment corresponding to the current frame image according to the current frame image; the video segment includes the current frame image and adjacent frame images corresponding to the current frame image;

[0028] A processing module, configured to input the video segment and the current frame image into a preset image classification network, and perform fusion processing on the video feature information of the video segment and the image feature information of the current frame image through the image classification network to obtain a classification result of the current frame image.

[0029] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the content of any one of the image classification methods in the first aspect is implemented.

[0030] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the content of any one of the image classification methods in the first aspect is implemented.

[0031] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the content of any one of the image classification methods in the first aspect is implemented.

[0032] For the above image classification method, device, equipment, storage medium and program product, the current frame image of the current part is obtained, and according to the current frame image, the video segment corresponding to the current frame image is determined; the video segment includes the current frame image and adjacent frame images corresponding to the current frame image. Feature information in the video segment and the current frame image is respectively extracted through a feature extraction network to obtain video feature information and image feature information. The image classification network performs fusion processing on the video feature information and the image feature information, and classifies the result of the fusion processing to obtain the classification result of the current frame image. This method uses the feature extraction network to respectively extract the feature information in the image and the video segment. This process not only considers the feature information of the current frame, but also considers the features of the adjacent frames corresponding to the current frame, and fuses the two feature information. The information content contained in the fused features is more comprehensive. By classifying the result of the fusion processing, the classification result of the current frame image is made more accurate. Description of the Drawings

[0033] Figure 1 It is an application environment diagram of the image classification method in an embodiment;

[0034] Figure 2 It is a schematic flowchart of the image classification method in an embodiment;

[0035] Figure 3Schematic diagram of the process of image classification in an embodiment;

[0036] Figure 4 Schematic diagram of the process of an image classification method in an embodiment;

[0037] Figure 5 Schematic diagram of extracting local feature information in an embodiment;

[0038] Figure 6 Schematic diagram of extracting global feature information in an embodiment;

[0039] Figure 7 Schematic diagram of the process of extracting video feature information in an embodiment;

[0040] Figure 8 Schematic diagram of the process of an image classification method in an embodiment;

[0041] Figure 9 Schematic diagram of extracting local feature information in an embodiment;

[0042] Figure 10 Schematic diagram of extracting global feature information in an embodiment;

[0043] Figure 11 Schematic diagram of the process of extracting the current frame image in an embodiment;

[0044] Figure 12 Schematic diagram of the process of an image classification method in an embodiment;

[0045] Figure 13 Schematic diagram of the process of an image classification method in an embodiment;

[0046] Figure 14 Schematic diagram of the process of an image classification method in an embodiment;

[0047] Figure 15 Structural block diagram of an image classification device in an embodiment. Detailed implementation manners

[0048] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0049] Before introducing the technical solutions of the present application in detail, the background technology of the present application will be introduced first.

[0050] In the process of classifying the image of the current part, a traditional neural network model can be used to classify different structures in the current part into one category to determine the specific structures included in the current part.

[0051] Taking the digestive tract part as an example of the current part, since the digestive tract contains a very large number of parts, just the upper digestive tract includes the pharynx, esophageal segment, gastric antrum, pylorus, upper body, middle body, lower body, incisura angularis, and duodenum, etc.

[0052] In the prior art, before classifying the digestive tract part, it was first proposed to take pictures of the gastric antrum, upper part of the gastric body, and lower part of the gastric body in the digestive tract during an upper gastrointestinal endoscopy (EsophagoGastro Duodenoscopy, EGD). Four images are saved for each field of view, and these four images respectively refer to the anterior wall, lesser curvature of the stomach, posterior wall, and greater curvature of the stomach. In addition, it is also necessary to take pictures of the gastric fundus, stomach, and gastric incisura angularis from a rotating perspective.

[0053] In the process of identifying the upper digestive tract part, a traditional neural network method can be used to classify based on some structures of the digestive tract. Based on 34513 pictures, on average, each structure contains 1000 - 2000 images, and its classification accuracy is 88.70% ± 0.23%, which can classify EGD images to a certain extent.

[0054] When the number of structures to be classified increases, the classification accuracy of the existing neural network for multiple structures is relatively low. To solve the problem of low classification accuracy, an image classification method is proposed in this application. By improving the neural network model, its classification of images becomes more accurate.

[0055] It should be noted that in this application, not only images of multiple structures are used, but also medical videos corresponding to the images of multiple structures. The images and medical videos are both taken using the same device and are operated by the same doctor. The image classification model in this solution can detect images and medical videos in real time at a speed of 30.8 FPS, so as to conduct more in-depth learning and improve the accuracy of image classification. Different methods are used in this application to classify multiple structures, and the classification accuracies are shown in Table 1. The classification accuracy using the neural architecture for fine-grained vision classification (Transformer for Fine-grained, TransFG) is 78.89%, the classification accuracy using the Residual Network 50 (Resnet50) is 73.07%, the classification accuracy using Resnet 101 is 80.44%, the classification accuracy using Resnet 152 is 80.59%, and the classification accuracy of the image classification method provided in this application is 89.59%. It can be seen that the classification accuracy of the method in this application is the highest.

[0056] Method Accuracy TransFG 78.89% Resnet 50 73.07% Resnet 101 80.44% Resnet 152 80.59% Image Classification Method 89.59%

[0057] The image classification method provided by the embodiments of the present application can be applied to, for example, Figure 1 the application environment shown. The computer device may include, but is not limited to, various terminals. For example, the terminal may be an image processor. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data during the image classification process. The network interface of the computer device is used to communicate with other external devices through a network connection. When the computer program is executed by the processor, it implements an image classification method. Among them, the computer device can be implemented by an independent computer device or a computer device cluster composed of multiple computer devices.

[0058] In one embodiment, as Figure 2 shown, an image classification method is provided. Taking the computer device in Figure 1 as an example, the method includes the following steps:

[0059] S201, obtain the current frame image of the current part; the current frame image is collected by an endoscopic device.

[0060] Among them, the current part may be any cavity communicating with the outside of the human body. The current frame medical image may be any frame of medical image scanned by an endoscopic device, and is not limited to the image scanned at the current moment.

[0061] Optionally, the computer device may search for historical images with the same identification information as the identification information of the current part from the image database according to the identification information of the current part, and arbitrarily select one frame of image from the historical images as the current frame image of the current part. Optionally, the computer device may send a scanning instruction to the scanning device. After receiving the scanning instruction, the scanning device scans the current part and sends the scanned image to the computer device, and the computer device may obtain the current frame image. The present embodiment does not limit the manner of obtaining the current frame image of the current part.

[0062] S202, determine the video segment corresponding to the current frame image according to the current frame image; the video segment includes the current frame image and the adjacent frame images corresponding to the current frame image.

[0063] For a frame of image, the information it contains is only local information. It is impossible to classify the current frame of image more precisely and accurately based on local information alone. Therefore, it is also necessary to obtain the context information of multiple consecutive frames of images corresponding to the current frame of image to obtain the global information of the image, so as to classify the current frame of image more accurately.

[0064] In this embodiment, after the computer device obtains the current frame of image, it can search for multiple frames of images adjacent to the sequence identifier from the image database according to the sequence identifier of the current frame of image, and use the segment composed of the adjacent multiple frames of images and the current frame of image as a video segment. For example, if the current frame of image is the 4th frame, the adjacent frames of images can be the 2nd, 3rd, 5th, and 6th frames, and the images from the 2nd to 6th frames are used as the video segment corresponding to the current frame of image. It should be noted that the number of adjacent frames of images can be any number, and the image segment is determined according to the number of adjacent frames of images.

[0065] S203. Respectively extract the feature information in the video segment and the current frame of image through the feature extraction network to obtain video feature information and image feature information. The image classification network performs fusion processing on the video feature information and the image feature information, and classifies the result of the fusion processing to obtain the classification result of the current frame of image.

[0066] In this embodiment, the computer device can extract the video feature information in the video segment through the feature extraction network. This video feature information is three-dimensional feature information in the time series. Also, the computer device can extract the image feature information in the current frame of image through the feature extraction network. This image feature information is two-dimensional feature information. Then, the image classification network extracts the two-dimensional feature information corresponding to the image feature information from the three-dimensional feature information, and fuses the extracted two-dimensional feature information with the image feature information to obtain more abundant fusion feature information. After classifying the fusion feature information, the classification result of the current frame of image obtained is more accurate. This classification result indicates the specific structure in the current part to which the current frame of image belongs. For example, the classification result can be the center of the gastric angle, the anterior angle of the gastric angle, or the posterior angle of the gastric angle, etc.

[0067] Figure 3 is a schematic flowchart of image classification. In Figure 3 both the current frame of image and the corresponding video segment are input into the image classification model. The feature extraction network in the image classification model extracts the feature information of the video segment to obtain video feature information, and extracts the feature information in the current frame of image through the feature extraction network to obtain image feature information. The video feature information and the image feature information are subjected to fusion processing, and the result of the fusion processing is classified by the classifier to obtain the classification result of the current frame of image.

[0068] It should be noted that both the feature extraction network and the image classification network belong to the image classification model. The image classification model is trained with a large number of sample images and sample video clips of the current part. Compared with the traditional method that only considers the features of a single-frame image, the training process of this image classification model takes into account the features in the time series. Therefore, the trained image classification model is more accurate, and the classification process of the current-frame image by this trained image classification model is also more accurate.

[0069] In the above image classification method, the current-frame image of the current part is obtained, and based on the current-frame image, the video clip corresponding to the current-frame image is determined; the video clip includes the current-frame image and the adjacent frame images corresponding to the current-frame image. The feature extraction network is used to extract the feature information in the video clip and the current-frame image respectively, obtaining video feature information and image feature information. The image classification network performs fusion processing on the video feature information and the image feature information, and classifies the result of the fusion processing to obtain the classification result of the current-frame image. This method uses the feature extraction network to extract the feature information in the image and the video clip respectively. This process not only considers the feature information of the current frame, but also considers the features of the adjacent frames corresponding to the current frame, and fuses the two feature information. The information content contained in the fused features is more comprehensive. By classifying the result of the fusion processing, the classification result of the current-frame image is more accurate.

[0070] Based on the above embodiments, when the feature extraction network includes a three-dimensional feature extraction network and a two-dimensional feature extraction network, this embodiment is an introduction and explanation of Figure 2 the relevant content of step S203 in

[0071] "Extract the feature information in the video clip and the current-frame image respectively through the feature extraction network to obtain video feature information and image feature information". The above step S203 may include the following content: extract the feature information in the video clip through the three-dimensional feature extraction network to obtain video feature information; and extract the feature information in the current-frame image through the two-dimensional feature extraction network to obtain image feature information.

[0072] In the above image classification method, feature information in a video clip is extracted through a three-dimensional feature extraction network to obtain video feature information; and feature information of the current frame image is extracted through a two-dimensional feature extraction network to obtain image feature information. In this method, features of the video clip are extracted through the three-dimensional feature extraction network, and features of the current frame image are extracted through the two-dimensional feature extraction network. Corresponding feature information is extracted through different feature extraction networks, making the feature extraction process more targeted and the extracted features more accurate.

[0073] Based on the above embodiments, this embodiment introduces and explains the relevant content of "extracting feature information in a video clip through a three-dimensional feature extraction network to obtain video feature information" in the above steps. As Figure 4 shown, the above steps may include the following content:

[0074] S301, extracting local feature information in a video clip through a three-dimensional feature extraction network to obtain local video feature information; and extracting global feature information in the video clip through the three-dimensional feature extraction network to obtain global video feature information.

[0075] Among them, the local feature information of a pixel point refers to the feature information jointly determined by the features in 8 directions adjacent to the pixel point, and the global feature information of a pixel point refers to the feature information determined by the features of all pixel points.

[0076] In this embodiment, for any pixel point, the computer device can use the three-dimensional feature extraction network to extract the feature information of the pixel point, and extract the feature information in 8 directions adjacent to the pixel point. By comprehensively considering the feature information of the pixel point, the feature information in 8 adjacent directions, and the feature information of the adjacent frames corresponding to the pixel point, the local feature information of the pixel point is obtained. In this way, the local feature information of all pixel points can be extracted. It should be noted that when extracting local feature information, by adding a time channel T to the traditional Convolutional Neural Networks (CNN), that is, using a convolutional kernel of W*H*T for convolutional operation. In this way, in addition to obtaining the local information of a frame of image, the context information between adjacent frames of images can also be obtained. By superimposing the local information and the context information, the local feature information in the video clip can be obtained. Wherein, W represents the length of any frame of image in the video clip, and H represents the width of any frame of image in the video clip. Figure 5 It is a schematic diagram for extracting local feature information. Select a pixel point whose features need to be extracted from the video clip, determine the adjacent pixel points of the pixel point and the corresponding pixel points in the adjacent frame images, and obtain the local feature information of the pixel point.

[0077] Further, the computer device can add a self-attention mechanism to the CNN. The advantage of the CNN lies in its ability to capture adjacent local information, while the advantage of the self-attention mechanism is that it can better obtain the global information of the image. Therefore, combining the CNN with the self-attention mechanism can combine local information and global information, better improving the feature expression ability of the model. During the feature extraction process, the global feature information in the video segment can be accurately obtained. Compared with the local feature extraction process, the extraction of global feature information only determines the information of a pixel point through the global feature information. In this way, the extracted features are more comprehensive and the information is richer. Among them, the process of combining the CNN with the self-attention mechanism for processing the video segment is as follows: For any frame of the image in the video segment, the features of each pixel point in the image will be calculated with the features of all other pixel points. The weights of other pixel points closer to the pixel point account for a larger proportion, and the weights of other pixel points farther from the pixel point account for a smaller proportion. The feature information of the pixel point can be obtained through the weights and the features of the pixel points. In this way, the global video feature information of the video segment can be obtained. Figure 6 It is a schematic diagram for extracting global feature information. Select the pixel points whose features need to be extracted from the video segment, and use the feature information of all pixel points to determine the global feature information of the pixel points.

[0078] S302, superimpose the local video feature information and the global video feature information to obtain video feature information.

[0079] In this embodiment, since the local video feature information focuses more on the extraction of local features and the extracted features are more accurate, but this feature does not consider the global features and has certain limitations. The global video feature information focuses more on the extraction of global features, and the extracted features are comprehensively considered with global information, but this global feature does not consider some more specific local feature information. Therefore, in order to make the obtained video feature information more comprehensive and accurate, after the computer device obtains the local video feature information and the global video feature information, it superimposes the local video feature information and the global video feature information. Only one of the overlapping parts of the local video feature information and the global video feature information is taken, and for different parts, all need to be retained, so that the obtained video feature information is more comprehensive and more accurate.

[0080] Figure 7 It is a schematic flowchart for extracting video feature information in an embodiment. The local feature information and the global feature information in the video segment are extracted through a W*H*T convolution kernel, and the two feature information are superimposed and processed to obtain video feature information.

[0081] In the above image classification method, local feature information in a video clip is extracted through a three-dimensional feature extraction network to obtain local video feature information; and global feature information in the video clip is extracted through the three-dimensional feature extraction network to obtain global video feature information. The local video feature information and the global video feature information are superimposed to obtain video feature information. This method extracts local video feature information and global video feature information in the video clip through the three-dimensional feature extraction network respectively, extracts different feature information from different perspectives, and performs a superimposing process on the local video feature information and the global video feature information, so that the obtained video feature information includes more comprehensive and accurate features.

[0082] Based on the above embodiments, this embodiment introduces and explains the relevant content of "extracting the feature information of the current frame image through a two-dimensional feature extraction network to obtain image feature information" in the above steps. As Figure 8 shown, the above steps may include the following content:

[0083] S401, extract local feature information in the current frame image through a two-dimensional feature extraction network to obtain local image feature information; and extract global feature information in the current frame image through the two-dimensional feature extraction network to obtain global image feature information.

[0084] In this embodiment, for any pixel point, the computer device can use the two-dimensional feature extraction network to extract the feature information of this pixel point, and, extract the feature information in 8 adjacent directions of this pixel point. By comprehensively considering the feature information of this pixel point, the feature information in 8 adjacent directions, and the feature information of the adjacent frames corresponding to this pixel point, the local feature information of this pixel point is obtained. In this way, the local feature information of all pixel points can be extracted. It should be noted that when extracting local feature information, a traditional CNN uses a convolution kernel of W*H for convolution operation, so that the local feature information in the current frame image can be obtained. Figure 9 FIG. is a schematic diagram for extracting local feature information. Select the pixel points whose features need to be extracted from the current frame image, determine the adjacent pixel points of this pixel point and the corresponding pixel points in the adjacent frame images, and obtain the local feature information of this pixel point.

[0085] Further, similar to extracting the global information in the video clip, a self-attention mechanism can be added to the traditional CNN algorithm to accurately obtain the global feature information in the current frame image during the feature extraction process. Figure 10 FIG. is a schematic diagram for extracting global feature information. Select the pixel points whose features need to be extracted from the current frame image, and use the feature information of all pixel points to determine the global feature information of this pixel point.

[0086] S402. Superimpose the local image feature information and the global image feature information to obtain the image feature information.

[0087] In this embodiment, after the computer device obtains the local video feature information and the global video feature information, it superimposes the local video feature information and the global video feature information. Only one of the overlapping parts of the local video feature information and the global video feature information is taken, and for different parts, all need to be retained, so that the obtained video feature information is more comprehensive and more accurate.

[0088] Figure 11 FIG. is a schematic flow chart for extracting the current frame image in an embodiment. The local feature information and the global feature information in the current frame image are extracted through a W*H convolution kernel, and the two feature information are superimposed to obtain the video feature information.

[0089] In the above image classification method, the local feature information in the current frame image is extracted through a two-dimensional feature extraction network to obtain the local image feature information; and the global feature information in the current frame image is extracted through a two-dimensional feature extraction network to obtain the global image feature information. The local image feature information and the global image feature information are superimposed to obtain the image feature information. This method extracts the local image feature information and the global image feature information in the current frame image respectively through a two-dimensional feature extraction network, extracts different feature information from different angles, and superimposes the local image feature information and the global image feature information, so that the obtained image feature information includes more comprehensive and more accurate features.

[0090] Based on the above embodiment, this embodiment introduces and explains the relevant content of "fusing the video feature information and the image feature information to obtain the fusion result" in step S203 above.

[0091] As Figure 12 shown, the above step S203 may include the following contents:

[0092] S501. Extract the reference feature information corresponding to the current frame image from the video feature information.

[0093] In this embodiment, since the video feature information is three-dimensional feature information, which includes the feature information corresponding to multiple frames of images, and the image feature information is two-dimensional feature information, in order to take into account the feature information of the video segment, it is necessary to extract the feature information corresponding to the current frame image from the video feature information and use this feature information as the reference feature information. For example, the video segment includes 7 frames of images, and the current frame image belongs to the 4th frame among all the frames in the video segment. The current frame image and the 4th frame image in the video segment are the same frame image in the time dimension. Since the video feature information is composed of the feature information of 7 frames of images, and the feature information of the 4th frame image in the video segment takes into account the feature information of adjacent frames, therefore, the feature information of the 4th frame image in the video segment may be different from the feature information of the current frame image. Therefore, it is necessary to extract the reference feature information corresponding to the current frame image from the video feature information, and this reference feature information is the feature information corresponding to the 4th frame image in the video feature information.

[0094] S502. Fuse the reference feature information and the image feature information to obtain a fusion result.

[0095] In this embodiment, in order to obtain more comprehensive information in the current frame image, it is necessary to fuse the reference feature information and the image feature information to obtain more comprehensive and accurate feature information. Therefore, the computer device can fuse the reference feature information and the image feature information, fix the parts with the same feature information, and fuse the parts with different features in a weighted summation manner to obtain the fusion result of the reference feature information and the image feature information. This fusion result can more comprehensively and accurately reflect the feature information of the current frame image.

[0096] It should be noted that the two-dimensional feature extraction network and the three-dimensional feature extraction network adopt the idea of the teacher-student network, taking the two-dimensional feature extraction network as the teacher network and the three-dimensional feature extraction network as the student network. The feature information output by the teacher network and the feature information output by the student network are weighted and fused to output the final fusion result.

[0097] In the above image classification method, the reference feature information is extracted from the video feature information; the reference feature information and the image feature information are information of the same feature. The reference feature information and the image feature information are fused to obtain a fusion result. This method extracts the reference feature information of the same feature as the image feature information from the video feature information, and fuses this reference information with the image feature information. The obtained fusion result includes not only the local image feature and the global image feature of the current frame image, but also the features of adjacent frame images, obtaining a more comprehensive fusion feature.

[0098] Based on the above embodiments, this embodiment introduces and explains the relevant content of "performing fusion processing on the video feature information and the image feature information to obtain a fusion result" in the above step S502.

[0099] As Figure 13 shown, the above step S502 may include the following content:

[0100] S601, obtain the similarity between the reference feature information and the image feature information.

[0101] In this embodiment, the computer device may determine whether the reference feature information of each pixel point is the same as the corresponding image feature information, obtain the number of pixel points with the same feature information, calculate the ratio of the number of pixel points with the same feature information to the total number of pixel points, and determine this ratio as the similarity between the reference feature information and the image feature information. For example, if the total number of pixel points is 100 and the number of pixel points with the same feature is 80, then the similarity between the reference feature information and the image feature information is 0.8.

[0102] S602, if the similarity is greater than the preset threshold, then use the reference feature information or the image feature information as the fusion result.

[0103] In this embodiment, the computer device may compare this similarity with the preset threshold. If the similarity is greater than the preset threshold, it means that the reference feature information and the image feature information are basically the same, and only a very small part of the features are different. Then, either the reference feature information can be used as the fusion result, or the reference feature information can be used as the fusion result.

[0104] S603, if the similarity is less than or equal to the preset threshold, then perform weighted summation on the reference feature information and the image feature information, and use the weighted summation result as the fusion result.

[0105] In this embodiment, when the similarity is less than or equal to the preset threshold, it means that there are some differences between the reference feature information and the image feature information. Therefore, it is necessary to perform weighted summation on the reference feature information and the image feature information, and use the weighted summation result as the fusion result. The weighting coefficients of the reference feature information and the image feature information can be obtained based on historical experience. For example, the weighting coefficient of the reference feature information can be 0.4, and the weighting coefficient of the image feature information can be 0.6.

[0106] In the above image classification method, the similarity between the reference feature information and the image feature information is obtained. If the similarity is greater than a preset threshold, the reference feature information or the image feature information is used as the fusion result. If the similarity is less than or equal to the preset threshold, the reference feature information and the image feature information are weighted and summed, and the weighted sum result is used as the fusion result. By obtaining the similarity between the reference feature information and the image feature information and comparing this similarity with the preset threshold, this method can obtain the fusion result in different ways, making the obtained fusion result more accurate.

[0107] In one embodiment, the image classification method is introduced in detail as follows. Figure 14 As shown, the method may include:

[0108] S701, obtaining the current frame image of the current part;

[0109] S702, determining the video segment corresponding to the current frame image according to the current frame image;

[0110] S703, extracting the local feature information in the video segment through a three-dimensional feature extraction network to obtain local video feature information; and extracting the global feature information in the video segment through a three-dimensional feature extraction network to obtain global video feature information;

[0111] S704, superimposing the local video feature information and the global video feature information to obtain video feature information;

[0112] S705, extracting the local feature information in the current frame image through a two-dimensional feature extraction network to obtain local image feature information; and extracting the global feature information in the current frame image through a two-dimensional feature extraction network to obtain global image feature information;

[0113] S706, superimposing the local image feature information and the global image feature information to obtain image feature information;

[0114] S707, extracting the reference feature information corresponding to the current frame image from the video feature information;

[0115] S708, obtaining the similarity between the reference feature information and the image feature information;

[0116] S709, if the similarity is greater than a preset threshold, using the reference feature information or the image feature information as the fusion result;

[0117] S710, if the similarity is less than or equal to the preset threshold, performing weighted summation on the reference feature information and the image feature information, and using the weighted sum result as the fusion result.

[0118] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0119] Based on the same inventive concept, an embodiment of the present application further provides an image classification device for implementing the above-mentioned image classification method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following image classification devices can refer to the limitations on the image classification method in the above text, and will not be repeated here.

[0120] In one embodiment, as Figure 15 shown, an image classification device is provided, including: an acquisition module 11, a determination module 12, and a processing module 13, where:

[0121] The acquisition module 11 is used to acquire the current frame image of the current part;

[0122] The determination module 12 is used to determine the video segment corresponding to the current frame image according to the current frame image; the video segment includes the current frame image and the adjacent frame images corresponding to the current frame image;

[0123] The processing module 13 is used to input the video segment and the current frame image into a preset image classification network, and perform fusion processing on the video feature information of the video segment and the image feature information of the current frame image through the image classification network to obtain the classification result of the current frame image.

[0124] In one embodiment, the above classification model includes a first extraction unit, where

[0125] The first extraction unit is used to extract the feature information in the video segment through a three-dimensional feature extraction network to obtain video feature information; and extract the feature information of the current frame image through a two-dimensional feature extraction network to obtain image feature information.

[0126] In one embodiment, the first extraction unit is further configured to extract local feature information in a video segment through a three-dimensional feature extraction network to obtain local video feature information; and extract global feature information in the video segment through the three-dimensional feature extraction network to obtain global video feature information; superimpose the local video feature information and the global video feature information to obtain video feature information.

[0127] In one embodiment, the first extraction unit is further configured to extract local feature information in the current frame image through a two-dimensional feature extraction network to obtain local image feature information; and extract global feature information in the current frame image through the two-dimensional feature extraction network to obtain global image feature information; superimpose the local image feature information and the global image feature information to obtain image feature information.

[0128] In one embodiment, the processing module further includes: a second extraction unit and a fusion unit, where:

[0129] The second extraction unit is configured to extract reference feature information corresponding to the current frame image from the video feature information;

[0130] The fusion unit is configured to fuse the reference feature information and the image feature information to obtain a fusion result.

[0131] In one embodiment, the fusion unit is further configured to obtain the similarity between the reference feature information and the image feature information; if the similarity is greater than a preset threshold, use the reference feature information or the image feature information as the fusion result; if the similarity is less than or equal to the preset threshold, perform weighted summation on the reference feature information and the image feature information, and use the weighted summation result as the fusion result.

[0132] Each module in the above image classification device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.

[0133] In one embodiment, a computer device is provided, including a memory and a processor, where a computer program is stored in the memory, and when the processor executes the computer program, it implements the content in any one of the above image classification method embodiments.

[0134] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the content in any one of the above image classification method embodiments.

[0135] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the content in any one of the above-described image classification method embodiments.

[0136] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties.

[0137] Those of ordinary skill in the art can understand that all or part of the processes in the above-described method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-described method embodiments. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memories can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0138] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0139] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. An image classification method, characterized in that, The method includes: Obtaining a current frame image of a current part; the current frame image is acquired by an endoscope device; the current part is a digestive tract part; Determining a video segment corresponding to the current frame image according to the current frame image; the video segment includes the current frame image and adjacent frame images corresponding to the current frame image; Respectively extracting feature information from the video segment and the current frame image through a feature extraction network to obtain video feature information and image feature information, extracting reference feature information corresponding to the current frame image from the video feature information through an image classification network, and obtaining a similarity between the reference feature information and the image feature information; if the similarity is greater than a preset threshold, using the reference feature information or the image feature information as a result of fusion processing; if the similarity is less than or equal to the preset threshold, performing weighted summation on the reference feature information and the image feature information, using the weighted summation result as the result of the fusion processing, and classifying the result of the fusion processing to obtain a classification result of the current frame image; the classification result represents a specific structure in the digestive tract part to which the current frame image belongs.

2. The method according to claim 1, characterized in that, The feature extraction network includes a three-dimensional feature extraction network and a two-dimensional feature extraction network; the respectively extracting feature information from the video segment and the current frame image through the feature extraction network to obtain video feature information and image feature information includes: Extracting feature information in the video segment through the three-dimensional feature extraction network to obtain the video feature information; and extracting feature information in the current frame image through the two-dimensional feature extraction network to obtain the image feature information.

3. The method according to claim 2, characterized in that The extracting feature information in the video segment through the three-dimensional feature extraction network to obtain the video feature information includes: Extracting local feature information in the video segment through the three-dimensional feature extraction network to obtain local video feature information; and extracting global feature information in the video segment through the three-dimensional feature extraction network to obtain global video feature information; Superimposing the local video feature information and the global video feature information to obtain the video feature information.

4. The method according to claim 2, wherein The extracting feature information in the current frame image through the two-dimensional feature extraction network to obtain the image feature information includes: Extracting local feature information in the current frame image through the two-dimensional feature extraction network to obtain local image feature information; and extracting global feature information in the current frame image through the two-dimensional feature extraction network to obtain global image feature information; Superimposing the local image feature information and the global image feature information to obtain the image feature information.

5. The method according to claim 1, characterized in that The current part is any cavity communicating with the outside of the human body.

6. The method according to claim 1, characterized in that The current frame image is any medical image scanned by an endoscope device.

7. An image classification device, characterized in that, The device includes: An acquisition module, configured to acquire a current frame image of a current part; the current frame image is acquired by an endoscopic device; the current part is a digestive tract part; A determination module, configured to determine a video segment corresponding to the current frame image according to the current frame image; the video segment includes the current frame image and adjacent frame images corresponding to the current frame image; A processing module, configured to respectively extract feature information in the video segment and the current frame image through a feature extraction network to obtain video feature information and image feature information, extract reference feature information corresponding to the current frame image from the video feature information through an image classification network, and obtain a similarity between the reference feature information and the image feature information; if the similarity is greater than a preset threshold, then use the reference feature information or the image feature information as the result of fusion processing; if the similarity is less than or equal to the preset threshold, then perform weighted summation on the reference feature information and the image feature information, use the weighted summation result as the result of the fusion processing, and classify the result of the fusion processing to obtain a classification result of the current frame image; the classification result represents a specific structure in the digestive tract part to which the current frame image belongs.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Video identification method and device based on adaptive reasoning

    CN114241360A

  • Video-based focus classification method and device, electronic equipment and medium

    CN114663372A

  • Video representation method, video classification method, electronic equipment and storage medium

    CN114996508A

  • Video processing method and apparatus

    WO2022179087A1