Improved neural network model for image recognition and image recognition method and system
Through the improved neural network model, dual-core convolution and Unified-IoU evaluation are adopted to solve the misdiagnosis and misdiagnosis of traditional bronchoscopy, efficient and accurate image recognition and lesion feature extraction, and the diagnostic efficiency of bronchoscopy detection is improved.
Patent Information
- Application Number
- CN202510482355.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional bronchoscopy relies on the experience of doctors and naked-eye observation, which is prone to misdiagnosis or misdiagnosis, and the large number of images leads to diagnostic fatigue, affecting diagnostic accuracy and efficiency.
A modified neural network model is constructed, dual-core convolution is used to process input feature channels, and a prediction box of different quality is trained. Unified-IoU is used to evaluate the prediction box to achieve multi-branch feature learning and efficient image recognition.
It improves the accuracy and efficiency of image recognition, can efficiently label weak features, is suitable for complex situations of bronchoscopy, and enhances the model's adaptability and detection efficiency.
Smart Images

Figure CN120339800A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of neural network models and image recognition, and particularly to an improved neural network model for image recognition, an image recognition method, and a system. Background Art
[0002] Bronchoscopes are a widely used diagnostic and therapeutic tool in clinical medicine, which can directly observe and collect images of bronchi and related parts by inserting into the respiratory tract. This method plays an important role in the detection, diagnosis, and treatment of lung diseases. However, traditional bronchoscopy mainly relies on doctors' experience and visual observation, and its diagnostic results are often limited by doctors' subjective judgments. There may be a risk of missed diagnosis or misdiagnosis, especially in the identification of complex or early-stage lesions. In addition, due to the large number of images generated by bronchoscope examinations, manual analysis is time-consuming and prone to diagnostic fatigue, further affecting the diagnostic accuracy and efficiency.
[0003] In recent years, the development of artificial intelligence (AI) technology, especially deep learning, has provided new solutions for the intelligent diagnosis of medical images. AI-based image analysis technology can extract key features from a large number of bronchoscope detection images by training a deep neural network model, and achieve automatic classification and early prediction of diseases. Compared with traditional methods, AI technology can greatly improve the efficiency and accuracy of image analysis, while reducing the burden on doctors.
[0004] In the application of bronchoscope images, using artificial intelligence technology to early identify and classify respiratory tract lesions such as lung cancer, chronic obstructive pulmonary disease (COPD), tuberculosis, etc. can significantly improve the diagnostic rate and reduce the risk brought by disease progression. However, the current technology still faces many challenges, such as how to efficiently annotate medical images, how to extract weak lesion features, and how to ensure the generalization ability of the model among different devices and patients. Therefore, developing an intelligent diagnosis system that combines bronchoscope images and artificial intelligence technology can not only make up for the deficiencies of traditional methods, but also provide strong technical support for precision medicine and disease management. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that traditional bronchoscopy mainly relies on doctors' experience and visual observation, and its diagnostic results are often limited by doctors' subjective judgments. There may be a risk of missed diagnosis or misdiagnosis in the identification of complex or early-stage lesions. At the same time, the large number of images generated by bronchoscope examinations makes manual analysis time-consuming and prone to diagnostic fatigue, further affecting the diagnostic accuracy and efficiency. In view of the above defects of the prior art, an improved neural network model for image recognition, an image recognition method, and a system are provided.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0007] Construct an improved neural network model for image recognition, including:
[0008] Obtain a real-time image of the first size;
[0009] Use the real-time image of the first size as the input feature map of the neural network model;
[0010] Simultaneously process the same input feature map channels using a first convolution kernel and a second convolution kernel;
[0011] Divide the input feature map channels to achieve multi-branch feature learning;
[0012] Concatenate the features extracted from all branches to obtain the finally output feature map;
[0013] Form an improved neural network model by aggregating multiple finally output feature maps.
[0014] Preferably, in the process of simultaneously processing the same input feature map channels using the first convolution kernel and the second convolution kernel, it further includes:
[0015] Use grouped convolution to arrange the convolution filters.
[0016] Preferably, in the process of simultaneously processing the same input feature map channels using the first convolution kernel and the second convolution kernel, it further includes:
[0017] Capture more spatial information of the features through the first convolution kernel, and perform interaction and information integration between the feature channels through the second convolution kernel.
[0018] Preferably, after forming an improved neural network model by aggregating multiple finally output feature maps, it further includes:
[0019] Use a first prediction box and a second prediction box to train and learn the model. The first prediction box is applied in the early stage of training, and the second prediction box is applied in the later stage of training;
[0020] Preferably, the first prediction box and the second prediction box are the geometric coincidence degrees between the prediction box and the ground truth box. The geometric coincidence degree lower than the set value is the first prediction box, and the set coincidence degree higher than the set value is the second prediction box.
[0021] Preferably, the set value is adjusted according to the ground truth box during the training process, and the prediction box includes at least one characteristic of shape, position or size.
[0022] Construct an image recognition method, including:
[0023] Obtain a real-time image of the first size;
[0024] Use the real-time image of the first size as an input value in the improved neural network model for image recognition according to any one of claims 1-6 to perform image recognition, and judge and mark the recognition information.
[0025] Preferably, in the feature map of the real-time image of the second size obtained by extracting features from the real-time image of the second size, it further includes:
[0026] Obtain the feature map in the real-time image of the second size by performing convolution and downsampling on the real-time image of the second size, and the size of the feature map is the first size;
[0027] In the obtaining of the real-time image of the first size, it further includes:
[0028] Perform noise filtering and cropping on the real-time image obtained by the endoscope, and perform photoelectric conversion and decoding to obtain the real-time image of the first size.
[0029] Construct an image recognition system, including:
[0030] An image acquisition module that acquires real-time image information and performs preprocessing;
[0031] An image recognition module that uses the same input feature map channels for dual-core convolution, divides the input channels to achieve multi-branch learning, cascades the features extracted from all branches to obtain the finally output feature image, thereby obtaining an improved neural network model, trains and learns the improved neural network model using prediction boxes of different qualities, obtains the trained neural network model, and uses the trained and learned model for image recognition;
[0032] An image reporting module that reports the recognized image information.
[0033] Construct a storage medium, on which program instructions are stored, and characterized in that: the program instructions are used to execute an image recognition method as described above when running.
[0034] The beneficial effects of the present invention are: the expressive power of the improved neural network model is improved, the computational efficiency is faster, and images and weak features can be efficiently annotated, which is suitable for real-time images of respiratory tract parts obtained by bronchoscopy, the recognition ability of small targets and weak lesion features under complex conditions of bronchoscopy, and the huge number of images generated by bronchoscopy detection, and the improved neural network can be trained to improve the adaptability and detection efficiency of the model, so as to realize large-scale extraction of key features in bronchoscopy detection images. The input real-time image is used as the input value of the improved neural network model, and the size of the input value is modified to process the same input feature channel using dual-core convolution to realize multi-branch feature learning, and the features extracted by all branches are cascaded to obtain the final output feature image, which not only captures more spatial information during feature extraction, but also interacts and integrates information between feature channels without increasing the computational complexity of too many parameters, greatly reducing the number of parameters, thereby improving the operation speed. Then, the model is trained using prediction frames of different qualities. Low-quality detection frames are used to reduce obviously wrong predictions, and high-quality prediction frames are used to improve accuracy. The characteristics of other prediction frames can also be combined to provide a more fine-grained evaluation standard to obtain a trained neural network model. The improved neural network model is used to identify real-time images and determine whether there are abnormalities. If there are abnormalities, the abnormal location is marked, thereby providing high-level results for intelligent recognition of bronchoendoscopic images and improving efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the present invention will be further described below in conjunction with the accompanying drawings and embodiments. The drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work:
[0036] Figure 1 A schematic diagram of a flow chart of an image recognition method according to a preferred embodiment of the present invention;
[0037] Figure 2 A schematic diagram of an improved neural network model of a preferred embodiment of the present invention;
[0038] Figure 3 A schematic diagram showing a comparison between an improved neural network model and an original neural network model according to a preferred embodiment of the present invention;
[0039] Figure 4 A schematic diagram of an image recognition system according to a preferred embodiment of the present invention;
[0040] Figure 5 Schematic diagram of the structure of a smart terminal according to a preferred embodiment of the present invention. Detailed implementation manners
[0041] In order to make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] An improved neural network model for image recognition and an image recognition method according to a preferred embodiment of the present invention; as Figure 1 shown, is a schematic flowchart of an image recognition method provided by an embodiment of the present invention. This method can be executed by a device, and the device can be implemented by software and / or hardware.
[0043] Specifically, in this embodiment, the image recognition method includes:
[0044] S10: Obtain real-time image information and perform preprocessing;
[0045] In a preferred embodiment of the present invention, a camera system is used to obtain real-time image information. For example, a bronchoscope can be used to collect real-time image information of a patient's respiratory tract, and it can also be used to obtain image information using a camera in other scenarios. In this application, the example of using a bronchoscope to obtain real-time image information of the respiratory tract will be described in detail. And the obtained real-time image is preprocessed. The preprocessing includes photoelectric conversion, and the converted data enters the endoscope handle for further decoding to form the original image format. And during the conversion process, noise filtering and cropping and other processing are performed on the obtained image information, so that the size of the obtained original image is 1280X720.
[0046] S20: Modify the obtained real-time image as the input image of the yolo10n model;
[0047] In a preferred embodiment of the present invention, a neural network model is used to identify target information in image information. Specifically, the yolo10n model is adopted. However, this model uses an image with a size of 640X640 as the input training image during training. Since images of this size often cannot meet the needs of actual production. Therefore, the obtained real-time image information is modified. First, the original image is filled so that the filled size becomes 1280X1280. Of course, it can also be filled to a larger image size according to needs. Such as Figure 2As shown, a Conv layer (convolution layer) and a downsampling layer are added before the two Conv layers in the original yolo10n model, so that the size of the feature map becomes 640X640 again after each subsequent convolution and downsampling, which is the same as the size of the original input image, and the clarity will be higher, which is more conducive to feature extraction. It should be noted that the modification of the real-time input image can be used as the input value during the training process of the improved yolo10n model, or as the input value during the final recognition process.
[0048] S30: Use DualConv to process the same feature image channels, divide the input channels, and concatenate the features extracted from all branches to obtain the finally output feature image;
[0049] In the preferred embodiment of the present invention, since the Conv layer is used for feature extraction in the original neural network model, its expression ability is low, and the calculation efficiency is also low. In order to improve the expression ability of the neural network model and improve the calculation efficiency. As Figure 2 shown, use DualConv (dual convolutional kernel) to replace the Conv in the original neural network model. DualConv combines 3X3 and 1X1 convolutional kernels to process the same input feature map channels simultaneously, and uses the grouped convolution technique to effectively arrange the convolutional filters. Among them, grouped convolution can realize dividing the input channels to achieve multi-branch feature learning, and then concatenate the features extracted from all branches to obtain the finally output feature image. As Figure 3 shown, in the DualConv structure, 3X3 and 1X1 convolutional kernels are combined and fused. The 3X3 convolutional kernel can capture more spatial information when extracting features, while the 1X1 convolutional kernel can perform the interaction and information integration between feature channels without increasing too many parameters and computational complexity. By using the DualConv method, the number of parameters can be greatly reduced, thereby improving the operation speed.
[0050] S40: Obtain the prediction boxes in the feature image through Uiou;
[0051] S50: Evaluate the geometric overlap degree between the prediction box and the ground truth box to train the neural network model;
[0052] In the preferred embodiment of the present invention, Unified-IoU (i.e., Uiou) is used as the loss of the target regression box. In the traditional model, IoU is used as a common metric in object detection to evaluate the geometric overlap degree between the prediction box and the ground truth box, and its calculation formula is:
[0053]
[0054] However, this calculation method has certain limitations. When the predicted box and the ground truth box partially overlap but do not completely coincide, the IoU may not be sufficient to distinguish the quality of the predicted box; or it is insensitive to the position and shape of the bounding box. IoU only measures the area overlap and does not consider whether the shape and position of the bounding box match; at the same time, there is also a problem of non-differentiable gradient. During the training process, the discrete nature of IoU will lead to unstable gradient updates.
[0055] Using Uiou can improve the above limitations. When obtaining the predicted box in the image from the extracted feature image, some new features are added in object detection, making the prediction of the object detection model more accurate. It mainly includes introducing dynamic weights in Uiou, which allows the model to focus on predicted boxes of different qualities during the training process. For example, in the early stage of training, the model gives priority to low-quality predicted boxes, thus reducing obvious incorrect predictions. In the later stage of training, it gradually focuses on the optimization of high-quality predicted boxes, further improving the accuracy. For example, if the overlap degree between the predicted box and the ground truth box is only 10%, at this time, because the predicted box contains less content of the ground truth box, it is prone to obvious errors, so it is named a low-quality predicted box; similarly, if the overlap degree between the predicted box and the ground truth box is 90%, at this time, because the predicted box contains more content of the ground truth box and already has a certain degree of accuracy, it is named a high-quality predicted box. The specific overlap degree can be set according to needs or dynamically adjusted according to the size of the ground truth box.
[0056] The original IoU loss formula is:
[0057]
[0058] where P is the predicted box, and P gt is the ground truth box.
[0059] Unified-loU adds a penalty term for the width and height error on the original basis, specifically:
[0060] (w t ―w p ) 2 / (w t 2 )
[0061] (h t ―h p ) 2 / (h 2 )
[0062] where w t and h t are the actual width and height, and w p and h p are the predicted width and height.
[0063] Then the Unified-loU loss is:
[0064]
[0065] where α and β are hyperparameters that can be adjusted during training.
[0066] UIoU not only evaluates geometric overlap but also incorporates characteristics such as the shape, position, and size of the predicted bounding boxes. With these additional factors, UIoU provides a more fine-grained evaluation criterion while training the neural network model during the prediction evaluation process. After multiple trainings, a trained neural network model can be obtained. Moreover, since the model's attention shifts from low-quality predicted bounding boxes to high-quality ones, it can enhance the model's detection performance on high-precision or dense datasets and achieve a balance in training speed.
[0067] S60: Use the trained neural network module to identify the acquired real-time image information, determine whether there is an abnormality, and mark the abnormal position;
[0068] S70: Generate an image analysis report based on the analysis results;
[0069] In the preferred embodiment of the present invention, the real-time image obtained by the bronchoscope is used as the input image of the trained neural network model to identify the image to determine whether there is an abnormality. If there is an abnormality, the abnormal position is marked. The acquired real-time image can also be preprocessed and / or converted into a feature map of size 640X640 and then input into the trained neural network model for identification.
[0070] Generating the image analysis report is mainly divided into two parts, one is the manual information writing module, and the other is the intelligent information production module. The intelligent information production module is mainly used to display and summarize the identified data to form a specific identification report, including the identification result, the credibility of the identification result, and the identification suggestions. The manual information writing module allows the operator to write the information corresponding to the image to be detected, including image information, generation time, operator, detection time, etc. At the same time, the information in the identification report generated by the intelligent information module can be further modified, and pictures or cases with different identification results for the operator and the image to be detected are stored for further analysis.
[0071] Through the above method, the expressive power of the improved neural network model is improved, the computational efficiency is faster, and the image and weak feature marking can be efficiently annotated. It is suitable for real-time images of respiratory tract obtained by bronchoscopy, the recognition ability of small targets and weak lesion features in complex bronchoscopic examinations, and the huge number of images generated by bronchoscopic detection. The improved neural network can be trained to improve the adaptability and detection efficiency of the model, so as to realize large-scale extraction of key features in bronchoscopic detection images. The input real-time image is used as the input value of the improved neural network model, and the size of the input value is modified to process the same input feature channel using dual-core convolution to realize multi-branch feature learning, and the features extracted by all branches are cascaded to obtain the final output feature image, which not only captures more spatial information during feature extraction, but also interacts and integrates information between feature channels without increasing the computational complexity of too many parameters, greatly reducing the number of parameters, thereby improving the operation speed. Then, the model is trained using prediction frames of different qualities. Low-quality detection frames are used to reduce obviously wrong predictions, and high-quality prediction frames are used to improve accuracy. The characteristics of other prediction frames can also be combined to provide a more fine-grained evaluation standard to obtain a trained neural network model. The improved neural network model is used to identify real-time images and determine whether there are abnormalities. If there are abnormalities, the abnormal location is marked, thereby providing high-level results for intelligent recognition of bronchoendoscopic images and improving efficiency.
[0072] Corresponding to the above-mentioned improved neural network model and image recognition method for image recognition, the present invention also provides an image recognition system, specifically, as Figure 4 As shown, the image recognition system includes: an image acquisition module 100, an image recognition module 200 and an image reporting module 300.
[0073] Image acquisition module 100, acquires real-time image information and performs preprocessing;
[0074] The camera system is used to obtain real-time image information, such as using a hysteroscope to collect real-time image information of the patient's uterine cavity. It can also be used to obtain image information using a camera in other scenarios. This application will take the hysteroscope to obtain real-time image information of the uterine cavity as an example for detailed description. The acquired real-time image is preprocessed, including photoelectric conversion, and the converted data is further decoded into the endoscope handle to form the original image format. During the conversion process, the acquired image information is subjected to noise filtering and cropping, so that the size of the obtained original image is 1280X720.
[0075] The image recognition module 200 adopts the same input feature map channels for dual-core convolution, divides the input channels to achieve multi-branch learning, cascades the features extracted from all branches to obtain the finally output feature image, thus obtaining an improved neural network model. The improved neural network model is trained and learned using prediction boxes of different qualities to obtain the trained neural network model, and the trained and learned model is used for image recognition;
[0076] Specifically, since the original neural network model uses the Conv layer for feature extraction, its expression ability is low, and at the same time, the calculation efficiency is also low. In order to improve the expression ability of the neural network model and improve the calculation efficiency. As Figure 2 shown, DualConv (dual convolution kernel) is used to replace the Conv in the original neural network model. DualConv combines 3X3 and 1X1 convolution kernels to process the same input feature map channels simultaneously, and uses the grouped convolution technique to effectively arrange the convolution filters. Among them, grouped convolution can achieve dividing the input channels to realize multi-branch feature learning, and then cascading the features extracted from all branches to obtain the finally output feature image. In the DualConv structure, 3X3 and 1X1 convolution kernels are combined and fused. The 3X3 convolution kernel can capture more spatial information when extracting features, while the 1X1 convolution kernel can perform interaction and information integration between feature channels without increasing too many parameters and computational complexity. By using the DualConv method, the number of parameters can be greatly reduced, thereby improving the operation speed.
[0077] Furthermore, Unified-IoU (i.e., Uiou) is used as the loss of the target regression box. In traditional models, IoU is used as a common metric in object detection to evaluate the geometric overlap between the prediction box and the ground truth box, and its calculation formula is:
[0078]
[0079] However, this calculation method has certain limitations. When the prediction box and the ground truth box partially overlap but do not completely coincide, IoU may not be sufficient to distinguish the quality of the prediction box; or it is not sensitive to the position and shape of the bounding box. IoU only measures the area overlap and does not consider whether the shape and position of the bounding box match; at the same time, there is also a problem of non-differentiable gradient. During the training process, the discrete nature of IoU will lead to unstable gradient updates.
[0080] However, adopting Uiou can improve the above limitations. When obtaining the prediction boxes in the image from the extracted feature images, some new features are added in object detection, making the prediction of the object detection model more accurate. This mainly includes introducing dynamic weights in Uiou, allowing the model to focus on prediction boxes of different qualities during the training process. For example, in the early stage of training, the model gives priority to low-quality prediction boxes, reducing significantly incorrect predictions. In the later stage of training, it gradually focuses on optimizing high-quality prediction boxes, further improving the accuracy. For instance, if the overlap between the prediction box and the ground truth box is only 10%, since the prediction box contains less content of the ground truth box, it is prone to obvious errors and is thus named a low-quality prediction box. Similarly, if the overlap between the prediction box and the ground truth box is 90%, since the prediction box contains more content of the ground truth box and already has a certain level of accuracy, it is named a high-quality prediction box. The specific overlap can be set as needed or dynamically adjusted according to the size of the ground truth box.
[0081] The original IoU loss formula is:
[0082]
[0083] where P is the prediction box, and P gt is the ground truth box.
[0084] Unified-loU adds a penalty term for width and height errors on the original basis, specifically:
[0085] (w t ― w p ) 2 / (w t 2 )
[0086] (h t ― h p ) 2 / (h 2 )
[0087] where w t and h t are the actual width and height, and w p and h p are the predicted width and height.
[0088] Then the Unified-loU loss is:
[0089]
[0090] where α and β are hyperparameters that can be adjusted during the training process.
[0091] UIoU not only evaluates geometric overlap but also incorporates characteristics such as the shape, position, and size of the predicted bounding boxes. With these additional factors, UIoU provides a more fine-grained evaluation criterion and trains the neural network model during the prediction evaluation process. After multiple trainings, a trained neural network model can be obtained. Moreover, since the attention of the model shifts from low-quality predicted bounding boxes to high-quality ones, it can enhance the detection performance of the model on high-precision or dense datasets and achieve a balance in training speed.
[0092] The image reporting module 300 reports the recognized image information.
[0093] Specifically, the image reporting module mainly consists of two parts, one is the manual information writing module, and the other is the intelligent information production module. The intelligent information production module is mainly used to display and summarize the recognized data to form a specific recognition report, including the recognition result, the credibility of the recognition result, and the recognition suggestions. The manual information writing module allows the operator to write the information corresponding to the image to be detected, including image information, generation time, operator, detection time, etc. At the same time, it can further modify the information in the recognition report generated by the intelligent information module. Also, pictures or cases with different recognition results for the operator and the image to be detected are stored for further analysis.
[0094] Based on the above embodiments, the present invention also provides an intelligent terminal, and its principle block diagram is as Figure 5 shown. The above intelligent terminal includes a processor, a memory, a network interface, and a display screen connected through a system bus. The processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an image recognition program. The memory provides an environment for the operation of the operating system and the image recognition program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. When the image recognition program is executed by the processor, it implements the steps of any of the above image recognition methods. The display screen of the intelligent terminal can be a liquid crystal display screen or other display screens.
[0095] Those skilled in the art can understand that Figure 5 the principle block diagram shown is only a block diagram of some structures related to the solution of the present invention and does not constitute a limitation on the intelligent terminal to which the solution of the present invention is applied. The specific intelligent terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0096] In an embodiment of the present invention, an intelligent terminal is provided. The intelligent terminal includes a memory, a processor, and an image recognition program stored on the memory and executable on the processor. When the image recognition program is executed by the processor, the following operation instructions are performed:
[0097] Obtain real-time image information;
[0098] Perform noise filtering and cropping on the obtained real-time image information to form an original image format;
[0099] Fill the original image format and then obtain a feature map through a convolutional layer and a downsampling layer;
[0100] Use dual convolutional kernels to simultaneously process the same input feature map channels;
[0101] Divide the input channels to obtain multi-branch feature learning;
[0102] Cascade the features extracted from all branches to obtain the final output to obtain an improved neural network model;
[0103] Train the improved neural network model using low-quality prediction boxes;
[0104] Train the improved neural network model using high-quality prediction boxes to obtain a trained and learned neural network model;
[0105] Input the obtained real-time image information into the trained and learned neural network model for image recognition and mark the differences;
[0106] Generate a corresponding recognition result report for the recognition result.
[0107] It should be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. An improved neural network model for image recognition, characterized in that, Including: Obtain a real-time image of the first size; Use the real-time image of the first size as the input feature map of the neural network model; Simultaneously process the same input feature map channel using the first convolutional kernel and the second convolutional kernel; Divide the input feature map channels to achieve multi-branch feature learning; Concatenate the features extracted from all branches to obtain the finally output feature map; Collect multiple finally output feature maps to form an improved neural network model.
2. The improved neural network model according to claim 1, characterized in that: During the process of simultaneously processing the same input feature map channel using the first convolutional kernel and the second convolutional kernel, it further includes: Use group convolution to arrange the convolutional filters.
3. The improved neural network model according to claim 2, characterized in that: During the process of simultaneously processing the same input feature map channel using the first convolutional kernel and the second convolutional kernel, it further includes: Capture more spatial information of the features through the first convolutional kernel, and perform interaction and information integration between the feature channels through the second convolutional kernel.
4. The improved neural network model according to claim 1, wherein: After forming the improved neural network model by collecting multiple finally output feature maps, it further includes: Use the first prediction box and the second prediction box to train and learn the model. The first prediction box is applied in the early stage of training, and the second prediction box is applied in the later stage of training.
5. The improved neural network model according to claim 4, wherein: The first prediction box and the second prediction box are the geometric coincidence degrees between the prediction box and the ground truth box. The geometric coincidence degree lower than the set value is the first prediction box, and the set coincidence degree higher than the set value is the second prediction box.
6. The improved neural network model according to claim 5, wherein: The set value is adjusted according to the ground truth box during the training process. The prediction box includes at least one characteristic of shape, position, or size.
7. An image recognition method, characterized in that, Including: Obtain a real-time image of the first size; Use the real-time image of the first size as the input value for the improved neural network model for image recognition as described in any one of claims 1-6, and judge and mark the recognition information.
8. The image recognition method according to claim 7, characterized in that: During the process of extracting features from the real-time image of the second size to obtain the feature map in the real-time image of the second size, it further includes: Perform convolution and downsampling on the real-time image of the second size to obtain the feature map in the real-time image of the second size. The size of the feature map is the first size; During the process of obtaining the real-time image of the first size, it further includes: Filter the noise and crop the real-time image obtained by the endoscope, and perform photoelectric conversion and decoding to obtain the real-time image of the first size.
9. An image recognition system, characterized in that, Including: An image acquisition module that acquires real-time image information and performs preprocessing; An image recognition module that uses the same input feature map channel for dual-core convolution, divides the input channels to achieve multi-branch learning, concatenates the features extracted from all branches to obtain the finally output feature image, thereby obtaining an improved neural network model, uses prediction boxes of different qualities to train and learn the improved neural network model to obtain the trained neural network model, and uses the trained and learned model for image recognition; An image reporting module that reports the recognized image information.
10. A storage medium, on which program instructions are stored, characterized in that: The program instructions are used to execute an image recognition method as described in any one of claims 7 to 8 when running.