Image recognition and detection method, image information detection method and image processing system

WO2026199807A1PCT designated stage Publication Date: 2026-10-01ZHUHAI LIVZON DIAGNOSTICS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/115730
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2025-08-19
Publication Date
2026-10-01

Smart Images

  • Figure CN2025115730_01102026_PF_FP_ABST
    Figure CN2025115730_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of medical instruments. Provided are an image recognition and detection method, an image information detection method and an image processing system. The image recognition and detection method comprises: inputting an image to be tested into a feature extraction unit; acquiring a feature map extracted by the feature extraction unit on the basis of said image; on the basis of the feature map, respectively enabling a first neural network to execute a first prediction process and enabling a second neural network to execute a second prediction process; and acquiring a category classification result about a target object in said image on the basis of a prediction result of the first prediction process, and acquiring quantization information about the target object on the basis of a prediction result of the second prediction process. The present invention can predict karyotypes on the basis of the microscopic morphology of cells in images, and takes into account both the overall presentation of the field of view and microscopic features of individual cells so as to identify titer information, thereby improving the accuracy of karyotype and titer identification and reducing occurrences of incorrect or failed predictions.
Need to check novelty before this filing date? Find Prior Art

Description

Image recognition and detection methods, image information detection methods, and image processing systems Technical Field

[0001] This invention relates to the field of medical device technology, and in particular to an image recognition and detection method, an image information detection method, and an image processing system. Background Technology

[0002] Pattern recognition of indirect immunofluorescence (IIF) images is a key technology for the diagnosis of autoimmune diseases, and it is commonly used for image recognition of antibodies such as antinuclear antibodies (ANA) and anti-double-stranded DNA antibodies (anti-dsDNA antibodies). While deep learning has made some progress in this field and can interpret fluorescence images of antinuclear antibodies (ANA) using indirect immunofluorescence, existing methods are typically limited to single karyotype recognition and are not accurate enough in identifying karyotype features, often resulting in misprediction or failure to predict. Furthermore, they lack joint modeling of titer information; karyotype and titer information are judged using separate models, failing to reflect the correlation between titer and karyotype. Summary of the Invention

[0003] To overcome the shortcomings of the prior art, the present invention aims to provide an image recognition and detection method, an image information detection method, and an image processing system that can predict karyotypes based on the microscopic morphology of cells in an image, while taking into account both the overall performance of the field of view and the microscopic feature identification titer information of individual cells. This avoids being limited to the identification and prediction of a single karyotype and can improve the accuracy of karyotype and titer identification, reducing prediction errors or unpredictable phenomena.

[0004] In a first aspect, the present invention provides an image recognition prediction method applied to an image processing model, the image processing model including a feature extraction unit, a first neural network, and a second neural network. The method includes: inputting an image to be detected into the feature extraction unit; obtaining a feature map extracted by the feature extraction unit based on the image to be detected; according to the feature map, causing the first neural network to execute a first prediction process and causing the second neural network to execute a second prediction process; obtaining a category classification result of a target object in the image to be detected based on the prediction result of the first prediction process, and obtaining quantitative information about the target object based on the prediction result of the second prediction process; wherein the first prediction process includes at least the following steps performed by the first neural network: generating a dynamic attention map related to the target object based on the feature map; forming a classification feature vector corresponding to the category based on the dynamic attention map and the feature map; and obtaining a category classification result of the target object from the classification feature vector.

[0005] The image recognition prediction method provided in the first aspect of this invention is applied in a deep learning-based image processing model. This model uses an extraction unit to detect and identify fluorescent images containing a target object (antinuclear antibody), and then extracts information that can be simultaneously applied to a first neural network for identifying the target object's category (e.g., karyotype) and a second neural network for predicting the target object's quantitative information (e.g., titer information) in the image. Therefore, the feature maps used by the two neural networks are consistent, and both can extract their respective predictive characteristic information from these feature maps without needing to be separately input into the two neural networks for individual processing. This allows the category classification prediction and quantitative information prediction processes to proceed synchronously. The combination of the two prediction results is beneficial for analyzing the disease status of the tested individual and greatly improves the overall prediction efficiency. Furthermore, this invention introduces a dynamic attention map in the first neural network, which can effectively extract features related to each category from the feature map. Compared to existing deep learning networks in this field, it has better micro-feature recognition results and significantly improves the accuracy of image recognition prediction.

[0006] In a preferred embodiment of the present invention, the step of forming a classification feature vector corresponding to the category based on the dynamic attention map and the feature map includes:

[0007] The second feature vector of the dynamic attention map is multiplied by the first feature vector of the feature map to obtain the feature matrix, and the classification feature vector is obtained by adding the features within the feature matrix.

[0008] In a preferred embodiment of the present invention, obtaining the category classification result of the target object from the classification feature vector includes: performing average pooling on the feature map to obtain a global feature vector, and obtaining the category classification result based on the global feature vector and the classification feature vector.

[0009] In a preferred embodiment of the present invention, the feature map includes a multi-scale feature map generated using an image processing network, and the first neural network forms the dynamic attention map based on the multi-scale feature map.

[0010] In a preferred embodiment of the present invention, the generation of the multi-scale feature map includes the following steps: obtaining an original feature map based on the image to be detected; dividing the original feature map into multiple sub-feature maps according to the number of channels; performing corresponding multi-scale processing on the multiple sub-feature maps and outputting them as multiple output subsets; concatenating all the output subsets and then performing channel adjustment and feature fusion through convolution to obtain the multi-scale feature map; wherein, the multi-scale processing includes: for the first sub-feature map, directly outputting it as an output subset; for the second sub-feature map, convolving it and outputting it as an output subset; for the i-th sub-feature map, adding the i-th sub-feature map to the convolved (i-1)-th sub-feature map to obtain a fused feature map, then convolving the fused feature map and outputting it as an output subset, i≥3.

[0011] In a preferred embodiment of the present invention, the second prediction process includes at least the following steps performed by the second neural network: imbuing the feature map with a dynamic mechanism by a dynamic network; generating a quantized feature vector for the quantized information of the target object based on the feature map with the dynamic mechanism; and obtaining the quantized information of the image to be detected about the target object from the quantized feature vector.

[0012] In a preferred embodiment of the present invention, the dynamic mechanism of the feature map includes: introducing dynamic weights into the feature map and forming a dynamic nonlinear layer, wherein each feature in the dynamic nonlinear layer is formed by combining each feature in the feature map with dynamic weights, and the dynamic weights are dynamically set to different weight values ​​according to the different feature values ​​of each feature in the feature map.

[0013] In a preferred embodiment of the present invention, the generation of the dynamic attention map from the feature map includes at least the following steps: converting the feature map into a transformed feature map with the same number of channels as the number of categories of the target object through a specific dynamic convolutional layer; and performing classification regression on the transformed feature map to obtain the dynamic attention map.

[0014] In a second aspect, the present invention provides an image information detection method, the method comprising: acquiring an image to be detected; executing an image recognition prediction method as described in the first aspect embodiment to acquire a prediction result about the image to be detected based on the image processing model, the prediction result including category classification prediction information and quantization prediction information about a target object in the image to be detected; acquiring a detection result based on the image to be detected input from an external source, the detection result including category classification detection information and quantization detection information about a target object in the image to be detected; and verifying the prediction result and the detection result to obtain a final detection result of the target object in the image to be detected.

[0015] The image information detection method provided by the second aspect of the present invention combines the image recognition prediction method of the first aspect embodiment, and on this basis, obtains the detection result of external input, and compares the external detection result with the prediction result obtained by the image recognition prediction method. This avoids the prediction error that may occur when the prediction result is used as the final detection result. When the two results are consistent, the accuracy of the result can be effectively guaranteed. At the same time, the prediction result can also be used as a reference for the detection result of external input, which can solve the problem of verification error that is easy to occur in the current manual detection process.

[0016] In a preferred embodiment of the present invention, acquiring the image to be detected includes: inputting the image to be detected into a system applying the image information detection method via an external input channel; displaying the image to be detected in a first area of ​​the display interface; the image information detection method further includes: displaying the prediction result and the detection result in a second area and a third area of ​​the display interface, respectively.

[0017] In a preferred embodiment of the present invention, the verification of the prediction result and the detection result includes a first-level verification process, which includes: inputting a first-level verification result of the category classification and a first-level verification result of the quantitative information of the target object in the image to be detected, based on the image to be detected displayed in the first area of ​​the display interface, the prediction result displayed in the second area, and the detection result displayed in the third area; the verification of the prediction result and the detection result also includes a second-level verification process, which includes at least one consecutive verification node, each verification node including: inputting a first-level verification result of the category classification and a first-level verification result of the quantitative information of the target object in the image to be detected, based on the image to be detected displayed in the first area of ​​the display interface, the prediction result displayed in the second area, and the detection result displayed in the third area; The detection results displayed in the third area, along with the verification information from the previous verification node, are used to input the category classification verification results and quantification information verification results for the target object in the image to be detected. Specifically, when the current verification node is the first verification node, the verification information from the previous verification node is replaced with the category classification first-level verification results and quantification information first-level verification results from the first-level verification process. If the current verification node is not the last verification node, the category classification verification results and quantification information verification results are input to the next verification node. If the current verification node is the last verification node, the category classification verification results and quantification information verification results input by the current verification node are the final detection results for the target object in the image to be detected.

[0018] In a preferred embodiment of the present invention, the verification of the prediction result and the detection result further includes a quality detection process, which includes: when the final detection result and the prediction result are consistent, positively stimulating the image processing model based on the prediction result; when the final detection result and the prediction result are inconsistent, recording the number of inconsistencies for the prediction result; and when the number of inconsistencies reaches a certain value, recording the current abnormal situation.

[0019] Thirdly, the present invention further proposes a training method for an image processing model according to the first aspect embodiment, the method comprising at least one training cycle, each training cycle comprising:

[0020] The training image set is input into the image processing model. The training image set contains multiple first images. Each first image has category information corresponding to the target object contained in the first image and quantification information about the target object in the first image.

[0021] The image processing model extracts feature maps based on the first image;

[0022] The feature maps are respectively input into the first neural network and the second neural network of the image processing model. The first neural network obtains the category classification result of the target object based on the feature maps, and the second neural network obtains the quantitative information of the target object based on the feature maps.

[0023] When acquiring the feature map, the first neural network performs at least the following steps:

[0024] Generate a dynamic attention map related to the category information based on the feature map;

[0025] Based on the dynamic attention map and the feature map, a classification feature vector corresponding to the category is formed;

[0026] The classification result of the target object is obtained from the classification feature vector;

[0027] When acquiring the feature map, the second neural network performs at least the following steps:

[0028] The feature map is given a dynamic mechanism by a dynamic network;

[0029] A quantized feature vector for the target object is generated based on the feature map with a dynamic mechanism.

[0030] The quantization information of the first image corresponding to each feature map is obtained by combining the quantization feature vectors of each feature map.

[0031] According to the training method provided in the third aspect of the present invention, the image processing model in the first aspect embodiment is trained in a training cycle manner. In each training cycle, the feature map is trained to extract the extraction strategy of each classification category in the image to be detected, so that the attention is focused on the relevant micro-features of each kernel type, and a dynamic attention map is generated on the first neural network to further focus on the dynamic attention phenomenon in the feature map. This can enable the kernel type prediction to achieve high accuracy and identify all kernel types, including rare kernel types that are difficult to identify manually. A dynamic mechanism is introduced into the second neural network so that the dynamic mechanism dynamically obtains the quantity information of the target object in the entire field of view of the feature map. The quantitative information such as titer information obtained by training has higher accuracy and can accurately know the detection status of the detected object.

[0032] In a preferred embodiment of the present invention, the step of forming a classification feature vector corresponding to the category based on the dynamic attention map and the feature map includes:

[0033] The second feature vector of the dynamic attention map is multiplied by the first feature vector of the feature map to obtain the feature matrix, and the classification feature vector is obtained by summing the features within the feature matrix.

[0034] The step of obtaining the category classification result of the target object from the classification feature vector includes:

[0035] The feature map is subjected to average pooling to obtain a global feature vector, and the category classification result is obtained based on the global feature vector and the classification feature vector.

[0036] In a preferred embodiment of the present invention, the method further includes a verification stage, the verification stage comprising:

[0037] The test image set is input into the image processing model trained by the training image set. The test image set contains multiple second images, each of which has category information corresponding to the target object contained in the second image and quantization information about the target object in the second image.

[0038] Obtain the verification results output by the image processing model based on the test image set;

[0039] The performance of the image processing model is evaluated based on the verification results.

[0040] In a preferred embodiment of the present invention, the category information includes common category information and rare category information, and the training image set, before being input into the image processing model, further includes:

[0041] The weight of the first image corresponding to the rare category information in the training image set is adjusted so that the probability of the image processing model selecting the first image containing the rare category information in the training image set is higher than the probability of the first image containing the rare category information in the training image set before adjustment.

[0042] In a preferred embodiment of the present invention, the method further includes:

[0043] Obtain the first loss function value of the first neural network on each of the feature maps with respect to the category classification result, and obtain the second loss function value of the second neural network on each of the feature maps with respect to the quantized information;

[0044] Obtain the adaptive loss function value based on the first loss function value and the second loss function value;

[0045] The first neural network and the second neural network optimize the performance of the image processing model based on the adaptive loss function value.

[0046] In a preferred embodiment of the present invention, obtaining the adaptive loss function value based on the first loss function value and the second loss function value includes:

[0047] Calculate the standardized weights of the first loss function value and the second loss function value, and then adaptively adjust the standardized weights to obtain the adaptive function value;

[0048] The calculation of the standardized weights for the first loss function value and the second loss function value includes:

[0049] The standardized weights are calculated based on the relative rate of change of the first loss function value and the second loss function value at adjacent time steps;

[0050] The adaptive adjustment method for the standardized weights includes at least one of the following: (1) weighting the standardized weights; (2) dynamically adjusting the standardized weights according to the magnitude of the first loss function value and the second loss function value; (3) dynamically adjusting the standardized weights according to the error type of the first loss function value and the second loss function value.

[0051] Fourthly, the present invention also proposes an image processing system applying an image processing model, wherein the image processing model includes a feature extraction unit, a first neural network, and a second neural network, and the image processing system includes:

[0052] The input module is used to acquire the input image to be detected;

[0053] The prediction module inputs the feature map of the image to be predicted into the image processing model and obtains the prediction result of the image processing model.

[0054] The image processing model, when processing the image to be detected, includes:

[0055] The feature extraction unit extracts a feature map based on the image to be detected;

[0056] The first neural network performs a first prediction process based on the feature map, and the second neural network performs a second prediction process based on the feature map;

[0057] The first prediction process includes:

[0058] Generate a dynamic attention map of the target object in the image to be detected based on the feature map;

[0059] A classification feature vector for the target object is formed based on the dynamic attention map and the feature map;

[0060] The classification result of the target object is obtained from the classification feature vector;

[0061] The second prediction process includes:

[0062] The feature map is given a dynamic mechanism by a dynamic network;

[0063] A quantized feature vector for the target object is generated based on the feature map with a dynamic mechanism.

[0064] Quantization information about the target object is obtained from the quantization feature vector.

[0065] Fifthly, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, performs the image recognition prediction method as described in the first aspect embodiment, the image information detection method as described in the second aspect embodiment, or the image processing model training method as described in the third aspect embodiment.

[0066] Other features and advantages of the invention will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures and / or processes particularly pointed out in the description, claims, and drawings. Attached Figure Description

[0067] Figure 1 is a flowchart of the image recognition prediction method provided in an embodiment of the present invention;

[0068] Figure 2 is a schematic diagram of the image processing model provided in an embodiment of the present invention;

[0069] Figure 3 is a schematic diagram of the feature extraction module of the image processing model provided in an embodiment of the present invention;

[0070] Figure 4 is a schematic diagram of the structure of the first neural network of the image processing model provided in an embodiment of the present invention;

[0071] Figure 5 is a schematic diagram of the structure of the second neural network of the image processing model provided in an embodiment of the present invention;

[0072] Figure 6 is a schematic diagram of the display interface provided in an embodiment of the present invention;

[0073] Figure 7 is a schematic diagram of the verification system used in the image information detection method provided in an embodiment of the present invention;

[0074] Figure 8 is a schematic diagram of the image processing system provided in an embodiment of the present invention;

[0075] Figure 9 is a performance evaluation diagram of the image processing model provided in the embodiment of the present invention;

[0076] Figure 10 shows the category classification prediction results of the image processing model provided in the embodiment of the present invention;

[0077] Figure 11 shows the quantization information prediction results of the image processing model provided in the embodiment of the present invention.

[0078] Figure 12 shows the performance evaluation of the image processing model provided in this embodiment of the invention under different backbones. Detailed Implementation

[0079] The following detailed description of the embodiments of the present invention, in conjunction with the accompanying drawings, will provide a thorough understanding of how the present invention uses technical means to solve technical problems and achieve technical effects, enabling its implementation. It should be noted that these specific descriptions are merely intended to facilitate a clearer understanding of the present invention by those skilled in the art, and are not intended to limit the scope of the invention. For example, the terms "first" and "second" mentioned in the embodiments of the present invention are not intended to limit the invention, but are merely used to indicate the sequence numbers of multiple identical or similar devices or mechanisms. Those skilled in the art can readjust these sequence numbers for ease of description or during the organization of technical solutions. Furthermore, alternative solutions are described for some mechanisms in different embodiments, and these alternatives can be applied to other identical or similar devices or mechanisms. As long as there is no conflict, the various embodiments and features in each embodiment of the present invention can be combined with each other, and the resulting technical solutions are all within the protection scope of the present invention.

[0080] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0081] Firstly, to improve karyotype recognition and titer prediction accuracy, this invention proposes an image recognition prediction method. This method is applied in an image processing model, which includes three modules (see Figure 2): a feature extraction unit, a first neural network, and a second neural network. The feature extraction unit extracts features from the image to be detected. The first neural network forms a category classification result based on the extracted features. The second neural network predicts the quantitative information of the image to be detected based on the extracted features. Typically, the image to be detected in this embodiment can be represented as a fluorescent image obtained by indirect immunofluorescence (IIF), such as Figure 1. The fluorescent image can display information about the cell nucleus, cytoplasm, cell membrane, and certain components produced during the cell division cycle. The target object (antinuclear antibody, ANA) can show corresponding microscopic and quantitative information within the fluorescent image. The corresponding ANA karyotype can be identified through this microscopic information, and the titer information of the target object can be obtained based on the entire field of view. Combining the karyotype and titer information, the detection result of the subject can be obtained.

[0082] The image processing model used in this embodiment will be described in detail below.

[0083] Referring to Figure 2, the feature extraction unit is used to extract feature information from the image to be detected and form a feature map. Specifically, the feature map extracts microscopic feature information about the target object ANA corresponding to each category of karyotype in the image to be detected. This microscopic feature information does not refer to single feature information such as specific karyotype outline, cell number, or cell distribution characteristics, but rather the feature information corresponding to each category of karyotype. This feature information may be formed from a portion of the features in the example above. Thus, the generated feature map contains the corresponding feature information of each category of karyotype, enabling the first neural network to focus on calculating the information of each category of karyotype within the feature map when performing feature recognition, thereby obtaining the relevant category classification results. The second neural network obtains relevant quantitative information from the same feature map as the first neural network, improving prediction efficiency and fully utilizing the feature information extracted from the feature map, making the karyotype titer result accurate.

[0084] Specifically, in this embodiment, the feature extraction unit serves as the backbone of the image processing model. The feature map includes a multi-scale feature map generated using an image processing network, which is either part of the backbone or a component of the backbone. The information in the multi-scale feature map contains relevant information about the image to be detected under multiple information dimensions. For example, the generated multi-scale feature map contains the original feature information of the image to be detected, the feature information after one feature extraction, and the feature information after multiple sequential feature extractions. Thus, the information used by the subsequent two neural networks to predict the feature map includes both the original image features and the relevant features after extraction and multiple extractions. As a result, the subsequent two neural networks can obtain more information during the prediction process, and can identify more accurately when the relevant features under the corresponding scale feature meet the category classification requirements. For example, a certain kernel type may exhibit significant features in a one-dimensional feature map, thus the neural network will have higher accuracy in recognizing that one-dimensional feature. Conversely, another kernel type may exhibit significant features in a three-dimensional feature map but not in a one-dimensional feature map, thus the neural network will have higher accuracy in recognizing that three-dimensional feature. Multi-scale feature maps contain feature information at multiple scales, allowing the neural network to more accurately identify the corresponding kernel type within the corresponding scale feature based on different scale information. This ensures that each kernel type can be identified and classified based on sufficient feature information. Here, the aforementioned dimension refers to the dimension represented by the feature map. The information combination consists of relevant scale feature information obtained from the number of processing iterations. For example, the multi-scale feature map includes original image feature information, which is described as one-dimensional feature information without convolution or other information processing; it also includes extracted feature information after one convolution or other information processing of the original image feature information, described as two-dimensional feature information; and it includes extracted feature information after two convolution or other information processing of the extracted information and / or the original feature information, described as three-dimensional feature information. Similarly, after each information processing, multi-dimensional feature information at the corresponding scale can be extracted, and all multi-dimensional feature information is combined to form a multi-scale feature map containing multi-scale information. Likewise, quantification information, such as titer, may be more pronounced at different scales at different quantities, which can also effectively improve the prediction accuracy of quantification information.

[0085] Referring to Figure 3, the backbone employs an extraction module comprising grouping, an input layer, an extraction layer, and an output layer. The grouping layer divides the input features of the original image into multiple groups. Specifically, it processes the input features of the original image using a 1×1 convolution to form an input feature matrix. Subsequently, the input feature matrix is ​​normalized (e.g., using the self.bn1() function) to accelerate training convergence and improve model stability. Then, the processed input feature matrix is ​​input into an activation function (e.g., the ReLU function) to introduce nonlinearity, increase the sparsity of the network, reduce parameter dependence, and alleviate gradient vanishing. The processed input feature matrix is ​​then divided into s groups, denoted as ai, i∈1,2,3,...,s. If s is the number of channels in the input feature map of the original image, then the grouping is a sub-feature map corresponding to a single channel number. Each sub-feature map undergoes corresponding multi-scale processing and outputs multiple output subsets, which are bi, i∈1,2,3,...,s in Figure 3. These multiple output subsets form the output layer. When grouping, the number of sub-feature maps corresponds to the number of 1×1 convolution kernels on the channels. When the number of convolution kernels is s, the number of channels of the sub-feature map is 1 / s. Then, the input feature map is reduced to 1 dimension and has s corresponding sub-feature maps.

[0086] In this embodiment, the extraction layer performs multi-scale processing on multiple obtained sub-feature maps. During the extraction process, for the first sub-feature map, the first output subset, b1, is directly output. This output subset is the output subset under the corresponding one-dimensional scale feature. For the second sub-feature map, it is convolved and output as the second output subset b2. For the i-th (i>2) sub-feature map, it is added to the (i-1)-th sub-feature map after convolution to obtain a fused feature map. This fused feature map includes feature information from the previous several scales. Then, the fused feature map is convolved and output as the i-th output subset bi. For example, when i is 4, after iterative fusion, it encompasses one-dimensional, two-dimensional, and three-dimensional scale features. Then, the fused feature map obtained above is convolved again to contain four-dimensional scale features and output as the fourth output subset b4. Therefore, this multi-scale feature map contains all the feature information after each convolution, which can extract features from different receptive fields and multiple scales, effectively extract global and local features, meet the prediction requirements of neural networks, and reduce the impact of noise.

[0087] Furthermore, after obtaining multiple output subsets, all output subsets are concatenated and channel adjustment and feature fusion are performed through convolution. The convolution can be set to 1×1 convolution. In this embodiment, the 1×1 convolution concatenates all output subsets to form a multi-scale feature map with s channels. Of course, the multi-scale feature map can also have other channel numbers to meet the prediction requirements of the two neural networks, and no further limitations are imposed here.

[0088] In this embodiment, the extraction module in the backbone can use a residual function, h(x) = x + f(x, θ), which, after training, can obtain more effective feature extraction information. Furthermore, the number of extraction modules in this backbone can be set to multiple, for example, c. The number of these c can correspond to the sorting number of sub-feature maps ai (i∈1,2,3,...,s) for each channel, such as c!. Finally, the output feature maps obtained from multiple extraction modules are fused to form the output feature map P. Feature fusion includes concatenation and convolution, ensuring that the feature information of each sub-feature map at each scale is reflected at each scale, making the information richer and more effective.

[0089] It should be noted that the convolution used in the extraction layer in this embodiment employs a 3×3 convolution kernel, which reduces the computational cost and retains more extracted features, forming effective feature information. Of course, 5×5 or 7×7 convolution kernels can also be used to achieve multi-scale feature extraction. Therefore, the extraction module in this embodiment references the Res2Net module to implement multi-scale extraction, forming a feature map with both global and local features.

[0090] Of course, backbones can also use ResNet residual network, Swin Transformer V2, Efficinet V2, CBAM-Res2Net50, etc. for feature extraction, without too many restrictions.

[0091] After obtaining the feature maps, the first and second neural networks will be explained in detail.

[0092] The first neural network receives the feature map extracted from the backbone, introduces a dynamic attention mechanism to generate a dynamic attention map related to the target object, and forms a classification feature vector corresponding to the category based on the dynamic attention map and the feature map. Unlike existing technologies, this first neural network, after introducing the dynamic attention map, fuses it with the originally received feature map to form a classification feature vector related to the corresponding category. Then, it obtains the category classification result of the target object ANA based on this classification feature vector. The dynamic attention map can be obtained after performing a corresponding query operation on the feature map.

[0093] Specifically, the feature matrix is ​​obtained by multiplying the second feature vector of the dynamic attention map with the first feature vector of the feature map, i.e., feature matrix L = ∑aij × pij. The first feature vector is the dynamic attention distribution element aij (i∈1,2,3,...,m, j∈1,2,3,...,n) on the dynamic attention map A, and the second feature vector is the feature element pij (i∈1,2,3,...,m, j∈1,2,3,...,n) on the feature map P. After that, the obtained feature matrix is ​​internally summed by weighted summation, etc. The weighted summation can be obtained by linearly transforming the input feature matrix L through the weight matrix W to obtain the classification feature vector E, i.e., E = W × L + b, where b is the bias.

[0094] Subsequently, the classification feature vector formed during the fusion of the dynamic attention map and the feature map is fused with the feature map using the concept of residuals. For example, after average pooling, the feature map yields a global feature vector, such that the D×H×W (depth, height, and width) of the global feature vector corresponds to the classification feature vector E. The sum of the two yields the total feature vector U, i.e., the total feature vector U = P' + E = P + W × L(A, P') + b = P + E(A, P'), where A is the dynamic attention map, and P' is the global feature vector obtained after average pooling of feature map P. Therefore, the classification feature vector E is based on the feature information of the original feature map P. Furthermore, it also possesses information related to the dynamic attention map based on the feature map P. This allows the subsequent activation function to further combine the dynamic attention map with the original feature map when calculating the classification feature vector. This means that it takes into account the feature information of attention without losing the relevant features of the original feature map, obtaining as many features as possible while reducing noise. The receptive field is mapped to the global features and attention features, which can significantly reduce the risk of prediction inaccuracy. Moreover, the concept of residual connection is used to make the convolution A and P better approximate the ideal feature map, avoiding the excessive suppression of original information by the attention mechanism and preserving multi-scale context.

[0095] After obtaining the total feature vector U, the first neural network further uses an activation function to calculate the category classification result. In this embodiment, the activation function is the sigmoid function. The sigmoid function converts the total feature vector U into the conditional probability corresponding to each category. That is, when the number of channels corresponds to the number of categories, the higher the conditional probability is when the predicted category of the image to be detected matches the relevant kernel type, thus obtaining the category classification result.

[0096] The attention mechanism adopted in the first neural network can be a self-attention mechanism, using the QKV pattern. The input feature map is divided into three feature matrices by the weight matrix W2: the key matrix K, the attention matrix Q, and the value matrix V. When the attention matrix Q is used as a query, the scoring function transforms the key matrix K through the attention matrix Q. The transformed matrix is ​​then added to the value matrix V and subjected to a softmax function operation to obtain the dynamic attention map A. In this embodiment, the scoring function can be a dot product function, specifically S(Q,K)=K. T ×Q, or alternatively, a scaled dot product function can be used, S(Q,K)=K T ×Q / √D, or additive functions, bilinear functions, etc., can all be used to obtain the dynamic attention graph A. Furthermore, for classification of multiple target objects, this invention can also introduce a multi-head attention strategy. That is, a QKV module is used as a dynamic attention subgraph corresponding to one category. When Q is introduced as Q1, Q2, ..., Qn into a QKV module with the same number of categories, dynamic attention subgraphs corresponding to each category can be obtained, and after fusion, a global dynamic attention graph is obtained.

[0097] Referring to Figure 4, the first neural network in a specific embodiment includes a normalization module, an attention mechanism module, and a processing module. The normalization module includes a dynamic convolutional layer, which uses a 1×1 convolutional kernel to transform the original D×H×W feature map P into a C×H×W normalized feature map P0, where C represents the number of categories. This ensures that the number of channels in the normalized feature map P0 matches the number of each category, thereby enabling the acquisition of the corresponding results for each category in subsequent results. The attention mechanism module employs a dynamic attention mechanism, introducing a fully connected layer. In this fully connected layer, the normalized feature map P0 generates multiple K dynamic attention maps Ak based on the softmax operation, where Ak∈R(H×W), such that each attention map corresponds to a category. Subsequently, the K dynamic attention maps Ak are dynamically weighted, using the argmax function to select the dynamic attention map Ak that is relevant to the predicted category. k* and Ak* are the dynamic attention maps with the largest values ​​among all dynamic attention maps Ak, i.e., the dynamic attention map with the highest probability in each category prediction process is selected. Then, the processing module enhances the residual features and outputs them in a classified manner. Specifically, the fully connected layer multiplies the dynamic attention map Ak* with the normalized feature map P0 point by point to obtain the enhanced feature map Pa. In the next connected layer, the residual connection is used to output P0+λ×Ak*×P0. Before the output, the feature map P0 is processed by global average pooling to obtain the global feature vector P'. At this time, the total feature vector U=P'+E(A,P')=P'+F(Ak*,P0)=P'+(P0+λ×Ak*×P0), where P' and P0 are obtained by simple transformation of the feature map P. The F(Ak*,P0) function can be obtained by referring to the E(A,P') transformation in the previous embodiment. Therefore, both vectors can be regarded as simple transformations of the feature map P. This formula is consistent with the concept of residual. When selecting the dynamic attention map Ak corresponding to the kernel type, the aforementioned process can also be achieved by weighting all dynamic attention maps Ak to obtain the feature map used for multiplication with the normalized feature map P0. Thus, the first neural network combined with the feature extraction unit significantly improves the model's ability to capture subtle differences, especially in tasks that distinguish highly similar categories, resulting in accurate output category classification results and effectively identifying all kernel types, including rare kernel types.

[0098] The second neural network introduces a dynamic mechanism, which endows the feature maps with the ability to dynamically adjust. Under this dynamic adjustment, the feature maps generate quantized feature vectors that represent the quantized information of the target object (ANA). These quantized feature vectors then provide the quantized information of the image being detected regarding the target object. The dynamic mechanism can dynamically adjust the weights, architecture, and activation of certain neurons or layers within the neural network. Implementation methods include dynamic weight generation, dynamic width / depth, spatial adaptation, and dynamic activation functions.

[0099] In this embodiment of the invention, when the dynamic mechanism is dynamic weight generation, methods such as meta-learning, conditional computation, weight generation networks, and dynamic routing can be selected. In this embodiment, a weight generation network is selected to dynamically adjust the second neural network. The weight generation network (such as a fully connected layer or a super network) predicts the convolutional kernel weights based on the input features. Assuming g is the weight generator, then wdynamic=g(x), which enables the same layer to use different weights for different inputs, thereby improving the model capacity. For example, a self-attention mechanism is introduced. This self-attention mechanism dynamically adjusts the attention weights based on the input feature map. Specifically, it can refer to the QKV mode in the previous embodiment, that is, using the input feature map as the attention query vector Q, so that the attention weights can be dynamically adjusted according to the query vector Q, and combined with the key matrix K, the scoring function, and the value matrix V to obtain the self-attention map. The self-attention map is combined with the original feature map to obtain the feature map required for prediction. This self-attention mechanism uses a lightweight weight generation network, which can dynamically change the attention weights according to the input features and extract the attention features in the feature map based on the attention weights.

[0100] In one embodiment, referring to Figure 5, the second neural network includes an input layer, a dynamic nonlinear layer, a linear layer, a normalization layer, a hidden layer, and an output layer connected in sequence. The linear layer, normalization layer, and hidden layer constitute a processing module, which performs two iterations within this module. In implementation, the feature map is input to the input layer. The dynamic nonlinear layer adaptively adjusts the dynamic weights w based on the input feature map, so that the dynamic weights w are input to the linear layer. The linear layer obtains a dynamic feature map based on the dynamic weights w and a bias b. The dynamic feature map undergoes channel normalization in the normalization layer and then enters the hidden layer where the activation function f introduces a nonlinear representation. After obtaining the initial feature matrix, it returns to the input node of the processing module, i.e., before the linear layer. The linear layer obtains a secondary dynamic feature map based on the dynamic weights w and a bias b. The secondary dynamic feature map undergoes channel normalization in the normalization layer to obtain a secondary feature matrix. Finally, the secondary feature matrix is ​​reduced in dimension by a fully connected layer, mapping the optimized feature vector into the output layer to conform to the titer regression prediction. The processing module, which uses a two-loop configuration to process the feature map, can better extract relevant features from the feature map, making the relevant information more accurate. Of course, more loops can be set, all within the scope of this embodiment. In this embodiment, the activation function is ReLU, which can filter out feature information that meets the requirements of nonlinear sparse features, thus satisfying the prediction of quantized information.

[0101] In the embodiment shown in Figure 5, the dynamic nonlinear layer can be implemented as a dynamic weighting mechanism. Dynamic weights are generated in real-time based on the input data, i.e., the multi-scale feature map. These dynamic weights are then input into the linear layer for further processing, forming an effective perceptual system. The second neural network in this embodiment is an improved multilayer perceptron (MLP) architecture. It enhances the model's adaptability to complex data distributions by introducing dynamic mechanisms (such as weight generation, adaptive input calculation, or dynamic network structure). Alternatively, the dynamic nonlinear layer of the second neural network in this embodiment can also incorporate a self-attention mechanism layer, where the adjusted parameter is the query vector Q, which can also effectively extract feature information.

[0102] Besides self-attention mechanisms, dynamic nonlinear layers can also dynamically adjust their width and depth. For example, in a neural network architecture with multiple neurons, after the input feature map is fed into the input layer, the activated neurons in the next layer are adjusted according to the feature map's characteristics to partially meet the sparsity requirements of the feature map, thereby obtaining more accurate titer information in subsequent steps. Alternatively, the number of network layers activated can be determined based on the complexity of the input, reducing computation and making the quantization results more accurate. Dynamic nonlinear layers can also dynamically adjust activation functions. For example, in the structure of the above embodiment, a hidden layer is set before the normalization layer, and the activation function in this hidden layer is adjusted. Or, the slope and intercept of the ReLU activation function in the hidden layer after the normalization layer can be adjusted according to the input.

[0103] Combining the multi-scale feature maps mentioned in the previous embodiments, the dynamic nonlinear layer in the second neural network can be dynamically adjusted according to the relevant information in the multi-scale feature maps. For example, when the image to be detected is a fluorescent image, the information such as the number of cells and cell distribution corresponding to the fluorescent points in the fluorescent image will be continuously amplified during the convolution process of the multi-scale feature maps, so that the dynamic nonlinear layer can better rely on the feature information in the feature maps. For example, in the self-attention adjustment mechanism, after the dynamic nonlinear layer projects the feature information to various scales, the attention of the attention mechanism is not a fixed attention weight. The attention can be amplified or reduced according to the information at different scales, so as to obtain more accurate prediction results. Or, some kernel type information may have different manifestations in feature maps at different scales. Therefore, the feature map includes not only the relevant density information of the corresponding target object, but also hides the microscopic information of different kernel types. Different kernel types may have different titers at the same density. Therefore, the dynamic adjustment on which the self-attention mechanism is based can be extended to multi-scale feature maps at different scales. It pays attention to both density and distribution features, as well as kernel type features, so that the receptive field is projected to a higher feature dimension. While maintaining the indirectness of MLP, it greatly improves the expressive power and efficiency, and can better achieve quantitative information prediction.

[0104] According to the image processing model in the embodiment, referring to Figure 1, the image recognition prediction method of the first aspect embodiment includes the following steps:

[0105] S100, input the image to be detected into the feature extraction unit;

[0106] In this embodiment, the image to be detected is an indirect immunofluorescence image, which displays the fluorescent cell outlines, cell nuclei, and certain components produced during the cell division cycle. Based on these fluorescence characteristics, the karyotype and titer of the subject can be identified, and subsequent information about the subject can be obtained based on the karyotype and titer. The input process can be externally fed into the backbone of the image processing model. External inputs can include wireless transmission, wired transmission, external devices, mobile terminals, etc., or from the ANA detection device used by the image processing model. This ANA detection device can perform automated IIF detection based on the collected samples; this has already been implemented in existing technologies and will not be described in detail here. After the detection is performed, the corresponding indirect immunofluorescence image is obtained, and this indirect immunofluorescence image is input into the image processing model.

[0107] S200, Obtain the feature map extracted by the feature extraction unit based on the image to be detected;

[0108] This feature map can be implemented based on networks such as ResNet, EfficientNet, Swing Transformer, CBAM-Res2Net, and Res2Net, as described in the previous embodiments, and will not be repeated here.

[0109] The feature map extracted in this embodiment includes multi-scale feature maps, specifically including the following steps:

[0110] S201, Obtain the original feature map based on the image to be detected. The original feature map is the indirect immunofluorescence image in the aforementioned embodiment.

[0111] S202, the original feature map is divided into multiple sub-feature maps according to the number of channels. During the division, a 1×1 convolution kernel is used to group the original feature map into multiple sub-feature maps, such as ai (i∈1,2,3,...,s) in the previous embodiment.

[0112] S203, so that the multiple sub-feature maps are respectively processed by the corresponding multi-scale process and the corresponding output is a multiple output subset, such as bi (i∈1,2,3,...,s) in the previous embodiment;

[0113] S204: After concatenating all output subsets, channel adjustment and feature fusion are performed through convolution to obtain a multi-scale feature map. This channel adjustment and feature fusion uses 1×1 convolution to reconstruct a feature map with s channels containing multi-scale feature information.

[0114] This embodiment includes the following when performing multi-scale processing:

[0115] S2031, For the first sub-feature map, directly output as the output subset;

[0116] S2032, For the second sub-feature map, the output after convolution of the sub-feature map is the output subset;

[0117] S2033, for the i-th sub-feature map, add the i-th sub-feature map to the (i-1)-th sub-feature map after convolution to obtain the fused feature map, and then convolve the fused feature map to output the output subset, where i is greater than or equal to 3.

[0118] Therefore, each sub-feature map corresponds to feature information at different scales. For example, the first sub-feature map corresponds to a one-dimensional scale, the second sub-feature map corresponds to a two-dimensional scale, and the i-th sub-feature map corresponds to an i-dimensional scale. The multi-dimensional scale includes relevant features from all the previous dimensions, encompassing feature fusion at different scales. It can map both global and local features. Furthermore, multi-scale processing can also set up multiple extraction modules that can independently execute the above processing methods, thus completing the multi-scale feature map within the entire feature map view.

[0119] S300, based on the feature map, respectively, the first neural network executes the first prediction process and the second neural network executes the second prediction process;

[0120] That is, after the backbone extracts the feature map, it transmits the feature map to the first neural network and the second neural network respectively, which can be an internal transmission process within the image processing model.

[0121] S400, obtain the category classification result of the target object in the image to be detected based on the prediction result of the first prediction process, and obtain the quantitative information of the target object based on the prediction result of the second prediction process;

[0122] When applied to ANA (antinuclear antibody) in target groups, the classification results are represented by the karyotype of the antinuclear antibody, such as H - homogeneous, DFS - dense fine granular, ACA - centromere, S - granular, nucleoid, N - nucleolar, M - nuclear membrane, polymorphic linear, cytoplasmic fibrillary, cytoplasmic granular, cytoplasmic reticulum & mitochondrial-like, G - cytoplasmic polarity & Golgi apparatus-like, RR - cytoplasmic rod-like, and karyotype of the mitotic stage. Quantitative information is usually represented by titer information, with the titer being the reciprocal of the dilution. Karyotype indicates antibody type, and titer reflects antibody level; the two results can jointly indicate the probability and severity of a specific disease.

[0123] When the first neural network performs the first prediction process, it includes at least the following steps:

[0124] S310, Generate a dynamic attention map related to the target object based on the feature map;

[0125] S320, which forms a classification feature vector for the corresponding category based on the dynamic attention map and feature map;

[0126] S330: Obtain the category classification result of the target object from the classification feature vector.

[0127] An attention mechanism is introduced, which focuses attention on effective classification features in the feature map during the first prediction process, such as cell number and cell distribution. Since the feature map contains feature information at multiple scales, the attention may focus on a set of unique parts of various feature types on the corresponding karyotype, such as the relevant parts of cell distribution for independent karyotypes, the unique parts of cell clarity, etc., so that it notices the unique feature parts of each karyotype.

[0128] In this embodiment, the dynamic attention map can be generated by the following steps:

[0129] S311, the feature map is converted into a transformed feature map with the same number of channels as the number of target object categories through a specific dynamic convolutional layer, that is, the feature map is reduced in dimensionality so that the transformed feature map has a feature sequence corresponding to each kernel type in each dimension, and the transformed feature map is the normalized feature map in the aforementioned embodiment;

[0130] S312, perform classification regression on the transformed feature map to obtain a dynamic attention map, that is, introduce the query vector Q or the corresponding probability weight matrix p, and then perform the Softmax function operation to obtain a dynamic attention map in which the sum of each probability element is 1.

[0131] After generating the dynamic attention map, it is combined with the feature map to form classification feature vectors corresponding to each category. Specifically, the second feature vector of the dynamic attention map is multiplied by the first feature vector of the feature map to obtain a feature matrix. Then, the feature matrices are summed internally using a weight matrix to obtain the classification feature vector. This classification feature vector can display the attention probability corresponding to each category kernel type within the view of the original feature map, so as to better achieve kernel type prediction.

[0132] In the process of obtaining the category classification result of the target object from the classification feature vector, the feature map is averaged to obtain the global feature vector. Average pooling can gather some features in the feature map, thereby obtaining the global feature vector corresponding to the dimension of the classification feature vector. The global feature vector and the classification feature vector are added to obtain the total feature vector corresponding to the category classification result. Finally, after the total feature vector is processed by the sigmoid function for probability conversion, the kernel with the highest probability on the corresponding kernel type is the predicted category classification result.

[0133] The second neural network, when executing the second prediction process, includes at least the following steps:

[0134] S340, which uses a dynamic network to give feature maps a dynamic mechanism;

[0135] S350 generates quantized feature vectors for quantized information of the target object based on feature maps with dynamic mechanisms.

[0136] S360 obtains quantization information about the target object in the image to be detected from the quantization feature vector.

[0137] In this embodiment, the second neural network achieves dynamic adjustment through a dynamic network. The dynamic network can be implemented as a weight generation network, which can automatically adjust the weights based on the feature information input to the second neural network. For example, in a self-attention mechanism, the attention weights are adjusted based on the input feature information so that the attention weights focus on the corresponding titer features, such as cell distribution and cell number. Of course, it can also be a dynamic network with corresponding methods such as dynamic width / depth adjustment, spatial adaptive adjustment, and dynamic activation function adjustment, without too many limitations.

[0138] With the addition of a dynamic mechanism, the feature map can dynamically generate a quantized feature vector for the quantified information of the target object. This quantized feature vector can more accurately reflect the titer-related features, and quantified information about the target object can be obtained based on this quantified feature vector.

[0139] The image recognition prediction method in this embodiment simultaneously performs category classification prediction and quantitative information prediction. The combination of the two prediction results is beneficial for analyzing the disease status of the tested person and greatly improves the overall prediction efficiency. In addition, compared with manual judgment, experienced medical staff need about 2 minutes to judge a single karyotype and about 5 minutes to judge a complex karyotype. The image recognition prediction method in this embodiment can complete the prediction in a few seconds or even less than 1 second, which significantly reduces the detection time. For a single image, it can not only identify a single karyotype and a single titer, but also identify multiple complex karyotypes and multiple complex titers from the image, with rich recognition types.

[0140] In a second aspect, embodiments of the present invention also propose an image information detection method that takes into account both model prediction and external detection, comprising:

[0141] S1, acquire the image to be detected;

[0142] When acquiring an image to be inspected, the image can be displayed on a corresponding interface so that it can be inspected by external entities such as humans or external systems. Specifically:

[0143] S101, The image to be detected is input into the system using the image information detection method via an external input channel;

[0144] S102, The image to be detected is displayed in the first area of ​​the display interface;

[0145] S103, the prediction results and detection results are displayed in the second and third areas of the display interface, respectively.

[0146] The display interface is the display module of the device used in the information detection method, such as a monitor. This interface has at least three areas: a first area, a second area, and a third area. The first area displays the image to be detected, allowing inspectors and reviewers to view it. The second area displays the prediction results of the image processing model in the image recognition prediction method. The third area displays the detection results from external input. Referring to Figure 6, the first area R1 is located in the center of the display interface, and the second area R2 and the third area R3 are located to the right of the first area R1.

[0147] S2, execute the image recognition and detection method described in the first aspect embodiment, and obtain the prediction result of the image to be detected based on the image processing model. The prediction result includes category classification prediction information and quantization prediction information of the target object in the image to be detected.

[0148] After obtaining the image to be detected, the image processing model outputs the prediction results based on the image to be detected, and displays the category classification prediction information and the quantization prediction information in the second region R2.

[0149] S3, obtain the detection results based on the image to be detected from the external input. The detection results include classification detection information and quantization detection information about the target object in the image to be detected.

[0150] The detection results from external input can be entered from an external system or manually input into the system from an external input module (signal input components such as keyboard and mouse), and displayed in the third area R3. The external system can be a separately configured ANA detection device, both of which are within the scope of this invention.

[0151] S4. Verify the prediction results and detection results to obtain the final detection result of the target object in the image to be detected.

[0152] During the verification process, the prediction results and the detection results can be checked to see if they are consistent. If they are consistent, it means that the detection accuracy of the target object ANA is relatively high, and the final detection result can be obtained through verification. If the final detection results are inconsistent, an additional verification node can be set up for verification. For example, the verification program can be set up to confirm that after obtaining the prediction results and the detection results, the two results are automatically compared. If the two results are consistent, the final detection result is output. If they are inconsistent, an error result is output or input into the next verification node.

[0153] In this embodiment, a verification system is embedded in S4 to further ensure the accuracy of the final detection result. For example, even if the prediction result and the detection result are consistent, there is still a possibility that both detections are incorrect. Therefore, a verification system is used to check the prediction result and the detection result. The verification of the prediction result and the detection result includes a first-level verification process, which includes:

[0154] S41, based on the image to be detected displayed in the first area of ​​the display interface, the prediction result displayed in the second area, and the detection result displayed in the third area, input the first-level verification result of the category classification and the first-level verification result of the quantification information of the target object in the image to be detected.

[0155] This input is an external input process, meaning that the detection and prediction results are reviewed by the reviewers. After obtaining the model prediction and detection results, further verification is performed based on the image to be detected to ensure the accuracy of the final detection results.

[0156] In addition, verifying the prediction results and the test results also includes a secondary verification process. The secondary verification process includes at least one consecutive verification node, and each verification node includes:

[0157] S42, based on the image to be detected displayed in the first area of ​​the display interface, the prediction result displayed in the second area, the detection result displayed in the third area, and the verification information based on the previous verification node, input the category classification verification result and the quantification information verification result of the target object in the image to be detected.

[0158] Specifically, when the verification node is the first verification node, the verification information of the previous verification node is replaced with the verification information of the first-level verification node, which is then replaced with the category classification first-level verification result and the quantitative information first-level verification result in the first-level verification process. If the current verification node is not the last verification node, the category classification verification result and the quantitative information verification result are input into the next verification node. If the current verification node is the last verification node, the category classification verification result and the quantitative information verification result input by the current verification node are the final detection results of the target object in the image to be detected.

[0159] Referring to Figure 7, the image processing model has processing flow A, detection result input processing flow B, first-level verification flow C, and second-level verification flow D. After A and B obtain the prediction result 'a' and detection result 'b' respectively, they are input into the first-level verification flow C. Flow C outputs the first-level verification result 'c'. Flow D consists of one or more verification nodes Dn (n∈1,2,3,...,N) connected sequentially. The results input to D1 are a, b, and c. At this point, the output result d1 is passed to the next node D2. D2 uses the information a, b, c, and d1 to further output the verification result d2 to the next node, until the final node, obtaining the final detection result. Of course, the current verification node can be implemented to have access to the relevant information of all previous nodes. For example, in node Dn, the information used is a, b, c, d1,..., dn-1, providing sufficient verification information. If the verification results of each node are inconsistent, the verification result of the last node is taken as the final detection result. The verification system can also be configured with a query module, which displays the processing nodes of each image to be detected in the fourth area R4 of the display interface.

[0160] This image information detection method is matched with the personnel structure within the review process unit, and the layer-by-layer review and verification effectively ensures the validity and accuracy of the final output.

[0161] In this embodiment, verifying the prediction results and detection results also includes a quality inspection process, which includes:

[0162] S43, when the final detection result and the prediction result are consistent, positively stimulate the image processing model based on the prediction result; when the final detection result and the prediction result are inconsistent, record the number of inconsistencies; record the number of inconsistencies based on the prediction result; when the number of inconsistencies reaches a certain value, record the current abnormal situation.

[0163] The final detection result is the category classification verification result and quantization information verification result input to the final verification node. The prediction result is the category classification prediction information and quantization prediction information of the target object in the image to be detected, obtained based on the image processing model. Consistency between the final detection result and the prediction result means that the category classification verification result and the category classification prediction information are consistent, and the quantization information verification result and the quantization prediction information are consistent. If any information is inconsistent, the final detection result and the prediction result are inconsistent. When the final detection result and the prediction result are consistent, this information is fed back to the trained image processing model, and the model is positively incentivized based on the current prediction result, such as by adding weights, to affirm the accuracy of the image processing model. When the final detection result and the prediction result are inconsistent, the number of inconsistencies corresponding to the current prediction result is recorded. When the number of inconsistencies reaches a certain value, this information is fed back to the trained image processing model, and the current anomaly is recorded for subsequent verification of the cause.

[0164] For example, when the category classification verification result and the category classification prediction information are consistent, this information is fed back to the trained image processing model, and the image processing model is positively stimulated for the current category's category classification prediction information. When the category classification verification result and the category classification prediction information are inconsistent, the number of inconsistencies corresponding to the current category's category classification prediction information is recorded; when the number of inconsistencies reaches a certain value, this information is fed back to the trained image processing model, and the current anomaly is recorded.

[0165] Next, regarding the image processing model mentioned in the first aspect of the embodiments described above, the present invention proposes a training method for the image processing model in a third aspect. This method includes at least one training cycle, and each training cycle includes:

[0166] S500, the training image set is input into the image processing model. The training image set contains multiple first images. Each first image has category information corresponding to the target object contained in the first image and quantization information about the target object in the first image.

[0167] The training image set is used to train the image processing model. It contains predetermined detection results, including category information and quantization information. During training, the images are used as input features to the image processing model. After obtaining the prediction results, the image processing model is compared with the corresponding detection results and the parameters in the image processing model are adjusted to obtain a trained model.

[0168] S600, the image processing model extracts feature maps based on the first image;

[0169] The image processing model architecture is the same as that in the first aspect embodiment. Therefore, the training image set is processed according to the processing steps for the image to be detected in the previous embodiment to extract relevant features from the first image and obtain a feature map.

[0170] S700, the feature maps are input into the first neural network and the second neural network of the image processing model respectively. The first neural network obtains the category classification result of the target object based on the feature maps, and the second neural network obtains the quantitative information of the target object based on the feature maps.

[0171] The first neural network performs at least the following steps when acquiring the feature map:

[0172] S710 generates a dynamic attention map related to category information based on the feature map;

[0173] S710 may include the following steps:

[0174] The feature matrix is ​​obtained by multiplying the second feature vector of the dynamic attention map with the first feature vector of the feature map. The feature matrix is ​​then summed to obtain the classification feature vector.

[0175] S720 generates classification feature vectors for corresponding categories based on dynamic attention maps and feature maps;

[0176] S730 obtains the category classification result of the target object from the classification feature vector.

[0177] S730 may include: performing flat pooling on the feature map to obtain a global feature vector, and obtaining the category classification result based on the global feature vector and the classification feature vector.

[0178] When acquiring feature maps, the second neural network performs at least the following steps:

[0179] S740, which uses a dynamic network to provide a dynamic mechanism for feature maps;

[0180] S750 generates quantized feature vectors for quantized information of the target object based on feature maps with dynamic mechanisms.

[0181] S760 combines the quantized feature vectors of each feature map to obtain the quantized information of the first image corresponding to each feature map.

[0182] The image processing model processes the first image in the training image set in the same way as the image to be detected, so we will not go into details here.

[0183] Specifically, after the image processing model obtains the prediction result for the corresponding first image, the parameters within the image processing model are optimized and adjusted using backpropagation based on the loss function and the actual detection results. For example, the weight parameters of the weight matrix between each connection layer in the image extraction unit, the first neural network, and the second neural network, or the relevant parameters of the dynamic network in the second neural network, are all within the adjustment range.

[0184] In one embodiment, the training method further includes a verification phase, which includes:

[0185] S800, input the test image set into the image processing model trained with the training image set. The test image set contains multiple second images, each of which has category information corresponding to the target object contained in the second image and quantization information about the target object in the second image.

[0186] S900, obtain the verification results output by the image processing model based on the test image set;

[0187] S1000 evaluates the performance of the image processing model based on the validation results.

[0188] The test image set is used to verify the performance of the image processing model trained on the training image set. It also contains established actual detection results, such as those provided by the gold standard method, including category and quantization information. Therefore, the results can be compared with the verification results output by the image processing model based on the test image set. Inconsistencies indicate prediction errors, while consistency indicates correct verification. Since the test image set contains numerous second images, predictions are performed on all second images in the test image set during each verification process, and the corresponding error rate can be obtained based on the error patterns, thereby further evaluating the model's performance.

[0189] This invention utilizes stochastic gradient descent, where, after dividing the image set into training and testing sets, after each training iteration on the first image, the model is validated on the second image in the testing set. The error rate of the image processing model on the testing set is calculated; if the error rate no longer decreases, the model has reached its optimal performance. In this embodiment, the optimizer (SGD, stochastic gradient descent) has a learning rate of 0.01 and a weight decay rate of 1e-4, with a regularization term to prevent overfitting. This embodiment also allows for dynamic adjustment of the learning rate. Initially, the learning rate as a hyperparameter is 0.01, and the StepLR learning rate scheduler is used to set the step size to 4 and the learning rate decay coefficient (gamma) to 0.1. Furthermore, this embodiment employs a learning rate warmup strategy, increasing or decreasing the learning rate during the first two training cycles to gently initiate the training process and ensure stable initialization of the model parameters. In addition, as can be extended, "error rate no longer decreasing" means that the error rate no longer decreases significantly within ten consecutive periods. That is, if the error rate of the ten periods following the current period is greater than that of the current period, based on the current period where the minimum value exists, it indicates that the model has reached its optimal performance state.

[0190] In one embodiment, the category information includes common category information and rare category information. Common category information includes common kernel types, and rare category information includes rare kernel types. Therefore, the number of first and second images corresponding to rare kernel types in the obtained training and test image sets is relatively small. When the corresponding sample size is small, it may affect the model's recognition and prediction performance for that category, biasing towards learning features of high-frequency categories while ignoring low-frequency types. Therefore, before the training image set is input into the image processing model, the training method in this embodiment further includes:

[0191] Adjust the weighting of the first image corresponding to rare category information in the training image set so that the probability of the image processing model selecting the first image containing rare category information in the training image set is higher than the probability of selecting the first image containing rare category information in the training image set before adjustment.

[0192] For example, in rare kernel types, the first image corresponding to the rare kernel type can be copied into the training image set, or the number of corresponding first images can be increased through data augmentation. Data augmentation methods can include resizing, random horizontal flipping, random vertical flipping, and normalization. This increases the weight of rare information in the training image set. Similarly, the second image corresponding to a rare category in the test image set can also have its weight increased through direct copying or data augmentation methods such as resizing and normalization. Furthermore, to increase the number of training samples, the above data augmentation methods can be used to increase the number of training samples, allowing the model to achieve optimal performance.

[0193] Furthermore, in this embodiment of the invention, an adaptive loss function can also be used to optimize model performance. Since model performance is used to evaluate two prediction results, and the architectures of the first and second neural networks corresponding to each prediction result are different, this embodiment designs an adaptive loss function that can simultaneously optimize and adjust the parameters of the first and second neural networks to achieve the optimal state. Specifically, this training method also includes:

[0194] S1100, obtain the first loss function value of the first neural network on each feature map regarding the category classification result, and obtain the second loss function value of the second neural network on each feature map regarding the quantization information.

[0195] Specifically, when the first neural network processes the first prediction process, the image processing model is optimized based on the first loss function corresponding to the category classification result. After optimization, the weight parameters obtained by the image processing model when processing the feature maps of each first image can be obtained. Similarly, when the second neural network processes the prediction process, the image processing model is optimized based on the second loss function corresponding to the quantization information result. After optimization, the weight parameters obtained by the image processing model when processing the feature maps of each first image can be obtained. If L1(y,x;θ) is set as the first loss function and L2(y,x;θ) is set as the second loss function, then based on the input feature x (feature map), parameter θ, and actual result y, the first loss function value of the first loss function is L1, and the second loss function value of the second loss function is L2.

[0196] S1200, Obtain the adaptive loss function value based on the first loss function value and the second loss function value;

[0197] This may include the following acquisition methods:

[0198] S1210, calculate the standardized weights of the first loss function value and the second loss function value, and then adaptively adjust the standardized weights to obtain the adaptive function value;

[0199] Continuing with the above embodiment as an example, after obtaining the first loss function value and the second loss function value, the standardized weights can be directly obtained through normalization. Alternatively, in another embodiment, the standardized weights are calculated based on the relative rate of change of the first loss function value and the second loss function value at adjacent time steps. An adjacent time step refers to the interval between two adjacent training cycles; for example, in the aforementioned formula, the current weight and the adjusted weight differ by one time step. After obtaining the standardized weights, softmax is then performed to normalize the weights and obtain the adaptive function value. Therefore, the adaptive adjustment of the standardized weights is performed through at least one of the following methods:

[0200] (1) Weighted average of the standardized weights;

[0201] (2) The standardized weights are dynamically adjusted based on the magnitudes of the first loss function value and the second loss function value;

[0202] (3) The standardized weights are dynamically adjusted according to the error types of the first loss function value and the second loss function value.

[0203] The above methods can all be used to adaptively adjust the loss terms of the loss function, namely the first loss function value and the second loss function value, to obtain the adaptive function value. Then, based on the adaptive function value, the two loss functions can be combined to obtain a new adaptive loss function.

[0204] Specifically, in this embodiment, the first loss function is the cross-entropy loss function, and the second loss function is the mean squared error loss function. After calculating the standardized weights of the two losses, they are weighted and averaged to obtain the final adaptive loss function. This method enables the model to learn in a balanced manner across different tasks, optimizing overall performance. Alternatively, the first loss function can also be one of the following: Focal loss function, Asymmetric loss function; and the second loss function can also be one of the following: mean absolute error loss function, Huber loss function.

[0205] Of course, the above two loss functions can also be obtained using other adaptive loss functions in existing technologies, and there are no further restrictions on this.

[0206] Based on the image processing model trained in the aforementioned embodiments, the overall performance evaluation of this invention during the verification phase is shown in Figure 9. It exhibits very high accuracy, prediction precision, and F1-score. The model's classification performance is shown in Figure 10. The trained image processing model can identify 14 karyotypes: H-homogeneous, DFS-dense fine granular, ACA-centromere, S-granular, nucleosite, N-nucleolar, M-nuclear membrane, nuclear polymorphism, cytoplasmic fibrils, cytoplasmic granules, cytoplasmic reticulum & mitochondrial-like, G-cytoplasmic polarity & Golgi apparatus-like, RR-cytoplasmic rod ring, and mitotic karyotypes. The model provides a comprehensive range of karyotypes and demonstrates high accuracy for each karyotype. Above 90%, the optimal model's F1-score, which is the harmonic mean of recall and precision, is a high value. For example, for the H-homogeneous type, the F1-score is 0.91 with 95% precision; for the ACA-centromere type, the F1-score is 0.96 with approximately 100% precision. As shown in Figure 11, the quantified information prediction performance shows that the trained image processing model can identify five titer information: 80, 160, 320, 640, and 1280. Within a certain error range, the accuracy is approximately 100% for each titer, and the F1-score is above 0.9 for each. Therefore, this image processing model has very high accuracy and precision.

[0207] The evaluation metrics include: recall, precision, accuracy (i.e., clinical data compliance rate), and F1-score.

[0208] In this embodiment, referring to Figure 12, five types of networks (ResNet, EfficientNet, Swin Transformer, CBAM-Res2Net, and Res2Net50) are used in the feature map extraction unit. All of them achieve excellent performance. The predicted f1-score corresponding to the kernel type reaches above 0.8, and the R2-score corresponding to the kernel titer reaches above 0.55. Among them, the backbone can obtain the highest score when using Res2Net50, with an f1-score of 0.8759 corresponding to the kernel type and an R2-score of 0.6625 corresponding to the titer, both of which have very accurate prediction effects.

[0209] In a fourth aspect, embodiments of the present invention also provide an image processing system 10 applying the image processing model of the foregoing embodiments, as shown in FIG8, the system comprising:

[0210] Input module 11 is used to acquire the input image to be detected;

[0211] The prediction module 12 inputs the feature map of the image to be predicted into the image processing model and obtains the prediction result of the image processing model.

[0212] The image processing model, when processing the image to be detected, includes:

[0213] The feature extraction unit extracts feature maps based on the image to be detected;

[0214] The first neural network executes a first prediction process based on the feature map, and the second neural network executes a second prediction process based on the feature map.

[0215] In a fifth aspect, embodiments of the present invention also provide a computer-readable storage medium comprising a stored program, wherein, when the program is running, it controls the computer-readable storage medium to execute the image recognition prediction method of the first aspect embodiment, the image information detection method of the second aspect embodiment, and / or the training method of the third aspect embodiment of the present invention when the computer-readable storage medium is in a device.

[0216] This invention also provides a computer device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When executed by the processor, the computer program implements the image recognition prediction method of the first aspect embodiment, the image information detection method of the second aspect embodiment, and / or the training method of the third aspect embodiment. To avoid repetition, these methods are not described in detail here. Alternatively, when executed by the processor, the computer program implements the functions of each model / unit of the control device in the embodiments. To avoid repetition, these functions are not described in detail here.

[0217] Computer devices include, but are not limited to, processors and memory. Those skilled in the art will understand that the above are merely examples of computer devices and do not constitute a limitation on computer devices. A computer device may include more or fewer components than illustrated, or a combination of certain components, or different components. For example, a computer device may also include input / output devices, network access devices, buses, etc.

[0218] The processor referred to can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0219] Memory can be an internal storage unit of a computer device, such as a hard drive or RAM. Memory can also be an external storage device of a computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal and external storage units. Memory is used to store computer programs and other programs and data required by the computer device. Memory can also be used to temporarily store data that has been output or will be output.

[0220] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).

[0221] While specific embodiments of the invention have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of the invention. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of the invention. The scope of the invention is defined by the appended claims.

[0222] Finally, it should be noted that the above description is merely the preferred embodiment of the present invention and is not intended to limit the present invention in any way. Any person skilled in the art can make many possible variations and simple substitutions to the technical solutions of the present invention using the disclosed methods and techniques without departing from the scope of the present invention; all of these variations fall within the protection scope of the present invention.

Claims

1. An image recognition prediction method applied to an image processing model, characterized in that, The image processing model includes a feature extraction unit, a first neural network, and a second neural network; the method includes: The image to be detected is input into the feature extraction unit; Obtain the feature map extracted by the feature extraction unit based on the image to be detected; Based on the feature map, the first neural network is made to execute a first prediction process and the second neural network is made to execute a second prediction process, respectively. Based on the prediction results of the first prediction process, the category classification result of the target object in the image to be detected is obtained, and based on the prediction results of the second prediction process, quantitative information about the target object is obtained. The first prediction process includes at least the following steps performed by the first neural network: Generate a dynamic attention map related to the target object based on the feature map; Based on the dynamic attention map and the feature map, a classification feature vector corresponding to the category is formed; The category classification result of the target object is obtained from the classification feature vector.

2. The method of claim 1, wherein, The process of forming a classification feature vector corresponding to a category based on the dynamic attention map and the feature map includes: The second feature vector of the dynamic attention map is multiplied by the first feature vector of the feature map to obtain the feature matrix, and the classification feature vector is obtained by adding the features within the feature matrix.

3. The method of claim 2, wherein, The step of obtaining the category classification result of the target object from the classification feature vector includes: The feature map is subjected to average pooling to obtain a global feature vector, and the category classification result is obtained based on the global feature vector and the classification feature vector.

4. The method of claim 1, wherein, The feature map includes a multi-scale feature map generated using an image processing network, and the first neural network forms the dynamic attention map based on the multi-scale feature map.

5. The method of claim 4, wherein, The generation of the multi-scale feature map includes the following steps: Obtain the original feature map from the image to be detected; The original feature map is divided into multiple sub-feature maps according to the number of channels; The multiple sub-feature maps are respectively subjected to corresponding multi-scale processing and output as multiple output subsets; After concatenating all the aforementioned output subsets, channel adjustment and feature fusion are performed through convolution to obtain the multi-scale feature map; The multi-scale processing includes: For the first sub-feature map, directly output it as the output subset; For the second sub-feature map, the output after convolution of the sub-feature map is the output subset; For the i-th sub-feature map, add the i-th sub-feature map to the (i-1)-th sub-feature map after convolution to obtain the fused feature map. Then convolve the fused feature map and output the output subset, i≥3.

6. The method of claim 1, wherein, The second prediction process includes at least the following steps performed by the second neural network: The feature map is given a dynamic mechanism by a dynamic network; A quantized feature vector for the target object is generated based on the feature map with a dynamic mechanism. The quantization information of the image to be detected regarding the target object is obtained from the quantization feature vector.

7. The method of claim 6, wherein, The dynamic mechanisms assigned to the feature map include: Dynamic weights are introduced into the feature map to form a dynamic nonlinear layer, wherein each feature in the dynamic nonlinear layer is formed by combining each feature in the feature map with dynamic weights, and the dynamic weights are dynamically set to different weight values ​​according to the different feature values ​​of each feature in the feature map.

8. The method of claim 1, wherein, The dynamic attention map is generated from the feature map by at least the following steps: The feature map is transformed into a transformed feature map with the same number of channels as the number of categories of the target object through a specific dynamic convolutional layer; The transformed feature map is subjected to classification regression to obtain the dynamic attention map.

9. An image information detecting method characterized by comprising: The method includes: Acquire the image to be detected; The image recognition prediction method as described in any one of claims 1 to 8 is executed to obtain a prediction result about the image to be detected based on the image processing model, wherein the prediction result includes category classification prediction information and quantization prediction information about the target object in the image to be detected; Obtain detection results based on the image to be detected, which are input from an external source. The detection results include category classification detection information and quantization detection information about the target object in the image to be detected. By verifying the prediction results and the detection results, the final detection result of the target object in the image to be detected is obtained.

10. The method of claim 9, wherein, The acquisition of the image to be detected includes: The image to be detected is input into the system applying the image information detection method via an external input channel; The image to be detected is displayed in the first area of ​​the display interface; The image information detection method further includes: The prediction result and the detection result are displayed in the second and third areas of the display interface, respectively.

11. The method of claim 10, wherein, The verification of the prediction results and the detection results includes a first-level review process, which includes: Based on the image to be detected displayed in the first area of ​​the display interface, the prediction result displayed in the second area, and the detection result displayed in the third area, input the first-level verification result of the category classification and the first-level verification result of the quantitative information of the target object in the image to be detected; The verification of the prediction result and the detection result further includes a secondary verification process, which includes at least one consecutive verification node, each of which includes: Based on the image to be detected displayed in the first area of ​​the display interface, the prediction result displayed in the second area, the detection result displayed in the third area, and the verification information of the previous verification node, input the category classification verification result and the quantification information verification result of the target object in the image to be detected; Wherein, when the current review node is the first review node, the review information based on the previous review node is replaced with the category classification first-level verification result and the quantitative information first-level verification result in the first-level review process; If the current review node is not the last review node, the category classification review result and the quantitative information review result will be input to the next review node; If the current verification node is the last verification node, then the category classification verification result and the quantization information verification result input by the current verification node are the final detection result of the target object in the image to be detected.

12. The method of claim 11, wherein, The verification of the prediction results and the detection results also includes a quality inspection process, which includes: When the final detection result is consistent with the prediction result, the image processing model is positively stimulated based on the prediction result; When the final detection result and the prediction result are inconsistent, the number of times the inconsistency occurs is recorded for the prediction result; when the number of times the inconsistency reaches a certain value, the current abnormal situation is recorded.

13. A method of training an image processing model as claimed in claims 1 to 8, characterized in that, The method includes at least one training cycle, each training cycle comprising: The training image set is input into the image processing model. The training image set contains multiple first images. Each first image has category information corresponding to the target object contained in the first image and quantification information about the target object in the first image. The image processing model extracts feature maps based on the first image; The feature maps are respectively input into the first neural network and the second neural network of the image processing model. The first neural network obtains the category classification result of the target object based on the feature maps, and the second neural network obtains the quantitative information of the target object based on the feature maps. When acquiring the feature map, the first neural network performs at least the following steps: Generate a dynamic attention map related to the category information based on the feature map; Based on the dynamic attention map and the feature map, a classification feature vector corresponding to the category is formed; The classification result of the target object is obtained from the classification feature vector; When acquiring the feature map, the second neural network performs at least the following steps: The feature map is given a dynamic mechanism by a dynamic network; A quantized feature vector for the target object is generated based on the feature map with a dynamic mechanism. The quantization information of the first image corresponding to each feature map is obtained by combining the quantization feature vectors of each feature map.

14. The method of claim 13, wherein, The process of forming a classification feature vector corresponding to a category based on the dynamic attention map and the feature map includes: The second feature vector of the dynamic attention map is multiplied by the first feature vector of the feature map to obtain the feature matrix, and the classification feature vector is obtained by summing the features within the feature matrix. The step of obtaining the category classification result of the target object from the classification feature vector includes: The feature map is subjected to average pooling to obtain a global feature vector, and the category classification result is obtained based on the global feature vector and the classification feature vector.

15. The method of claim 13, wherein, The method further includes a verification phase, which includes: The test image set is input into the image processing model trained by the training image set. The test image set contains multiple second images, each of which has category information corresponding to the target object contained in the second image and quantization information about the target object in the second image. Obtain the verification results output by the image processing model based on the test image set; The performance of the image processing model is evaluated based on the verification results.

16. The method of claim 13, wherein, The category information includes common category information and rare category information. Before being input into the image processing model, the training image set also includes: The weight of the first image corresponding to the rare category information in the training image set is adjusted so that the probability of the image processing model selecting the first image containing the rare category information in the training image set is higher than the probability of the first image containing the rare category information in the training image set before adjustment.

17. The method of claim 13, wherein, The method further includes: Obtain the first loss function value of the first neural network on each of the feature maps with respect to the category classification result, and obtain the second loss function value of the second neural network on each of the feature maps with respect to the quantized information; Obtain the adaptive loss function value based on the first loss function value and the second loss function value; The first neural network and the second neural network optimize the performance of the image processing model based on the adaptive loss function value.

18. The method of claim 17, wherein, The step of obtaining the adaptive loss function value based on the first loss function value and the second loss function value includes: Calculate the standardized weights of the first loss function value and the second loss function value, and then adaptively adjust the standardized weights to obtain the adaptive function value; The calculation of the standardized weights for the first loss function value and the second loss function value includes: The standardized weights are calculated based on the relative rate of change of the first loss function value and the second loss function value at adjacent time steps; The adaptive adjustment method for the standardized weights includes at least one of the following: (1) Weighted average of the standardized weights; (2) The standardized weights are dynamically adjusted based on the magnitudes of the first loss function value and the second loss function value; (3) The standardized weights are dynamically adjusted according to the error types of the first loss function value and the second loss function value.

19. An image processing system that applies an image processing model, the system comprising: The image processing model includes a feature extraction unit, a first neural network, and a second neural network; the image processing system includes: The input module is used to acquire the input image to be detected; The prediction module inputs the feature map of the image to be predicted into the image processing model and obtains the prediction result of the image processing model. The image processing model, when processing the image to be detected, includes: The feature extraction unit extracts a feature map based on the image to be detected; The first neural network performs a first prediction process based on the feature map, and the second neural network performs a second prediction process based on the feature map; The first prediction process includes: Generate a dynamic attention map of the target object in the image to be detected based on the feature map; A classification feature vector for the target object is formed based on the dynamic attention map and the feature map; The classification result of the target object is obtained from the classification feature vector; The second prediction process includes: The feature map is given a dynamic mechanism by a dynamic network; generating a quantization feature vector of quantization information of the target object based on the feature map endowed with the dynamic mechanism; obtaining quantization information of the target object from the quantization feature vector.

20. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is run by the processor to execute the image recognition prediction method of any one of claims 1 to 8, the image information detection method of any one of claims 9 to 12, or the training method of the image processing model of claims 13 to 18.