Image processing method and pathological image processing method

By performing multiple sub-image encoding processing and coding relationship learning on the image, an accurate image encoding sequence is generated, which solves the problem of inaccurate image processing results and improves the accuracy of image processing.

CN119919353APending Publication Date: 2025-05-02ZHEJIANG LAKESIDE DATA SCIENCE & APPLICATION LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411875038.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

In the prior art, the image processing model has the problem that the image processing results are inaccurate when processing image data.

Method used

By performing multiple sub-image encoding processing on the target object image, a target image encoding sequence is generated, which includes an initial image encoding sequence and inserted image information encoding. Then, these coding sequences are input into the coding processing sub-model of the image processing model for encoding relationship learning to generate the final target object image coding sequence.

Benefits of technology

By accurately characterizing multiple aspects of the target object image, the accuracy of image processing results is improved, and the problem of inaccurate image processing results is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919353A_ABST
    Figure CN119919353A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and a pathological image processing method, and the image processing method comprises the steps: determining a plurality of sub-images of a target object image, carrying out the coding processing of each sub-image, and obtaining a target image coding sequence, the target image coding sequence comprises an initial image coding sequence corresponding to each sub-image and an image information code inserted in each initial image coding sequence, and each image information code is used for representing a global feature of the corresponding sub-image; inputting the target image coding sequence into a coding processing sub-model of an image processing model, and performing coding relation learning on each initial image coding sequence and the image information codes inserted in each initial image coding sequence by using the coding processing sub-model to obtain a target object image coding sequence; and determining an image processing result corresponding to the target object image according to the target object image coding sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and in particular to an image processing method. One or more embodiments of this specification also relate to a pathological image processing method, an image processing model training method, a cancer computer-aided diagnosis method, two other image processing methods, a computing device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the continuous development of computer technology and artificial intelligence technology, neural network models can be applied to various data processing scenarios for data processing, for example, using image processing models to process image data.

[0003] The current image processing model has the problem of inaccurate image processing results when processing image data. Therefore, how to improve the accuracy of image processing results becomes a problem that needs to be solved. Summary of the invention

[0004] In view of this, an embodiment of this specification provides an image processing method. One or more embodiments of this specification also relate to a pathological image processing method, an image processing model training method, a cancer computer-aided diagnosis method, two other image processing methods, an image processing device, a pathological image processing device, an image processing model training device, a cancer computer-aided diagnosis device, two other image processing devices, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of this specification, there is provided an image processing method, including:

[0006] Determine multiple sub-images of the target object image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted in each initial image encoding sequence, and each image information code is used to represent a global feature of a corresponding sub-image;

[0007] Inputting the target image coding sequence into the coding processing sub-model of the image processing model, using the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, to obtain the target object image coding sequence;

[0008] An image processing result corresponding to the target object image is determined according to the target object image coding sequence.

[0009] According to a second aspect of the embodiments of this specification, there is provided an image processing apparatus, including:

[0010] A first encoding processing module is configured to determine a plurality of sub-images of a target object image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted in each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image;

[0011] A second encoding processing module is configured to input the target image encoding sequence into an encoding processing sub-model of an image processing model, and use the encoding processing sub-model to learn the encoding relationship of each initial image encoding sequence and the image information encoding inserted in each initial image encoding sequence to obtain a target object image encoding sequence;

[0012] The result determination module is configured to determine the image processing result corresponding to the target object image according to the target object image coding sequence.

[0013] According to a third aspect of the embodiments of this specification, a pathological image processing method is provided, comprising:

[0014] Determine multiple sub-images of the pathological image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted in each initial image encoding sequence, and each image information code is used to characterize the global features of the corresponding sub-image;

[0015] Inputting the target image coding sequence into the coding processing sub-model of the image processing model, using the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, to obtain a pathological image coding sequence;

[0016] An image processing result corresponding to the pathological image is determined according to the pathological image coding sequence.

[0017] According to a fourth aspect of the embodiments of this specification, a pathological image processing device is provided, comprising:

[0018] A first encoding processing module is configured to determine a plurality of sub-images of a pathological image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted into each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image;

[0019] A second coding processing module is configured to input the target image coding sequence into a coding processing sub-model of an image processing model, and use the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence to obtain a pathological image coding sequence;

[0020] The result determination module is configured to determine the image processing result corresponding to the pathological image according to the pathological image coding sequence.

[0021] According to a fifth aspect of the embodiments of this specification, there is provided an image processing model training method, comprising:

[0022] Determine a plurality of sub-images of a sample object image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted into each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image;

[0023] Inputting the target image coding sequence into the coding processing sub-model of the image processing model, using the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, to obtain a sample object image coding sequence;

[0024] Determining a sample image processing result corresponding to the sample object image according to the sample object image coding sequence;

[0025] A sample label corresponding to the sample object image is determined, and the sample label and the sample image processing result are used to perform model training on the encoding processing sub-model of the image processing model to obtain a trained encoding processing sub-model.

[0026] According to a sixth aspect of the embodiments of this specification, there is provided an image processing model training device, comprising:

[0027] A first encoding processing module is configured to determine a plurality of sub-images of a sample object image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted into each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image;

[0028] A second encoding processing module is configured to input the target image encoding sequence into an encoding processing sub-model of an image processing model, and use the encoding processing sub-model to learn the encoding relationship of each initial image encoding sequence and the image information encoding inserted in each initial image encoding sequence to obtain a sample object image encoding sequence;

[0029] A result determination module, configured to determine a sample image processing result corresponding to the sample object image according to the sample object image coding sequence;

[0030] The model training module is configured to determine the sample label corresponding to the sample object image, and use the sample label and the sample image processing result to perform model training on the encoding processing sub-model of the image processing model to obtain the trained encoding processing sub-model.

[0031] According to a seventh aspect of the embodiments of this specification, a method for computer-aided diagnosis of cancer is provided, comprising:

[0032] Determine multiple sub-images of the tumor image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted in each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image;

[0033] Inputting the target image coding sequence into a coding processing sub-model of an image processing model, and using the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, to obtain a tumor image coding sequence;

[0034] A tumor image processing result corresponding to the tumor image is determined according to the tumor image coding sequence.

[0035] According to an eighth aspect of the embodiments of this specification, a cancer computer-aided diagnosis device is provided, comprising:

[0036] A first encoding processing module is configured to determine a plurality of sub-images of a tumor image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted into each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image;

[0037] A second encoding processing module is configured to input the target image encoding sequence into an encoding processing sub-model of an image processing model, and use the encoding processing sub-model to learn the encoding relationship of each initial image encoding sequence and the image information encoding inserted in each initial image encoding sequence to obtain a tumor image encoding sequence;

[0038] The result determination module is configured to determine the tumor image processing result corresponding to the tumor image according to the tumor image coding sequence.

[0039] According to a ninth aspect of the embodiments of this specification, there is provided an image processing method, which is applied to a client of a medical system, comprising:

[0040] In response to a user's clicking operation on a display interface of the client, determining a medical image to be processed;

[0041] The medical image to be processed is sent to the server of the medical system, and the image processing result corresponding to the medical image to be processed returned by the server is received, wherein the image processing result is determined according to a medical image coding sequence, and the medical image coding sequence is obtained by learning the coding relationship of an initial image coding sequence corresponding to each sub-image included in a target image coding sequence and an image information coding inserted in each initial image coding sequence using a coding processing sub-model of an image processing model, the image information coding is used to characterize the global features of the corresponding sub-image, and the target image coding sequence is obtained by coding multiple sub-images of the medical image to be processed.

[0042] According to a tenth aspect of the embodiments of this specification, there is provided an image processing device, applied to a client of a medical system, comprising:

[0043] An image determination module, configured to determine a medical image to be processed in response to a user's selection operation on the display interface of the client;

[0044] The result receiving module is configured to send the medical image to be processed to the server of the medical system, and receive the image processing result corresponding to the medical image to be processed returned by the server, wherein the image processing result is determined according to the medical image coding sequence, and the medical image coding sequence is obtained by learning the coding relationship of the initial image coding sequence corresponding to each sub-image included in the target image coding sequence and the image information coding inserted in each initial image coding sequence using the coding processing sub-model of the image processing model, the image information coding is used to characterize the global features of the corresponding sub-image, and the target image coding sequence is obtained by encoding multiple sub-images of the medical image to be processed.

[0045] According to an eleventh aspect of the embodiments of this specification, there is provided an image processing method, which is applied to a cloud-side device, including:

[0046] A plurality of sub-images of a target object image sent by a receiving end-side device are encoded to obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information encoding inserted into each initial image encoding sequence, and each image information encoding is used to represent a global feature of a corresponding sub-image;

[0047] Inputting the target image coding sequence into the coding processing sub-model of the image processing model, using the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, to obtain the target object image coding sequence;

[0048] Determining an image processing result corresponding to the target object image according to the target object image coding sequence;

[0049] The image processing result is sent to the terminal side device.

[0050] According to a twelfth aspect of the embodiments of this specification, there is provided an image processing apparatus, applied to a cloud-side device, comprising:

[0051] A first encoding processing module is configured to receive multiple sub-images of a target object image sent by a terminal-side device, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted in each initial image encoding sequence, and each image information code is used to represent a global feature of a corresponding sub-image;

[0052] A second encoding processing module is configured to input the target image encoding sequence into an encoding processing sub-model of an image processing model, and use the encoding processing sub-model to learn the encoding relationship of each initial image encoding sequence and the image information encoding inserted in each initial image encoding sequence to obtain a target object image encoding sequence;

[0053] A result determination module is configured to determine an image processing result corresponding to the target object image according to the target object image coding sequence;

[0054] The result sending module is configured to send the image processing result to the terminal device.

[0055] According to a thirteenth aspect of the embodiments of this specification, a computing device is provided, including:

[0056] Memory and processor;

[0057] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned image processing method are implemented.

[0058] According to the fourteenth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and the computer program / instruction implements the steps of the above-mentioned image processing method when executed by a processor.

[0059] According to a fifteenth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above-mentioned image processing method when executed by a processor.

[0060] The image processing method in one or more embodiments of the present specification, during the process of processing the target object image, will encode multiple sub-images of the target object image, so as to obtain a target image coding sequence of the target object image, the target image coding sequence includes an initial image coding sequence corresponding to each sub-image, and an image information code inserted in each initial image coding sequence, each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence can characterize the local features of the corresponding sub-image; thereby, the features of each sub-image of the target object image can be accurately represented from multiple aspects through the initial image coding sequence and the image information coding.

[0061] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, and obtain the target object image coding sequence. Finally, according to the target object image coding sequence that can accurately express the image features of the target object image and takes into account the correlation between the codings, the image processing result corresponding to the target object image is accurately determined; the problem of inaccurate image processing results is avoided, and the accuracy of the image processing results is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is a schematic diagram of a technical solution for image processing provided by an embodiment of this specification;

[0063] Figure 2 is an application schematic diagram of an image processing method provided by an embodiment of this specification;

[0064] Figure 3 is a flow chart of an image processing method provided by an embodiment of this specification;

[0065] Figure 4 is a processing flow chart of an image processing method provided by an embodiment of this specification;

[0066] Figure 5 is a flow chart of a pathological image processing method provided by an embodiment of this specification;

[0067] Figure 6 is a flowchart of an image processing model training method provided by an embodiment of this specification;

[0068] Figure 7 is a flow chart of a cancer computer-aided diagnosis method provided by one embodiment of this specification;

[0069] Figure 8 is a flow chart of an image processing method provided by an embodiment of this specification;

[0070] Fig. 9 is a flow chart of an image processing method provided by an embodiment of this specification;

[0071] Fig.10 is a structural schematic diagram of an image processing device provided by an embodiment of this specification;

[0072] Fig.11 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0073] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0074] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0075] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0076] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0077] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, which usually contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than 10 trillion model parameters. A large model can also be called a foundation model / foundation model. The large model is pre-trained with large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (LLM), a multi-modal pre-training model, etc.

[0078] When the big model is used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. The big model can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the big model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0079] First, the terms involved in one or more embodiments of this specification are explained.

[0080] SSM (State Space Model, SSM) is a mathematical model used to describe the behavior of dynamic systems. It is applicable to systems that change over time, such as control systems, economic systems, biological systems, etc. The state space model can provide a framework to analyze the internal states of the system and how these states evolve over time.

[0081] Mamba: A deep learning model based on the state space model (SSM) that uses the power of deep neural networks to improve the ability of traditional SSM in processing complex data. Specifically, Mamba combines deep learning techniques with state space models to more effectively capture nonlinear relationships and long-term dependencies in time series data.

[0082] WSI (Whole Slide Imaging) refers to whole slide images, that is, images obtained using whole slide imaging technology, which is a technology for digitizing pathological histological sections. This technology converts microscope slide specimens into high-resolution digital images through a high-resolution digital scanner, so that these images can be viewed, stored, managed and analyzed on a computer screen.

[0083] MIL (Multiple Instance Learning): refers to multiple instance learning, which is mainly used to process "bags" in a dataset rather than individual instances. In MIL, each bag contains multiple instances, and the label is usually given to the entire bag rather than a single instance. This setting is very useful in many practical applications, especially in the fields of medical image analysis, text classification, and chemical identification.

[0084] Pixel-Mamba: A new deep learning model based on Mamba, which is particularly suitable for survival analysis and other medical imaging tasks of whole-view digital slides (WSI). Different from the existing MIL learning framework, pixel-mamba directly regards the entire WSI as a sample and learns the representation of the entire WSI end-to-end, without the concepts of bag and instance in the MIL framework. The core idea of ​​Pixel-Mamba is to extract features directly from pixel-level data through an end-to-end learning process and generate a representation at the level of the whole slide image (WSI) for survival prediction.

[0085] CLS Token (Classification Token): refers to a classification token that is used to summarize the information of an entire sentence or text sequence.

[0086] region fusion: regional fusion.

[0087] Token expansion: Token expansion.

[0088] mean pooling: mean pooling.

[0089] Transformer: It is a sequence model based on the self-attention mechanism; the Transformer architecture mainly consists of an input part, a multi-layer encoder, a multi-layer decoder, and an output part.

[0090] LongViT model: is a visual Transformer model optimized for billion-pixel images.

[0091] With the continuous development of computer technology and artificial intelligence technology, neural network models can be applied to various data processing scenarios for data processing, for example, using image processing models to process image data. In the process of processing image data, the current image processing model has the problem of inaccurate image processing results.

[0092] For example, pathological diagnosis is a relatively accurate standard for diagnosing tumors and other related diseases. Pathological diagnosis using pathological tissue sections requires experienced doctors. The entire pathological examination process is very long, and the degree of digitization and intelligence is not high. AI analysis of digital pathological images can effectively reduce the burden on pathologists, improve the efficiency and accuracy of pathological diagnosis, and provide patients with high-quality treatment basis.

[0093] Pathological images represent the microenvironment in which tumor cells survive. Through artificial intelligence (AI) algorithms, the immune microenvironment of tumor cells can be reconstructed to achieve accurate prognosis analysis of cancer patients and further achieve precision medicine. If the entire pathology inspection process is digitized and AI analysis is performed on the pathology sections of each patient, the entire pathology department will be greatly transformed. Training a pathology inspection doctor requires years of training by experienced doctors. The misdiagnosis and missed diagnosis rates are high, and it is difficult to give quantitative analysis results for pathological diagnosis. Once the entire pathology diagnosis is intelligent, the AI ​​algorithm will provide doctors with accurate quantitative diagnosis reports and basis. Doctors only need to review the results, which can greatly improve the work efficiency of doctors. Compared with manual observation by doctors, AI algorithms can also provide quantitative analysis of tumor cells that doctors could not do before. Furthermore, intelligent pathology algorithms will also change the extremely unbalanced distribution of medical resources, so that people in medically backward areas will have the opportunity to enjoy more accurate medical services.

[0094] However, in view of the above technical objectives, there is a problem of inaccurate image processing results in the process of processing image data using the image processing model; for example, this specification provides a solution, Figure 1 is a schematic diagram of a technical solution for image processing provided by an embodiment of this specification, based on Figure 1 It can be seen that in the process of performing WSI analysis in this scheme, the paradigm adopts a two-stage process and a two-stage training method to process gigapixel-level panoramic pathology images (WSIs).

[0095] Specifically, the process of this scheme is as follows: 1) offline extraction of features of millions of small blocks cut from WSI to reduce the computational requirements, i.e. Figure 1 Part a of the .

[0096] By dividing a WSI into thousands of small blocks, the feature vector is extracted using a pre-trained network.

[0097] 2) Use algorithms such as multiple instance learning (MIL) to perform feature aggregation to learn task-specific features.

[0098] This paradigm can be encapsulated in an encoder-decoder framework, using a frozen encoder (i.e. Figure 1 The main focus of these works is to develop small-patch aggregation strategies to enhance the decoder’s capabilities. For example, the MIL scheme treats each WSI as a packet ( Figure 1 ), and form multiple aggregation models or embedding learning models ( Figure 1 Aggregator in to learn package-level representations.

[0099] By taking these features as a whole for multi-instance learning, or taking these small block vectors as Transformer Tokens to model the sequence relationship, the survival probability can be further predicted. Since the feature extractor is not optimized for survival analysis, the extracted features may lack the visual information that is critical to the task and may not be suitable for survival analysis.

[0100] The above methods free up the process of processing WSIs, so that we can focus on the design of the decoder component. However, these methods exhibit inherent performance bottlenecks: their performance ceiling is limited by the offline extracted features. Therefore, the features extracted by the commonly used pre-trained networks are not suitable for survival analysis, and end-to-end training can learn more suitable feature representations for survival prediction. However, end-to-end training of a stronger extractor on WSIs requires a GPU with huge memory, which is unrealistic for most researchers. The reason is that the huge number of small blocks in WSIs and the gradient calculation of the extractor require huge video memory. In addition, even if end-to-end training is achieved through large-scale distributed training or model segmentation techniques, limited annotated prognostic data (such as a few hundred cases) may cause the model to not converge or seriously overfit.

[0101] Based on the above solution, the disadvantages of this solution are:

[0102] Difficulty in efficiently locating key data: When processing gigapixel-level panoramic pathology images, accurately locating data that is critical for survival prediction remains a major obstacle.

[0103] Performance bottleneck: The performance ceiling of this solution is limited by offline feature extraction and cannot achieve better survival prediction.

[0104] High computing resource requirements: End-to-end training of a powerful feature extractor on WSIs requires huge GPU memory, which is difficult for ordinary researchers to meet.

[0105] Data scarcity problem: Even if end-to-end training is achieved by splitting the data across multiple machines and multiple cards, the annotated data used for survival prediction is very scarce, which may lead to model non-convergence or overfitting problems.

[0106] In addition, the above scheme can also provide a LongViT model ( Figure 1 The Memory Optimized ViT in LongViT is an end-to-end solution based on the Transformer architecture. This solution uses an optimized attention mechanism and uses 64 GPUs for pre-training on pathological images. This type of solution relies on a lot of computing power and cannot be used on a large scale. In other words, in the process of sending the entire WSI into the Transformer model for learning, since up to 64 A100 graphics cards are required for calculation, data segmentation to different graphics cards and local attention mechanism optimization are also used to optimize computing power; the entire solution process is very complicated, and the effect is average, making it difficult to use industrially.

[0107] Based on this, in this specification, an image processing method is provided. One or more embodiments of this specification also involve a pathological image processing method, an image processing model training method, a cancer computer-aided diagnosis method, two other image processing methods, an image processing device, a pathological image processing device, an image processing model training device, a cancer computer-aided diagnosis device, two other image processing devices, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.

[0108] See also Figure 2 , Figure 2 A schematic diagram of an application of an image processing method provided according to an embodiment of the present specification is shown. Figure 2 In the application scenario shown, the image processing model is deployed in the server 10, and the server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The client device 20 here may include but is not limited to: a smart phone, a tablet computer, a laptop computer, a PDA, a personal computer, a smart home device, a vehicle-mounted device, etc. The client device 20 can interact with the user through a graphical user interface to implement the call of the image processing model (such as a large model), thereby implementing the method provided in the embodiment of this specification.

[0109] In the embodiment of the present specification, the system composed of the client device 20 and the server 10 can execute the following steps: the client device 20 executes the operation of sending the pathological image to the server, and the server 10 executes: 1. Pathological image serialization; 2. Using Pixel-Mamba to obtain the slice-level representation of WSI; 3. Using the slice-level representation of WSI to execute the steps of downstream tasks.

[0110] Among them, pathological image serialization refers to dividing the whole slice image into multiple image regions, encoding the pixels in each image region into a token sequence, and inserting CLS tokens into the token sequence. Using Pixel-Mamba to obtain slice-level representation of WSI refers to using the pixel-mamba network in the pixel-mamba model architecture to model long-range dependencies on the token sequence inserted with CLS tokens and learn slice-level representation of WSI; using slice-level representation of WSI to perform downstream tasks refers to performing downstream tasks through slice-level representation of WSI to obtain corresponding image processing results.

[0111] It should be noted that, when the operating resources of the client device can meet the deployment and operating conditions of the image processing model, the embodiments of the present application can be carried out in the client device.

[0112] See also Figure 3 , Figure 3 A flowchart of an image processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0113] Step 302: Determine multiple sub-images of the target object image, encode each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted in each initial image encoding sequence, and each image information code is used to characterize the global features of the corresponding sub-image.

[0114] Among them, the target object image can be understood as an image that needs to be processed using a data processing model. The target object image can be an image of a target object located at a target site. The target site can be any part of the human body. The target object can be a tumor, human tissue, organ, etc. When the image processing method is applied to different scenarios, the target object image can be different; for example, when the image processing method is applied to a cancer diagnosis scenario, the target object image can be a tumor image; for example, when the image processing method is applied to a medical scenario, the target object image can be a medical image to be processed; for example, when the image processing method is applied to a pathology analysis scenario, the target object image can be a pathology image.

[0115] The tumor image can be understood as an image of a tumor located at a target site, and the target site may be a nasopharyngeal site, a pelvic site, a chest, or the like.

[0116] The medical image to be processed can be understood as a medical image used for medical diagnosis. For example, the medical image to be processed can be an X-ray image, a microscopic image, an ultrasound image, a nuclear magnetic resonance image, a radionuclide image, etc.

[0117] A pathological image can be understood as an image of a biological tissue sample obtained through a microscope or other imaging technology for medical diagnosis, research or teaching purposes. The pathological image can be a high-resolution image of a single or a series of smaller areas obtained from a biological tissue sample; or, the pathological image can be a panoramic pathological image.

[0118] The panoramic pathology image can be understood as a high-resolution digital image panoramic pathology image that can display the entire biological tissue structure. The panoramic pathology image can contain all the information of the biological tissue slice. For example, the panoramic pathology image can be a whole slice image (WSI) or a pathological tissue slice.

[0119] The sub-image can be understood as a local image divided or extracted from the target object image, for example, a local image area in a tumor image, a local image area in a panoramic pathology image.

[0120] The target image coding sequence can be understood as the pathological image representation of the target object image. The target image coding sequence is a sequence including multiple initial image coding sequences and image information codes inserted in each initial image coding sequence. That is to say, the target image coding sequence is composed of multiple initial image coding sequences and image information codes inserted in each initial image coding sequence.

[0121] The initial image coding sequence can be understood as a pathological image representation of a sub-image. The initial image coding sequence includes multiple initial image codes. The initial image code can be a code generated by encoding pixels in a sub-image. For example, the initial image code can be a Token, a matrix, etc.

[0122] The image information code can be understood as a code for representing the global features of the sub-image, and the image information code can be a learnable embedding vector or a token. For example, the image information code can be a CLS Token. The image information code can be inserted in the middle of the initial image code sequence; or, the image information code can be inserted at the beginning of the initial image code sequence as the first code of the initial image code sequence; or, the image information code can be inserted at the end of the initial image code sequence as the last code of the initial image code sequence.

[0123] In one or more embodiments provided in this specification, the step of determining a plurality of sub-images of a target object image, encoding each sub-image, and obtaining a target image encoding sequence includes:

[0124] The image processing unit of the image processing model is used to determine the multiple sub-images from the target object image, and perform encoding processing on each sub-image to obtain a target image encoding sequence.

[0125] Among them, the image processing unit can be understood as a unit in the image processing model for processing the target object image, and the image processing unit can be an image processing sub-model or one or more network layers in the image processing model.

[0126] Specifically, the image processing method provided in this specification utilizes the image processing unit of the image processing model to determine multiple sub-images from the target object image, and encodes each sub-image separately to obtain an initial image coding sequence corresponding to each sub-image; inserts the image information code into the initial image coding sequence; and constitutes a target image coding sequence with multiple initial image coding sequences into which the image information code is inserted.

[0127] Taking the application of the image processing method provided in this specification in the panoramic pathology image analysis scenario as an example, the image processing method is explained. The image processing method can be a panoramic pathology image analysis method, wherein the target object image is a panoramic pathology image, the sub-image is a local image area in the panoramic pathology image, the target image coding sequence is a Token sequence corresponding to the panoramic pathology image, the initial image coding sequence is a Token sequence corresponding to the local image area, the image information is encoded as a CLS Token, and the image processing model is a pixel-mamba model architecture; the pixel-mamba model architecture includes three parts: a WSI full-image serialization module, a pixel-mamba network structure, and a pathology task head.

[0128] Based on this, the image processing method provided in this specification is in the process of serializing the WSI full image, and the WSI full image serialization module (i.e., image processing unit) is responsible for serializing the WSI full image into a sequence. The serialization method is to divide the WSI into several 224x224 image areas (i.e., sub-images), insert a CLS Token in the middle of the Token sequence (i.e., the initial image coding sequence) of each area (i.e., image area), and finally merge the Token sequences of all areas together to form a long sequence (i.e., the target image coding sequence).

[0129] In the above embodiment, the image processing unit encodes each sub-image to obtain a target image encoding sequence, thereby obtaining a target image encoding sequence that can accurately express the image features of the target object image, which facilitates the subsequent determination of accurate image processing results.

[0130] In one or more embodiments provided in this specification, determining the multiple sub-images from the target object image, and encoding each sub-image to obtain a target image encoding sequence includes:

[0131] In the image processing unit, dividing the target object image into the plurality of sub-images;

[0132] Performing encoding processing on pixels contained in each sub-image respectively to obtain initial image codes corresponding to the pixels, and using the initial image codes to obtain initial image code sequences corresponding to the sub-images, wherein the initial image codes are used to characterize local features of the corresponding sub-images;

[0133] The image information code is inserted into each initial image code sequence respectively, and the target image code sequence is obtained by using a plurality of initial image code sequences into which the image information code is inserted.

[0134] Among them, the initial image code is multiple; the use of the initial image code to obtain the initial image code sequence corresponding to each sub-image can refer to arranging and combining the initial image codes of each sub-image according to the position information of the pixels in the sub-image, so as to obtain the initial image code sequence corresponding to each sub-image.

[0135] Among them, encoding the pixels contained in each sub-image separately to obtain the initial image code corresponding to the pixel can mean that a Z-shaped scanning method is used to encode the pixels contained in each sub-image separately to obtain the initial image code corresponding to the pixel, wherein one pixel corresponds to one initial image code (Token).

[0136] Using the above example, the WSI full-image serialization module can first divide the WSI into n 224x224 (hxw) regions (i.e., sub-images); and use each region as a scanning window to obtain a local sequence of length hxw+1 (i.e., initial image encoding) using zigzag scanning. The local sequence contains all pixel tokens of each region (each pixel is used as a token separately);

[0137] This method uses a zigzag scan to scan the pixels in each area to determine the Token sequence (i.e., local sequence) corresponding to each area. The zigzag scan can ensure that adjacent pixels are as close as possible in the sequence, thereby preserving the local spatial relationship.

[0138] Secondly, insert a Class Token into the Token sequence corresponding to each of the n regions.

[0139] Specifically, within each scanning window, a zigzag scan is applied and a CLS token (i.e., Class Token) is inserted at the center of the local sequence, resulting in a sequence length of (h×w+1) for each region; the zigzag scan is then extended to all regions, ensuring the spatial proximity of adjacent tokens. For an initial pixel-level token of I, the total sequence length is M=H×W+n.

[0140] Finally, the local sequences corresponding to the n regions are connected to form a global long sequence, which facilitates the subsequent acquisition of accurate image processing results.

[0141] Step 304: Input the target image coding sequence into the coding processing sub-model of the image processing model, and use the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence to obtain the target object image coding sequence.

[0142] Among them, the image processing model can be understood as a model used to process the target object image. For example, the image processing model can be a pixel-mamba model architecture; when the image processing method is applied to different scenarios, the image processing model can be different; for example, when the image processing method is applied to a cancer diagnosis scenario, the image processing model can be a tumor detection model; for example, when the image processing method is applied to a medical scenario, the image processing model can be a medical image processing model; for example, when the image processing method is applied to a pathology analysis scenario, the image processing model can be a pathology image analysis model.

[0143] The encoding processing sub-model may be understood as a sub-model in the image processing model for processing a target image encoding sequence. For example, the encoding processing sub-model may be a pixel-mamba network structure.

[0144] Among them, the coding relationship can be understood as the association relationship between multiple initial image codes contained in the initial image coding sequence and the association relationship between the multiple initial image codes and the image processing codes inserted in the initial image coding sequence. For example, the coding relationship can be a distance relationship, a long-distance dependency relationship, a long-range dependency relationship, etc.

[0145] The target object image coding sequence can be understood as an image coding sequence obtained after coding relationship learning is performed on the target image coding sequence using the coding processing sub-model; the target image coding sequence may include multiple target image processing codes after coding relationship learning. Alternatively, the target object image coding sequence may include multiple initial object image codes after coding relationship learning, and image information codes inserted in each of the initial image coding sequences and after coding relationship learning.

[0146] It should be noted that when the image processing method is applied to different scenarios, the target object image coding sequence may be different; the details are as follows.

[0147] For example, when the image processing method is applied to a cancer diagnosis scenario, the target object image may be a tumor image, and correspondingly, the target object image coding sequence may be a tumor image coding sequence; for example, when the image processing method is applied to a medical scenario, the target object image may be a medical image to be processed, and the target object image coding sequence may be a medical image coding sequence; for example, when the image processing method is applied to a pathology analysis scenario, the target object image may be a pathology image, and the target object image coding sequence may be a pathology image coding sequence.

[0148] In one or more embodiments provided in this specification, the step of inputting the target image coding sequence into a coding processing sub-model of an image processing model, and using the coding processing sub-model to learn coding relationships between the initial image coding sequences and the image information coding inserted into the initial image coding sequences to obtain the target object image coding sequence includes:

[0149] The target image coding sequence is input into the coding processing sub-model of the image processing model, wherein the coding processing sub-model includes a relational learning network layer and a coding fusion network layer.

[0150] The relationship learning network layer is used to perform coding relationship learning on each of the initial image coding sequences and the image information coding inserted in each of the initial image coding sequences to obtain a plurality of image coding sequences to be fused.

[0151] The coding fusion network layer is used to perform coding fusion on the plurality of image coding sequences to be fused, so as to obtain the target object image coding sequence.

[0152] Among them, the relation learning network layer can be understood as a network layer that implements coding relation learning in the coding processing sub-model. The relation learning network layer can be multiple, and the multiple relation learning network layers can be configured in different positions (i.e., different layers) of the model structure of the coding processing sub-model or in different coding processing units according to the needs of actual applications, so as to better perform coding relation learning. For example, the relation learning network layer can be a relation learning sub-module (MambaBlock).

[0153] The coding fusion network layer can be understood as a network layer that implements coding fusion in the coding processing sub-model. There can be multiple coding fusion network layers, and the multiple coding fusion network layers can be configured in different positions (i.e., different layers) of the model structure of the coding processing sub-model or in different coding processing units according to the needs of actual applications, so as to better perform coding fusion. For example, the coding fusion network layer can be a regional fusion sub-module (region fusion).

[0154] It should be noted that the coding processing sub-model may include multiple coding processing units (e.g., pixel-mambalayer), each of which may include a relational learning network layer and / or a coding fusion network layer. Among the multiple coding processing units, the input data of one coding processing unit may be the input data of the next coding processing unit; and the output target object image coding sequence is obtained through the processing of the multiple coding processing units.

[0155] The image coding sequence to be fused can be understood as an image coding sequence obtained after coding relationship learning and requiring coding fusion.

[0156] Using the above example, the pixel-mamba network in this method can be composed of multiple network layers (pixel-mamba layer), each layer can include two sub-modules, namely: Mamba Block, region fusion, the MambaBlock and region fusion can be understood as a sub-network layer in a pixel-mamba layer. The pixel-mamba layer is composed of two sub-network layers, Mamba Block and region fusion.

[0157] Among them, the mamba block is responsible for modeling long sequences. By performing forward propagation on the input long sequence, it learns the long-distance dependency (encoding relationship) between tokens in the sequence during the forward propagation process, and then outputs a token sequence of the same length as the original sequence (i.e., the image encoding sequence to be fused).

[0158] The region fusion module is responsible for fusing two similar regions in each layer (pixel-mamba layer). When the mamba block completes the relationship learning and outputs the Token sequence, the Token sequence output by the mamba block is input into the region fusion module for encoding fusion to obtain the fused Token sequence (i.e. the target object image encoding sequence).

[0159] Based on the above embodiments, it can be seen that the mamba architecture used in this method to construct the basic model of pathological images can effectively reduce video memory, and mamba is also good at capturing long sequence relationships. In addition, the region fusion module designed in this method can effectively fuse the Token sequences (initial image coding sequences) corresponding to two regions (sub-images), thereby further reducing the amount of calculation.

[0160] In one or more embodiments provided in this specification, the encoding relationship is a long-distance dependency relationship;

[0161] The method of using the relation learning network layer to perform coding relation learning on the initial image coding sequences and the image information coding inserted into the initial image coding sequences to obtain a plurality of image coding sequences to be fused includes:

[0162] Inputting the initial image coding sequences into the relational learning network layer, and determining a coding sequence to be processed from the initial image coding sequences in the relational learning network layer, wherein the coding sequence to be processed is any one of the initial image coding sequences;

[0163] Determine the long-distance dependency relationship between multiple image codes in the code sequence to be processed by performing forward propagation and backward propagation on the code sequence to be processed, wherein the multiple image codes are at least one initial image code and the image information code contained in the code sequence to be processed;

[0164] Based on the long-distance dependency, encoding adjustment is performed on each of the initial image encoding sequences to obtain the multiple image encoding sequences to be fused.

[0165] Continuing with the above example, the input of the pixel-mamba network structure is a pixel-level token sequence T (target image encoding sequence) generated by serializing the full slice image (i.e., the target object image). The initial size of each token τ (i.e., the initial image encoding) is C=3, corresponding to the RGB channel. As the token receptive field in the forward process (i.e., the forward propagation process) expands, the dimension C gradually increases to accommodate a richer representation.

[0166] The specific implementation of the relationship learning submodule (Mamba Block) is as follows:

[0167] 1. Input the above pixel-level token sequence T into the relational learning submodule (Mamba Block) of the first layer (pixel-mamba layer);

[0168] 2. In the Mamba Block, the forward state space model of the bidirectional state space model (SSM) is used to forward propagate the pixel-level token sequence T, and the front-to-back dependency (i.e., long-distance dependency) is captured from the beginning to the end of the forward propagation of the pixel-level token sequence T;

[0169] 3. In the Mamba Block, the reverse state space model of the bidirectional state space model (SSM) is used to perform reverse propagation processing on the pixel-level token sequence T after the forward propagation processing, and the back-to-front dependency (i.e., long-distance dependency) is captured in the reverse processing process from the end of the forward propagation to the start of the forward propagation for the pixel-level token sequence T;

[0170] 4. According to the dependency from front to back and the dependency from back to front, the original pixel-level token sequence T is adjusted, and a new token sequence T with the same length as the original pixel-level token sequence T is output, wherein the new token sequence T includes multiple initial image coding sequences adjusted by the dependency (i.e., multiple image coding sequences to be fused); and the CLS Token (i.e., image information coding) adjusted by the dependency is inserted into the multiple image coding sequences to be fused.

[0171] Based on the above embodiments, it can be seen that the method can better capture long-distance dependencies through Mamba Block, thereby improving the accuracy of subsequent image processing results.

[0172] In one or more embodiments provided in this specification, the coding processing sub-model further includes a coding extension network layer;

[0173] The method of using the coding fusion network layer to perform coding fusion on the plurality of image coding sequences to be fused to obtain the target object image coding sequence includes steps 1 to 3:

[0174] Step 1: Using the coding fusion network layer, the coding fusion is performed on the plurality of image coding sequences to be fused to obtain a fused image coding sequence.

[0175] The fused image coding sequence can be understood as an image coding sequence obtained by coding and fusing multiple image coding sequences to be fused; the number of codes in the fused image coding sequence is less than the number of codes in the multiple image coding sequences to be fused. The fused image coding sequence can be one or more.

[0176] The coding extension network layer can be understood as a network layer that implements coding extension in the coding processing sub-model. There can be multiple coding extension network layers, and the multiple coding fusion network layers can be configured at different positions (i.e., different layers) of the model structure of the coding processing sub-model or in different coding processing units according to the needs of actual applications, so as to better perform coding extension; for example, the coding extension network layer can be a token expansion sub-module (Token expansion).

[0177] It should be noted that the coding processing sub-model may include multiple coding processing units, each of which may include a relational learning network layer, a coding fusion network layer and / or a coding extension network layer. Among the multiple coding processing units, the input data of one coding processing unit may be the input data of the next coding processing unit; through the processing of the multiple coding processing units, the output target object image coding sequence is obtained. For the model structure of the coding processing sub-model, please refer to the following Table 1:

[0178] Table 1

[0179]

[0180]

[0181] Among them, Layer id can be understood as the number of pixel-mamba layer, the Token refers to the size of the token, 1×1 refers to the Token corresponding to 1×1 pixels, and 32×32 refers to the Token corresponding to 32×32 pixels; the Channels refers to the number of channels; Pixel-Mamba Layers refers to the modules contained in each Pixel-Mamba Layers, among which Mamba refers to the relationship learning submodule (Mamba Block), RF refers to the region fusion submodule (region fusion), TE refers to the token expansion submodule (Token expansion), Horizontal TE-Cat refers to the token expansion submodule that uses the vertical expansion (Horizontal) method and the series (Cat) method for coding expansion, and Vertical TE-Cat refers to the token expansion submodule that uses the horizontal expansion (Horizontal) method and the series (Cat) method for coding expansion.

[0182] In one or more embodiments provided in this specification, the step of using the coding fusion network layer to perform coding fusion on the plurality of image coding sequences to be fused to obtain a fused image coding sequence includes:

[0183] Inputting the plurality of image coding sequences to be fused into the coding fusion network layer;

[0184] In the coding fusion network layer, the multiple image coding sequences to be fused are divided into a first coding sequence set and a second coding sequence set, wherein the first coding sequence set includes a first image coding sequence to be fused, the second coding sequence set includes a second image coding sequence to be fused, and the number of sequences of the first image coding sequence to be fused is the same as the number of sequences of the second image coding sequence to be fused;

[0185] Determining a first image information code inserted in the first image coding sequence to be fused, and determining a second image information code inserted in the second image coding sequence to be fused;

[0186] Determine a coding similarity between a target first image information coding and a target second image information coding, wherein the target first image information coding is any one of the first image information codings, and the target second image information coding is any one of the second image information codings;

[0187] Based on the coding similarity, the first coding sequence set and the second coding sequence set are fused to obtain the fused image coding sequence.

[0188] The first image coding sequence to be fused refers to the image coding sequence to be fused in the first coding sequence set; the second image coding sequence to be fused refers to the image coding sequence to be fused in the second coding sequence set.

[0189] The first image information coding refers to the image information coding inserted in the first image coding sequence to be fused; the second image information coding refers to the image information coding inserted in the second image coding sequence to be fused.

[0190] The coding similarity can be understood as a value representing the similarity between the target first image information coding and the target second image information coding. The coding similarity can be cosine similarity or Euclidean distance similarity.

[0191] The step of fusing the first coding sequence set and the second coding sequence set based on the coding similarity to obtain the fused image coding sequence comprises:

[0192] Based on the coding similarity, the first image coding sequence to be fused and the second image coding sequence to be fused whose coding similarity is greater than or equal to a preset similarity threshold in the first coding sequence set and the second coding sequence set are fused to obtain the fused image coding sequence.

[0193] Continuing with the above example, the region fusion submodule is responsible for fusing two similar image regions in each pixel-mamba layer, thereby identifying similar image regions in each layer and merging them.

[0194] The specific implementation of the region fusion submodule is as follows:

[0195] 1. For a given n image regions containing M tokens, the n regions are divided into two equal region sets, namely the gallery set (i.e., the first coding sequence set) and the detection set (i.e., the second coding sequence set). Each set contains 2 / n regions, which can also be understood as each set containing the Token sequences corresponding to 2 / n regions (i.e., the first image coding sequence to be fused and the second image coding sequence to be fused);

[0196] 2. Calculate the cosine similarity (i.e., encoding similarity) between the CLS tokens (i.e., first image information encoding) of each region in the detection set and the CLS tokens (i.e., second image information encoding) of all regions in the gallery set;

[0197] 3. Based on cosine similarity, the Token sequences corresponding to the regions in the gallery set and the detection set are fused to obtain the fused Token sequence (i.e., the fused image coding sequence).

[0198] Based on the content of the above embodiments, it can be seen that the present method takes into account that whole slice images (WSIs) usually contain hundreds of millions of pixels, which poses a significant computational challenge to model training. However, much of the information is redundant, which provides an opportunity to reduce complexity. Therefore, Pixel-Mamba introduces a regional fusion submodule, which iteratively fuses similar regions layer by layer to reduce redundancy and computational overhead, thereby improving memory efficiency.

[0199] In one or more embodiments provided in this specification, the step of fusing the first coding sequence set and the second coding sequence set based on the coding similarity to obtain the fused image coding sequence includes:

[0200] Determine a target first coding sequence from the first coding sequence set, and determine a target second coding sequence from the second coding sequence set, wherein the target first coding sequence is the first image coding sequence to be fused inserted into the target first image information coding, and the target second coding sequence is the second image coding sequence to be fused inserted into the target second image information coding;

[0201] Determine the target first coding sequence and the target second coding sequence as a coding sequence pair, and determine a similar coding sequence pair from the coding sequence pairs according to the coding similarity;

[0202] Determine, from the similar coding sequences, a plurality of first image codes to be fused and the target first image information coding in the target first coding sequence, and determine a plurality of second image codes to be fused and the target second image information coding in the target second coding sequence;

[0203] According to the sequence position information of each first image code to be fused and the sequence position information of each second image code to be fused, determining corresponding associated codes for each first image code to be fused from the second image code to be fused;

[0204] respectively calculating average values ​​between the first image codes to be fused and the corresponding associated codes, and obtaining multiple fused image codes based on multiple average values;

[0205] The target first image information code and the target second image information code are encoded to obtain a fused image information code, and the fused image code sequence is constructed based on the fused image information code and the multiple fused image codes.

[0206] Among them, the coding sequence pair can be understood as a pair of coding sequences consisting of a target first coding sequence and the target second coding sequence.

[0207] The similar coding sequence pairs may be understood as coding sequence pairs whose coding similarity is greater than or equal to a preset similarity threshold; or, among multiple coding sequence pairs, the K coding sequence pairs with the highest similarity.

[0208] Wherein, determining similar coding sequence pairs from the coding sequence pairs according to the coding similarity comprises:

[0209] From the coding sequence pairs, determine the coding sequence pairs whose coding similarity is greater than or equal to a preset similarity threshold as the similar coding sequence pairs; or,

[0210] According to the coding similarity, multiple coding sequence pairs are sorted to obtain coding sequence pair sorting results; and a preset number of coding sequence pairs are selected from the coding sequence pair sorting results in a top-down manner as the similar coding sequence pairs.

[0211] Among them, the sequence position information of the first image code to be fused can be understood as the position information of the first image code to be fused in the first image code sequence to be fused; the sequence position information of the second image code to be fused can be understood as the position information of the second image code to be fused in the second image code sequence to be fused.

[0212] The associated code corresponding to the first image code to be fused may be understood as the second image code to be fused corresponding to the position of the first image code to be fused among the plurality of second image codes to be fused.

[0213] The fused image coding can be understood as the image coding obtained after coding fusion.

[0214] The fused image information code may be understood as an image information code obtained by encoding and fusing the target first image information code and the target second image information code.

[0215] Using the above example, the region fusion submodule uses cosine similarity to fuse regions as follows:

[0216] 1. Select k region pairs (i.e. coding sequence pairs) with the highest similarity scores for fusion;

[0217] The k region pairs (Top-k) with the highest similarity scores are selected for fusion.

[0218] 2. Merge these k region pairs by averaging the corresponding tokens in the token sequences from each pair of regions;

[0219] For a certain image area in the detection set, the tokens at each position in the area with the highest similarity in the gallery set (i.e., the first image code to be fused and the corresponding associated code) are averaged, and the token average value (i.e., the average value) is used as the new token (i.e., the fused image code), so that the two tokens belonging to the two image areas are fused into one token (i.e., the fused image code), thereby realizing the merging of the token sequences of the two image areas.

[0220] At the same time, the CLS Tokens corresponding to the two regions also need to be merged to obtain the merged CLS Token.

[0221] 3. Output the number of fused regions M, each of which has a new token sequence (fused image encoding sequence). The fused CLS Token is inserted into the new token sequence.

[0222] Based on the content of the above embodiments, it can be seen that the present method takes into account that much information in whole slice images (WSIs) is redundant, which provides an opportunity to reduce complexity. Therefore, Pixel-Mamba introduces a regional fusion submodule, which iteratively fuses similar regions layer by layer to reduce redundancy and computational overhead, thereby improving memory efficiency.

[0223] Step 2: Utilize the coding extension network layer to perform coding extension on the fused image coding sequence to obtain multiple extended image coding sequences.

[0224] The extended image coding sequence may be understood as an image coding sequence obtained by coding and expanding the fused image coding sequence.

[0225] In one or more embodiments provided in this specification, the step of using the coding extension network layer to perform coding extension on the fused image coding sequence to obtain the extended image coding sequence includes:

[0226] Inputting the fused image coding sequence into the coding extension network layer, and generating a coding feature map in the coding extension network layer according to a plurality of fused image codes in the fused image coding sequence;

[0227] Determining a target coding feature from a plurality of coding features included in the coding feature map, and expanding the coding feature map according to the target coding feature to obtain an expanded coding feature map;

[0228] Performing a shift operation on a plurality of coding features included in the extended coding feature map to obtain a shift coding feature map, and performing coding expansion based on the extended coding feature map and the shift coding feature map to obtain a plurality of extended codes;

[0229] An extended image coding sequence is generated based on the fused image information coding in the fused image coding sequence and the multiple extended codings.

[0230] The coding feature map can be understood as a feature map corresponding to the fused image coding sequence, and the coding feature map can be one or more; when there is one fused image coding sequence, the coding feature map can be one; when there are multiple fused image coding sequences, the coding feature map can be multiple. It should be noted that when there are multiple fused image coding sequences, coding extension can be performed for each fused image coding sequence to obtain an extended image coding sequence corresponding to each fused image coding sequence.

[0231] The coding feature can be understood as a feature contained in the coding feature map, for example, a feature corresponding to multiple pixels in the coding feature map.

[0232] The target coding feature can be understood as the coding feature of any row or column in the coding feature map. For example, the target coding feature can be the coding feature of the first column in the coding feature map or the coding feature of the first row in the coding feature map.

[0233] The expanded coding feature map can be understood as a feature map after coding expansion; the expanding the coding feature map according to the target coding feature to obtain the expanded coding feature map includes: adding the target coding feature to the coding feature map to obtain the expanded coding feature map.

[0234] The displacement coding feature map can be understood as a coding feature map obtained by performing a displacement operation on the extended coding feature map; the displacement coding feature map is obtained by performing a displacement operation on the multiple coding features contained in the extended coding feature map, including:

[0235] A horizontal displacement method or a vertical displacement method is adopted to move a plurality of coding features included in the extended coding feature map by a target distance to obtain a displacement coding feature map.

[0236] The extended coding can be understood as the image coding obtained after performing the coding extension operation; the extended image coding sequence can be understood as the image coding sequence obtained after performing the coding extension operation.

[0237] Continuing with the above example, the token expansion submodule is responsible for performing token expansion in the h and w dimensions of the image in the forward process to expand the receptive field of the token. The token expansion submodule can expand the token receptive field, gradually increasing it from 1×1 to 32×32, which is crucial for learning hierarchical representations.

[0238] Token expansion submodule, the specific implementation is as follows:

[0239] 1. Remove the CLS token from the token sequence of each region in the n regions (i.e., the fused image encoding sequence), and reshape the other feature tokens (i.e., multiple fused image encodings) into an h×w×C feature map (i.e., the encoded feature map).

[0240] 2. Apply the shift and merge operation to expand the feature map vertically or horizontally to obtain a new feature map;

[0241] Specifically, for vertical expansion, the shift merging in this method involves filling the first row (i.e., the target encoding features);

[0242] First, a new row is filled for the feature map by copying the first row into a new row, and a filled feature map (expanded coded feature map) is obtained; in the case of vertical expansion, the feature map (i.e., expanded coded feature map) is moved down by one position to obtain a shifted feature map (i.e., displacement coded feature map) after average shifting;

[0243] In the case of horizontal expansion, the feature map (i.e., the expanded coding feature map) is shifted one position to the right to obtain a shifted feature map (i.e., the displacement coding feature map) after average shifting;

[0244] Secondly, the tokens in the shifted feature map are used to perform encoding expansion with the original tokens at the corresponding positions in the original feature map to obtain multiple extended codes.

[0245] Finally, the removed CLS tokens are added to multiple extended codes (i.e., fused image information codes) to obtain the final token sequence for each region (i.e., extended image code sequence).

[0246] Based on the above embodiments, it can be seen that the token expansion of this method gradually expands the receptive field of the token as the network deepens, so that the Mamba block can capture multi-scale representations at all levels. This process enables Pixel-Mamba to extract hierarchical representations from full-slice images.

[0247] In one or more embodiments provided in this specification, the coding extension is performed based on the extended coding feature map and the displacement coding feature map to obtain multiple extended codes, including:

[0248] Determining a plurality of extended coding features in the extended coding feature map, and determining a plurality of displacement coding features in the displacement coding feature map;

[0249] Determining a corresponding associated displacement coding feature for each of the extended coding features from the plurality of displacement coding features according to the extended coding feature map position of each of the extended coding features and the displacement coding feature map position of each of the displacement coding features;

[0250] Combining the target extended coding feature and the corresponding associated displacement coding feature to obtain a plurality of coding feature pairs, wherein the target extended coding feature is any one of the plurality of extended coding features;

[0251] Merging the expanded coding features and the displacement coding features contained in each coding feature pair respectively to obtain a plurality of merged coding features, and determining a merged coding feature map based on the plurality of merged coding features;

[0252] Two adjacent merged coding features in the merged coding feature map are subjected to feature fusion to obtain a coding extension feature map, and the coding extension feature map is encoded to obtain multiple extension codes.

[0253] Among them, the expanded coding feature map position of the expanded coding feature can be understood as the position information of the expanded coding feature in the expanded coding feature map; the displacement coding feature map position of the displacement coding feature can be understood as the position information of the displacement coding feature in the displacement coding feature map.

[0254] The associated displacement coding feature corresponding to the extended coding feature may be understood as a displacement coding feature corresponding to the position of the extended coding feature among multiple displacement coding features.

[0255] The combined coding feature map can be understood as a feature map composed of multiple combined coding features. The coding extension feature map refers to a feature map obtained by extending the combined coding feature map.

[0256] Continuing with the above example, for the token expansion submodule, the process of applying the shift-and-merge operation to expand the feature map vertically or horizontally is as follows:

[0257] 1. The tokens in the shifted feature map (i.e., displacement coded features) are averaged (Avg) or concatenated (Cat) with the original tokens at the corresponding positions (i.e., associated displacement coded features) in the original feature map to generate a new feature map (h×w×C) with shifted merged tokens (i.e., merged coded features). The new feature map is referred to as the merged coded feature map.

[0258] 2. In the case of vertical expansion, for the tokens in the new feature map (i.e., merged coding features), the token expansion operation increases the receptive field by concatenating or averaging adjacent tokens (i.e., two adjacent merged coding features) along the vertical dimension; the size of the output Token (i.e., multiple expanded codes) is h / 2×w×2C (if concatenated) or h / 2×w×C (if averaged).

[0259] It should be noted that horizontal token expansion follows a similar process, but for the tokens in the new feature map (i.e., merged encoding features), shift merging and token expansion are applied in the w dimension, so that adjacent tokens in the w dimension (i.e., two adjacent merged encoding features) are concatenated or averaged; the size of the output Token (i.e., multiple expanded codes) is h×w / 2×2C or h×w / 2×C, depending on the operation (concatenation or averaging).

[0260] Based on the above embodiments, it can be seen that the token expansion of this method gradually expands the receptive field of the token as the network deepens, so that the Mamba block can capture multi-scale representations at all levels. This process enables Pixel-Mamba to extract hierarchical representations from full-slice images.

[0261] Step three: Determine the target object image coding sequence based on the multiple extended image coding sequences.

[0262] Specifically, after using the coding extension network layer to perform coding expansion to obtain multiple extended image coding sequences, in the coding processing sub-model, the relational learning network layer and the coding fusion network layer can also be used to perform the processing operations in the above embodiments on the multiple extended image coding sequences until the output data of the last layer of coding processing unit in the coding processing sub-model is obtained, and the output data is the target object image coding sequence.

[0263] Alternatively, the method may also use multiple extended image coding sequences as the target object image coding sequences and output them.

[0264] In the above embodiment, Pixel-Mamba first serializes each pixel of the WSI into a tag by scanning the entire slice. Then, the Pixel-Mamba network processes these tags layer by layer, gradually expanding its receptive field to capture hierarchical information. At each level, it uses a state space model (SSM) to capture the global context between tags, ensuring that long-range dependencies are effectively modeled. Finally, Pixel-Mamba learns a slice-level representation of the WSI (i.e., the target object image encoding sequence), which facilitates various output heads (i.e., task heads) to accurately perform downstream tasks.

[0265] Step 306: Determine the image processing result corresponding to the target object image according to the target object image coding sequence.

[0266] The image processing result can be understood as a processing result corresponding to the target object image obtained by processing the target object image coding sequence. When the image processing method is applied to different scenes, the image processing result can be different; the details are as follows.

[0267] When the image processing method is applied to a cancer diagnosis scenario, the target object image may be a tumor image, and correspondingly, the image processing result may be a tumor image processing result.

[0268] When the image processing method is applied to a medical scenario, the target object image may be a medical image to be processed, and the image processing result may be an image processing result corresponding to the medical image to be processed.

[0269] In the case where the image processing method is applied to a pathological analysis scenario, the target object image may be a pathological image, and the image processing result may be an image processing result corresponding to the pathological image.

[0270] When the image processing method is applied to a pathology analysis scenario, the image processing result can be understood as a pathology report analysis result; that is, the image processing method in this specification provides a pathology report tool through one or more of the above embodiments, which can perform pathology report analysis based on pathology images to obtain pathology report analysis results.

[0271] When the image processing method is applied to the cancer cell ratio detection scenario, the image processing result can be understood as the cancer cell ratio detection result; the image processing method in this specification realizes cancer cell ratio detection based on tumor images and obtains cancer cell ratio detection results through one or more of the above embodiments.

[0272] When the image processing method is applied to a prognostic analysis scenario, the image processing result can be understood as a patient survival index prediction result; the image processing method in this specification realizes the prediction of the patient survival index based on the medical image (i.e., the medical image to be processed) through one or more of the above embodiments, and obtains the patient survival index prediction result.

[0273] When the image processing method is applied to the cancer TNM staging scenario, the image processing result can be understood as the cancer TNM staging result; the image processing method in this specification implements cancer TNM staging based on tumor images and obtains cancer TNM staging results through one or more of the above embodiments. It should be noted that the cancer TNM staging result can be used to support experimental parameters for specific diseases.

[0274] In one or more embodiments provided in this specification, determining the image processing result corresponding to the target object image according to the target object image coding sequence includes:

[0275] The target object image coding sequence is processed by utilizing the task execution unit in the image processing model to obtain an image processing result corresponding to the target object image, wherein the task execution unit is associated with the target image processing task.

[0276] Among them, the task execution unit can be understood as a unit in the image processing model that performs various target image processing tasks based on the target object image coding sequence. The task execution unit can be a sub-model, one or more network layers; for example, the task execution unit can be a task head or an output head.

[0277] Among them, the target image processing task can be understood as a task that needs to be performed according to the target object. The target image processing task can be a prognosis analysis task, a cancer cell ratio detection task, a pathological analysis task, a cancer TNM staging task, etc.

[0278] Continuing with the above example, Pixel-Mamba learns the slice-level representation of WSI and performs downstream tasks such as image classification, tumor staging, and survival prediction through various output heads (i.e., task execution units) to obtain the corresponding task execution results (i.e., image processing results).

[0279] In one or more embodiments provided in this specification, the task execution unit is a task execution network layer;

[0280] The step of using the task execution network layer in the image processing model to process the target object image coding sequence to obtain the image processing result corresponding to the target object image includes:

[0281] Determining a plurality of target image information codes from the target object image code sequence, and performing average pooling processing on the plurality of target image information codes using an average pooling network layer in the image processing model to obtain an average image information code;

[0282] The task execution network layer in the image processing model is utilized to execute the target image processing task for the target object image according to the average image information encoding to obtain the image processing result.

[0283] The target image information coding may be understood as the image information coding obtained after being processed by the coding processing sub-model in the image processing model.

[0284] Using the above example, for the pathology task head, after the input sequence (i.e., the target image coding sequence) is forward-calculated by pixel-mamba, a series of Tokens (i.e., the target object image coding sequence) will be output. The CLS Token (i.e., the target image information coding) will be taken out and mean pooling will be performed using the average pooling network layer as the representation of the series output (i.e., the average image information coding). Finally, the Token (i.e., the average image information coding) obtained by mean pooling is input into a fully connected layer (task head) to complete the downstream task and obtain the corresponding task execution result (i.e., the image processing result).

[0285] The image processing method in one or more embodiments of the present specification, during the process of processing the target object image, will encode multiple sub-images of the target object image, so as to obtain a target image coding sequence of the target object image, the target image coding sequence includes an initial image coding sequence corresponding to each sub-image, and an image information code inserted in each initial image coding sequence, each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence can characterize the local features of the corresponding sub-image; thereby, the features of each sub-image of the target object image can be accurately represented from multiple aspects through the initial image coding sequence and the image information code.

[0286] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, and obtain the target object image coding sequence. Finally, according to the target object image coding sequence that can accurately express the image features of the target object image and takes into account the correlation between the codings, the image processing result corresponding to the target object image is accurately determined; the problem of inaccurate image processing results is avoided, and the accuracy of the image processing results is improved.

[0287] The following combination Figure 4 , taking the application of the image processing method provided in this specification in the medical record image processing scenario as an example, the image processing method is further described. Figure 4 A processing flow chart of an image processing method provided by an embodiment of the present specification is shown. The image processing method can be understood as a method for end-to-end modeling and training of pathological images based on a state-space model.

[0288] Figure 4This is the pixel-mamba model architecture proposed in this method, which mainly includes three modules: 1. WSI full-image serialization module, 2. pixel-mamba network structure, and 3. pathology task head.

[0289] It should be noted that in order to address the challenges associated with gigapixel WSI, the Pixel-Mamba model architecture in this method first serializes each pixel of the WSI into a label through the entire slice scan. Secondly, the Pixel-Mamba network is used to process these labels layer by layer, gradually expanding its receptive field to capture hierarchical information. At each level, it uses a state-space model (SSM) to capture the global context between labels, ensuring that long-range dependencies are effectively modeled. Finally, Pixel-Mamba learns slice-level representations of WSI and performs downstream tasks through various output heads; such as image classification, tumor staging, and survival prediction.

[0290] Due to the linear computational complexity of SSM and the design including label extension and region fusion, the framework supports end-to-end training with full image WSI input and outperforms existing state-of-the-art two-stage MIL solutions on multiple tasks.

[0291] Among them, for WSI full-graph serialization:

[0292] The WSI full image serialization refers to the full slice image serialization; the WSI full image serialization module is responsible for serializing the WSI full image into a sequence. The serialization method is to divide the WSI into several 224x224 areas, insert a CLS Token in the middle of each area, and finally merge the tokens of all areas together to form a long sequence;

[0293] For a full-slice image I of size H×W×C, the Pixel-Mamba model architecture can tokenize I into visual tokens T = {τi,i∈[1,M]}, where H, W, and C represent the height, width, and number of channels of the image, respectively. Here τi represents the i-th token, and M represents the total number of visual tokens; the specific execution method of WSI full-image serialization is:

[0294] 1. Split the WSI into n 224x224 (hxw) regions;

[0295] 2. Take each region as a scanning window and use zigzag scanning to obtain a local sequence of length hxw+1, which contains all pixel tokens of each region (each pixel is used as a token separately);

[0296] This method uses a zigzag scan to scan the pixels in each area, so as to determine the Token sequence corresponding to each area. The zigzag scan can ensure that adjacent pixels are as close as possible in the sequence, thereby preserving the local spatial relationship. The specific zigzag scanning method can be:

[0297] Given the large height and width of WSI, a simple serialization strategy may cause adjacent tokens (e.g., tokens in the upper and lower rows) to be far apart in the sequence, which may disrupt the learning of inductive biases. To alleviate this problem, this method adopts a region-based zigzag scanning method, where the image I is divided into n regions of size h×w, where h and w are the height and width of each region, respectively, and each region acts as a scanning window.

[0298] 3. Insert a Class Token into the Token sequence corresponding to each of the n regions;

[0299] Specifically, within each scan window, a zigzag scan is applied and a CLS token is inserted at the center of the local sequence, resulting in a sequence length of (h×w+1) for each region. The zigzag scan is then extended to all regions, ensuring the spatial proximity of adjacent tokens. For an initial pixel-level token of I, the total sequence length is M=H×W+n.

[0300] 4. Connect the local sequences corresponding to the n regions to form a global long sequence.

[0301] It should be noted that, unlike previous methods, which usually segment images into fixed-size patches (e.g., 16×16), Pixel-Mamba initially tokenizes each pixel individually, generating 1×1 pixel-level tokens. When these tokens are processed in the Pixel-Mamba network, their receptive fields gradually expand, enabling the network to learn hierarchical representations. This is particularly important for capturing the multi-scale information inherent in pathological images.

[0302] Among them, for the pixel-mamba network structure:

[0303] The input of the pixel-mamba network structure is a pixel-level token sequence T generated by serializing the full slice image. Each token τ has an initial size of C = 3, corresponding to the RGB channels. As the token receptive field expands in the forward process, the dimension C gradually increases to accommodate richer representations.

[0304] The pixel-mamba network structure consists of 24 pixel-mamba layers. Each pixel-mamba layer includes three sub-modules: relationship learning sub-module (Mamba Block), region fusion sub-module (regionfusion), and token expansion sub-module (Token expansion).

[0305] The relational learning submodule (Mamba Block) is responsible for modeling the global context between tokens in a long sequence to capture long-distance dependencies. Specifically, it learns the long-distance dependencies between tokens in the sequence during the forward propagation and back-propagation processes, and then outputs a token sequence of the same length as the original sequence.

[0306] The region fusion submodule is responsible for fusing two similar image regions in each pixel-mamba layer; thus identifying similar image regions in each layer and merging them to reduce redundancy and improve memory efficiency.

[0307] The token expansion submodule is responsible for expanding the token receptive field of the token in the h and w dimensions of the image in the forward process. The token expansion submodule can expand the token receptive field, gradually increasing it from 1×1 to 32×32, which is crucial for learning hierarchical representations.

[0308] For the relationship learning submodule (Mamba Block), the specific implementation method is as follows:

[0309] 1. Input the above pixel-level token sequence T into the relational learning submodule (Mamba Block) of the first layer (pixel-mamba layer);

[0310] 2. In the Mamba Block, the forward state space model of the bidirectional state space model (SSM) is used to forward propagate the pixel-level token sequence T, and the dependency from the beginning to the end of the forward propagation of the pixel-level token sequence T is captured;

[0311] 3. In the Mamba Block, the reverse state space model of the bidirectional state space model (SSM) is used to perform reverse propagation processing on the pixel-level token sequence T after the forward propagation processing, and the back-to-front dependency relationship is captured from the reverse processing process from the end of the forward propagation to the start of the forward propagation for the pixel-level token sequence T;

[0312] 4. According to the front-to-back dependency and the back-to-front dependency, the original pixel-level token sequence T is adjusted, and a new token sequence T with the same length as the original pixel-level token sequence T is output.

[0313] It should be noted that Mamba Block uses a bidirectional state space model (SSM) to establish long-distance dependencies between multiple tokens in a sequence in each pixel-mambalayer layer. The forward propagation process of Mamba Block is formulated as the following formula (1):

[0314]

[0315] Among them, T in Formula 1 is the token sequence of Pixel-Mamba in the lth layer, fc represents the linear layer, and SSM f and SSM b denote the forward and backward branches of the bidirectional SSM, respectively, and ψ denotes the Conv1D operation. represents element-wise multiplication, SiLU is the SiLU activation function, and norm represents normalization.

[0316] For the region fusion submodule, the specific implementation is as follows:

[0317] Whole slice images (WSIs) typically contain hundreds of millions of pixels, which poses a significant computational challenge for model training. However, much of the information therein is redundant, which provides an opportunity to reduce complexity. To address this issue, Pixel-Mamba introduces a region fusion submodule that iteratively fuses similar regions layer by layer to reduce redundancy and computational overhead.

[0318] The specific execution steps are as follows:

[0319] 1. Divide the n regions into two equal region sets, namely the gallery set and the detection set, each set contains 2 / n regions;

[0320] Given n image regions containing M tokens, the region fusion submodule divides them into two equal sets: a gallery set and a detection set, each containing n / 2 regions.

[0321] 2. Calculate the cosine similarity between the CLS token of each region in the probe set and the CLS tokens of all regions in the gallery set;

[0322] For each region in the probe set, calculate the cosine similarity S between the CLS token of that region and the CLS tokens of all regions in the gallery set.ij :

[0323]

[0324] in, Represent the CLS tokens from the gallery collection and the probe collection, respectively.

[0325] 3. Select k region pairs with the highest similarity scores for fusion;

[0326] The k region pairs (Top-k) with the highest similarity scores are selected for fusion.

[0327] 4. Merge these k region pairs by averaging the corresponding tokens in the token sequences from each pair of regions;

[0328] For a certain image area in the detection set, the average value of the tokens at each position in the area with the highest similarity with the gallery set is calculated, and the average value of the tokens is used as the new token, so that the two tokens distributed in the two image areas are merged into one token, thereby merging the token sequences of the two image areas.

[0329] Specifically, these k pairs of regions are merged by averaging the corresponding tokens in the token sequences from each pair of regions. The number of merged regions k is defined as [α*n / L], where 0<α<1 is a hyperparameter that controls the proportion of regions retained in the final output, and L is the total number of layers in the network; after fusion, the number of remaining regions is reduced to nk.

[0330] 5. Output the number of fused regions M, each of which has a new token sequence.

[0331] For the token expansion submodule, the specific implementation is as follows:

[0332] Learning hierarchical token representations is crucial for effective pathology image analysis. Pixel-Mamba adopts a token expansion strategy to gradually expand the receptive field of tokens at key layers of the network. This enables subsequent Mamba blocks to capture multi-scale token representations at different levels.

[0333] It should be noted that there are two subtypes of token scaling: vertical scaling and horizontal scaling, which are applied alternately to each layer of the network.

[0334] The specific steps to perform token expansion are:

[0335] 1. Remove the CLS token from the token sequence of each of the n regions and reshape the other feature tokens into an h×w×C feature map.

[0336] Specifically, for a token sequence of a given image region, the CLS tokens are first removed, and then the remaining tokens are reshaped into a h×w×C feature map, where h and w are the height and width of the feature map, and C is the number of channels (i.e., the feature dimension of each token);

[0337] 2. Apply the shift-and-merge operation to expand the feature map vertically or horizontally to obtain a new feature map;

[0338] Specifically, for vertical expansion, the shift-merge method involves filling the first row;

[0339] First, a new row is filled for the feature map by copying the first row into a new row; the feature map is shifted down by one position to obtain a shifted feature map after average shift;

[0340] Secondly, the tokens in the shifted feature map are averaged (Avg) or concatenated (Cat) with the original tokens at the corresponding positions in the original feature map to generate a new feature map (h×w×C) with the shifted merged tokens.

[0341] Finally, in the case of vertical expansion, for a token in the new feature map, the token expansion operation increases the receptive field by concatenating or averaging adjacent tokens along the vertical dimension; the output size is h / 2×w×2C (if concatenated) or h / 2×w×C (if averaged).

[0342] It should be noted that horizontal token expansion follows a similar process, but shift-merge and token expansion are applied in the w dimension for the tokens in the new feature map; the output size is h×w / 2×2C or h×w / 2×C, depending on the operation (concatenation or averaging).

[0343] 3. Determine a new feature token based on the target feature map, add the CLS token removed in step 1 to the new feature token, and obtain the final token sequence for each region.

[0344] It should be noted that token expansion gradually expands the receptive field of the token as the network deepens, allowing the Mamba block to capture multi-scale representations at all levels. This process enables Pixel-Mamba to extract hierarchical representations from full-slice images. The final CLS tokens of the merged regions are averaged into one token for downstream tasks.

[0345] Among them, for the pathological task head:

[0346] After the input sequence is forward-calculated by pixel-mamba, a series of tokens will be output, and the CLSToken will be taken out and mean pooled as the representation of the series output. Finally, the token of mean pooling is input into a fully connected layer (task head) to complete the downstream task.

[0347] Both image classification and tumor staging are defined as classification tasks, processed by attaching a classification head and trained using cross entropy loss.

[0348] For survival analysis, the output CLS tokens of Pixel-Mamba are used to predict the risk function h(t), which is then used to calculate the survival function F(t). The negative log-likelihood loss L is defined as:

[0349]

[0350] Among them, (I, t, c) represents the full-slice image, survival time, and right-censoring status of the patients in the dataset. represents training data; log refers to the logarithmic function.

[0351] Based on the above steps, it can be seen that the mamba architecture used in this method to build the basic model of pathological images can effectively reduce video memory, and mamba is also good at capturing long sequence relationships. The region fusion module designed in this method can effectively fuse two similar regions, further reducing the amount of calculation. In addition, the Token expansion model can gradually expand the receptive field of the Token representation during the forward operation of the network, and at the same time, with the help of the mamba block, the global contextual relationship can be captured at each layer, so that the pixel-mamba model learns multi-level pathological representations, which is very beneficial for the task of full pathological image analysis.

[0352] In summary, the proposed method surpasses existing models in natural image classification tasks, and surpasses existing optimal solutions that use a large amount of pathology data for pre-training in pathology-related tasks such as tumor staging and prognosis without using pathology images for pre-training.

[0353] The development of AI pathology analysis algorithms can effectively improve the intelligence level of hospital pathology departments, and AI pathology algorithms can also provide important decision-making basis for the treatment of cancer patients. Therefore, the development of AI pathology algorithms can enhance the laboratory's ability in medical analysis and may provide detection interface services for precision medicine in the future.

[0354] Based on this, this method proposes Pixel-Mamba, a highly memory-efficient architecture designed for end-to-end WSI analysis; Pixel-Mamba introduces progressive tag expansion to incorporate hierarchical local inductive biases while modeling long-range dependencies at all scales, thereby achieving efficient modeling of gigapixel WSIs. Extensive experiments show that Pixel-Mamba achieves or exceeds the performance of base models pre-trained on millions of WSIs or WSI-text pairs without any pathology-specific pre-training.

[0355] This method solves the problem of modeling and training end-to-end pathology image analysis with efficient multi-modal input under billion-pixel pathology images, and achieves a significant improvement in the performance of multiple tasks of pathology analysis. The technical content of this method includes: 1) A lightweight end-to-end training model for pathology images based on the Pixel-mamba architecture is proposed; 2) The Pixel-mamba model is a hierarchical representation model that can effectively model multi-scale pathology images; 3) The pixel-mamba architecture is superior to the two-stage SOTA solution based on the pre-trained basic model.

[0356] It should be noted that the image processing method provided in this specification can be applied to the scenario of tumor staging based on pathological images to perform tumor staging tasks based on pathological images;

[0357] For tumor staging, this method uses Pixel-Mamba as the backbone network and adds a linear layer as the staging head (i.e., task head), which projects the features into the probabilities of the four stages. This staging model is named Pixel-Mamba-Stage. Three datasets from The Cancer Genome Atlas (TCGA) were used: bladder urothelial carcinoma (BLCA), breast invasive carcinoma (BRCA), and lung adenocarcinoma (LUAD). Each patient was divided into one of the four tumor stages. And, due to the imbalanced number of patients in each stage, this method was evaluated using a macro f1 score with 5-fold cross validation.

[0358] Tumor staging results for performing tumor staging tasks:

[0359] Compared with different feature extraction models such as ResNet-50, Gi-gaViT and CONCH, Pixel-Mamba-Stage achieves macro F1 scores of 0.5334, 0.3744 and 0.3917 on BLCA, BRCA and LUAD datasets, respectively, which is better than the two-stage MIL method, the two-stage hierarchical representation method and LongViT. In addition, all two-stage methods are pre-trained on large-scale pathological images or image-caption pairs, while LongViT is pre-trained on 10k wsi. In contrast, Pixel-Mamba-stage is only pre-trained on natural images, demonstrating the effectiveness of Pixel-Mamba.

[0360] The image processing method provided in this specification can also be applied to the scenario of survival analysis based on pathological images to perform survival analysis tasks based on pathological images;

[0361] For survival analysis, this method uses Pixel-Mamba as the backbone network and adds a linear layer as the survival head (i.e., task head) to predict the risk of patients. This model is named Pixel-Mamba-Surv. The datasets used for survival analysis are BLCA, BRCA, and LUAD. In addition, this method is evaluated using the consistency index (C-index) of 5-fold cross-validation.

[0362] Survival analysis results for performing the survival analysis task:

[0363] Similar to tumor staging, Pixel-Mamba-Surv, pre-trained on natural images, achieved C-index values ​​of 0.6507, 0.6707, and 0.6468 on the BLCA, BRCA, and LUAD datasets, respectively, outperforming all two-stage MIL methods, two-stage hierarchical methods, and end-to-end LongViT. These results show that learning slide representations in an end-to-end manner can outperform existing two-stage methods even without pre-training on pathological images. In addition, the performance of four different methods (ILRA-MIL, HIPT, LongViT, and Pixel-Mamba-Surv) was analyzed, and Pixel-Mamba-Surv was better at distinguishing between low-risk and high-risk groups compared to the other methods, and its lower p-value (0.0013) indicated that the prediction results were statistically significant.

[0364] See also Figure 5 , Figure 5 A flow chart of a pathological image processing method provided according to an embodiment of this specification is shown, which specifically includes the following steps:

[0365] Step 502: determining a plurality of sub-images of the pathological image, performing encoding processing on each sub-image, and obtaining a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image, and an image information code inserted into each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image;

[0366] Step 504: input the target image coding sequence into the coding processing sub-model of the image processing model, and use the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence to obtain a pathological image coding sequence;

[0367] Step 506: Determine the image processing result corresponding to the pathological image according to the pathological image coding sequence.

[0368] The pathological image processing method in one or more embodiments of the present specification, during the process of processing the pathological image, will encode multiple sub-images of the pathological image, so as to obtain a target image coding sequence of the pathological image, the target image coding sequence includes an initial image coding sequence corresponding to each sub-image, and an image information code inserted in each initial image coding sequence, each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence can characterize the local features of the corresponding sub-image; thereby, the features of each sub-image of the pathological image can be accurately represented from multiple aspects through the initial image coding sequence and the image information coding.

[0369] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, and obtain the pathological image coding sequence. Finally, based on the pathological image coding sequence that can accurately express the image features of the pathological image and considers the correlation between the codings, the image processing result corresponding to the pathological image is accurately determined; the problem of inaccurate image processing results is avoided, and the accuracy of the image processing results is improved.

[0370] The above is a schematic scheme of a pathological image processing method of this embodiment. It should be noted that the technical scheme of the pathological image processing method and the technical scheme of the above-mentioned image processing method belong to the same concept, and the details not described in detail in the technical scheme of the pathological image processing method can be referred to the description of the technical scheme of the above-mentioned image processing method.

[0371] See also Figure 6 , Figure 6A flowchart of an image processing model training method provided according to an embodiment of the present specification is shown, which specifically includes the following steps:

[0372] Step 602: determining a plurality of sub-images of the sample object image, performing encoding processing on each sub-image, and obtaining a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image, and an image information code inserted into each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image;

[0373] Step 604: input the target image coding sequence into the coding processing sub-model of the image processing model, and use the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence to obtain a sample object image coding sequence;

[0374] Step 606: determining a sample image processing result corresponding to the sample object image according to the sample object image coding sequence;

[0375] Step 608: Determine the sample label corresponding to the sample object image, and use the sample label and the sample image processing result to perform model training on the encoding processing sub-model of the image processing model to obtain the trained encoding processing sub-model.

[0376] Among them, the sample object image can be understood as the target object image as a sample; the sample label corresponding to the sample object image can be understood as the label data used for model training. For example, the sample label can be the real classification result corresponding to classification tasks such as image classification and tumor staging; or the real survival analysis result corresponding to the survival analysis task.

[0377] In the image processing model training method in one or more embodiments of the present specification, during the model training process, multiple sub-images of the sample object image are encoded to obtain a target image coding sequence of the sample object image, wherein the target image coding sequence includes an initial image coding sequence corresponding to each sub-image and an image information code inserted in each initial image coding sequence, wherein each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence can characterize the local features of the corresponding sub-image; thereby, the features of each sub-image of the sample object image can be accurately represented from multiple aspects through the initial image coding sequence and the image information code.

[0378] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, obtain the sample object image coding sequence that can accurately express the image features of the sample object image and take into account the correlation between the codings, and determine the sample image processing result based on the sample object image coding sequence.

[0379] The sample labels and sample image processing results are used to perform efficient model training on the encoding processing sub-model of the image processing model, thereby obtaining the encoding processing sub-model that can determine accurate image processing results.

[0380] The above is a schematic scheme of an image processing model training method of this embodiment. It should be noted that the technical scheme of the image processing model training method and the technical scheme of the above-mentioned image processing method belong to the same concept, and the details not described in detail in the technical scheme of the image processing model training method can be referred to the description of the technical scheme of the above-mentioned image processing method.

[0381] See also Figure 7 , Figure 7 A flowchart of a cancer computer-aided diagnosis method provided according to an embodiment of this specification is shown, which specifically includes the following steps:

[0382] Step 702: determining a plurality of sub-images of the tumor image, performing encoding processing on each sub-image, and obtaining a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image, and an image information code inserted into each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image;

[0383] Step 704: input the target image coding sequence into the coding processing sub-model of the image processing model, and use the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence to obtain a tumor image coding sequence;

[0384] Step 706: Determine a tumor image processing result corresponding to the tumor image according to the tumor image coding sequence.

[0385] In the computer-aided cancer diagnosis method in one or more embodiments of the present specification, in the process of implementing computer-aided cancer diagnosis, multiple sub-images of a tumor image are encoded to obtain a target image coding sequence of the tumor image, wherein the target image coding sequence includes an initial image coding sequence corresponding to each sub-image and an image information code inserted in each initial image coding sequence, wherein each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence can characterize the local features of the corresponding sub-image; thereby, the features of each sub-image of the tumor image can be accurately represented from multiple aspects through the initial image coding sequence and the image information code.

[0386] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, and obtain the tumor image coding sequence. Finally, according to the tumor image coding sequence that can accurately express the image characteristics of the tumor image and considers the correlation between the codings, the image processing result corresponding to the tumor image is accurately determined; the problem of inaccurate image processing results is avoided, and the accuracy of the image processing results is improved.

[0387] The above is a schematic scheme of a cancer computer-aided diagnosis method of this embodiment. It should be noted that the technical scheme of the cancer computer-aided diagnosis method and the technical scheme of the above-mentioned image processing method belong to the same concept, and the details not described in detail in the technical scheme of the cancer computer-aided diagnosis method can be referred to the description of the technical scheme of the above-mentioned image processing method.

[0388] See also Figure 8 , Figure 8 A flowchart of an image processing method provided according to an embodiment of the present specification is shown. The image processing method is applied to a client of a medical system and specifically includes the following steps:

[0389] Step 802: In response to a user's selection operation on the display interface of the client, determining a medical image to be processed;

[0390] Step 804: Send the medical image to be processed to the server of the medical system, and receive the image processing result corresponding to the medical image to be processed returned by the server, wherein the image processing result is determined according to a medical image coding sequence, and the medical image coding sequence is obtained by learning the coding relationship of an initial image coding sequence corresponding to each sub-image included in a target image coding sequence and an image information coding inserted in each initial image coding sequence using a coding processing sub-model of an image processing model, and the image information coding is used to characterize the global features of the corresponding sub-image, and the target image coding sequence is obtained by coding multiple sub-images of the medical image to be processed.

[0391] In one or more embodiments of the present specification, the image processing method applied to the client of the medical system may utilize the client to send the medical image to be processed to the server of the medical system during the process of processing the medical image to be processed.

[0392] The server will encode multiple sub-images of the medical image to be processed, so as to obtain a target image coding sequence of the medical image to be processed, wherein the target image coding sequence includes an initial image coding sequence corresponding to each sub-image, and an image information code inserted in each initial image coding sequence, wherein each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence can characterize the local features of the corresponding sub-image; thus, the features of each sub-image of the medical image to be processed can be accurately represented from multiple aspects through the initial image coding sequence and the image information code.

[0393] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, and obtain the coding sequence of the medical image to be processed. Finally, according to the coding sequence of the medical image to be processed that can accurately express the image features of the medical image to be processed and takes into account the correlation between the codings, the image processing result corresponding to the medical image to be processed is accurately determined.

[0394] Finally, the server will send the image processing results to the client of the medical system, thereby avoiding the problem of inaccurate image processing results and improving the accuracy of image processing results.

[0395] The above is a schematic scheme corresponding to the image processing method applied to the client of the medical system in this embodiment. It should be noted that the technical scheme of the image processing method and the technical scheme of the above-mentioned image processing method belong to the same concept, and the details not described in detail in the technical scheme of the image processing method applied to the client of the medical system can be referred to the description of the technical scheme of the above-mentioned image processing method.

[0396] See also Fig. 9 , Fig. 9 A flowchart of an image processing method provided according to an embodiment of the present specification is shown. The image processing method is applied to a cloud-side device and specifically includes the following steps:

[0397] Step 902: receiving multiple sub-images of a target object image sent by a terminal-side device, encoding each sub-image to obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted into each initial image encoding sequence, and each image information code is used to represent a global feature of a corresponding sub-image;

[0398] Step 904: input the target image coding sequence into the coding processing sub-model of the image processing model, and use the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence to obtain the target object image coding sequence;

[0399] Step 906: determining an image processing result corresponding to the target object image according to the target object image coding sequence;

[0400] Step 908: Send the image processing result to the terminal device.

[0401] In one or more embodiments provided in this specification, the cloud-side device may be a central cloud device of a distributed architecture or an edge cloud device of a distributed architecture. The cloud-side device may be a cloud-side device with a cloud desktop system or cloud desktop software installed and deployed, such as a cloud server, cloud host, etc. The terminal-side device may be understood as any terminal that interacts with the cloud-side device for data exchange. The terminal may be a laptop, desktop computer, tablet computer, smart device, server, etc.

[0402] In one or more embodiments of the present specification, an image processing method applied to a cloud-side device receives multiple sub-images of the target object image sent by an end-side device during processing of the target object image, and encodes the multiple sub-images of the target object image to obtain a target image coding sequence of the target object image. The target image coding sequence includes an initial image coding sequence corresponding to each sub-image, and an image information code inserted in each initial image coding sequence. Each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence can characterize the local features of the corresponding sub-image; thereby, the initial image coding sequence and the image information code can accurately represent the features of each sub-image of the target object image from multiple aspects.

[0403] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, and obtain the target object image coding sequence. Finally, according to the target object image coding sequence that can accurately express the image features of the target object image and considers the correlation between the codings, the image processing result corresponding to the target object image is accurately determined, and the image processing result is sent to the terminal side device, thereby avoiding the problem of inaccurate image processing results and improving the accuracy of the image processing results.

[0404] The above is a schematic scheme of the image processing method applied to the cloud-side device in this embodiment. It should be noted that the technical scheme of the image processing method applied to the cloud-side device and the technical scheme of the above-mentioned image processing method belong to the same concept, and the details of the technical scheme of the image processing method applied to the cloud-side device that are not described in detail can all be referred to the description of the technical scheme of the above-mentioned image processing method.

[0405] Corresponding to the above method embodiment, this specification also provides an image processing device embodiment, Fig.10 FIG. 2 shows a schematic diagram of the structure of an image processing device provided by an embodiment of the present specification. Fig.10 As shown, the device comprises:

[0406] The first encoding processing module 1002 is configured to determine a plurality of sub-images of the target object image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted in each initial image encoding sequence, and each image information code is used to represent a global feature of a corresponding sub-image;

[0407] The second coding processing module 1004 is configured to input the target image coding sequence into the coding processing sub-model of the image processing model, and use the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence to obtain the target object image coding sequence;

[0408] The result determination module 1006 is configured to determine the image processing result corresponding to the target object image according to the target object image coding sequence.

[0409] Optionally, the second encoding processing module 1004 is further configured to:

[0410] Inputting the target image coding sequence into the coding processing sub-model of the image processing model, wherein the coding processing sub-model includes a relation learning network layer and a coding fusion network layer;

[0411] Using the relationship learning network layer, the coding relationship learning is performed on the initial image coding sequences and the image information coding inserted into the initial image coding sequences to obtain a plurality of image coding sequences to be fused;

[0412] The coding fusion network layer is used to perform coding fusion on the plurality of image coding sequences to be fused, so as to obtain the target object image coding sequence.

[0413] Optionally, the encoding processing sub-model further includes an encoding extension network layer;

[0414] The second encoding processing module 1004 is further configured to:

[0415] Using the coding fusion network layer, coding and fusing the plurality of image coding sequences to be fused to obtain a fused image coding sequence;

[0416] Using the coding extension network layer, coding and extending the fused image coding sequence to obtain a plurality of extended image coding sequences;

[0417] Based on the multiple extended image coding sequences, the target object image coding sequence is determined.

[0418] Optionally, the encoding relationship is a long-distance dependency relationship;

[0419] The second encoding processing module 1004 is further configured to:

[0420] Inputting the initial image coding sequences into the relational learning network layer, and determining a coding sequence to be processed from the initial image coding sequences in the relational learning network layer, wherein the coding sequence to be processed is any one of the initial image coding sequences;

[0421] Determine the long-distance dependency relationship between multiple image codes in the code sequence to be processed by performing forward propagation and backward propagation on the code sequence to be processed, wherein the multiple image codes are at least one initial image code and the image information code contained in the code sequence to be processed;

[0422] Based on the long-distance dependency, encoding adjustment is performed on each of the initial image encoding sequences to obtain the multiple image encoding sequences to be fused.

[0423] Optionally, the second encoding processing module 1004 is further configured to:

[0424] Inputting the plurality of image coding sequences to be fused into the coding fusion network layer;

[0425] In the coding fusion network layer, the multiple image coding sequences to be fused are divided into a first coding sequence set and a second coding sequence set, wherein the first coding sequence set includes a first image coding sequence to be fused, the second coding sequence set includes a second image coding sequence to be fused, and the number of sequences of the first image coding sequence to be fused is the same as the number of sequences of the second image coding sequence to be fused;

[0426] Determining a first image information code inserted in the first image coding sequence to be fused, and determining a second image information code inserted in the second image coding sequence to be fused;

[0427] Determine a coding similarity between a target first image information coding and a target second image information coding, wherein the target first image information coding is any one of the first image information codings, and the target second image information coding is any one of the second image information codings;

[0428] Based on the coding similarity, the first coding sequence set and the second coding sequence set are fused to obtain the fused image coding sequence.

[0429] Optionally, the second encoding processing module 1004 is further configured to:

[0430] Determine a target first coding sequence from the first coding sequence set, and determine a target second coding sequence from the second coding sequence set, wherein the target first coding sequence is the first image coding sequence to be fused inserted into the target first image information coding, and the target second coding sequence is the second image coding sequence to be fused inserted into the target second image information coding;

[0431] Determine the target first coding sequence and the target second coding sequence as a coding sequence pair, and determine a similar coding sequence pair from the coding sequence pairs according to the coding similarity;

[0432] Determine, from the similar coding sequences, a plurality of first image codes to be fused and the target first image information coding in the target first coding sequence, and determine a plurality of second image codes to be fused and the target second image information coding in the target second coding sequence;

[0433] According to the sequence position information of each first image code to be fused and the sequence position information of each second image code to be fused, determining corresponding associated codes for each first image code to be fused from the second image code to be fused;

[0434] respectively calculating average values ​​between the first image codes to be fused and the corresponding associated codes, and obtaining multiple fused image codes based on multiple average values;

[0435] The target first image information code and the target second image information code are encoded to obtain a fused image information code, and the fused image code sequence is constructed based on the fused image information code and the multiple fused image codes.

[0436] Optionally, the second encoding processing module 1004 is further configured to:

[0437] Inputting the fused image coding sequence into the coding extension network layer, and generating a coding feature map in the coding extension network layer according to a plurality of fused image codes in the fused image coding sequence;

[0438] Determining a target coding feature from a plurality of coding features included in the coding feature map, and expanding the coding feature map according to the target coding feature to obtain an expanded coding feature map;

[0439] Performing a shift operation on a plurality of coding features included in the extended coding feature map to obtain a shift coding feature map, and performing coding expansion based on the extended coding feature map and the shift coding feature map to obtain a plurality of extended codes;

[0440] An extended image coding sequence is generated based on the fused image information coding in the fused image coding sequence and the multiple extended codings.

[0441] Optionally, the second encoding processing module 1004 is further configured to:

[0442] Determining a plurality of extended coding features in the extended coding feature map, and determining a plurality of displacement coding features in the displacement coding feature map;

[0443] Determining a corresponding associated displacement coding feature for each of the extended coding features from the plurality of displacement coding features according to the extended coding feature map position of each of the extended coding features and the displacement coding feature map position of each of the displacement coding features;

[0444] Combining the target extended coding feature and the corresponding associated displacement coding feature to obtain a plurality of coding feature pairs, wherein the target extended coding feature is any one of the plurality of extended coding features;

[0445] Merging the expanded coding features and the displacement coding features contained in each coding feature pair respectively to obtain a plurality of merged coding features, and determining a merged coding feature map based on the plurality of merged coding features;

[0446] Two adjacent merged coding features in the merged coding feature map are subjected to feature fusion to obtain a coding extension feature map, and the coding extension feature map is encoded to obtain multiple extension codes.

[0447] Optionally, the result determination module 1006 is further configured to:

[0448] The target object image coding sequence is processed by utilizing the task execution unit in the image processing model to obtain an image processing result corresponding to the target object image, wherein the task execution unit is associated with the target image processing task.

[0449] Optionally, the task execution unit is a task execution network layer;

[0450] The result determination module 1006 is further configured to:

[0451] Determining a plurality of target image information codes from the target object image code sequence, and performing average pooling processing on the plurality of target image information codes using an average pooling network layer in the image processing model to obtain an average image information code;

[0452] The task execution network layer in the image processing model is utilized to execute the target image processing task for the target object image according to the average image information encoding to obtain the image processing result.

[0453] Optionally, the first encoding processing module 1002 is further configured to:

[0454] The image processing unit of the image processing model is used to determine the multiple sub-images from the target object image, and perform encoding processing on each sub-image to obtain a target image encoding sequence.

[0455] Optionally, the first encoding processing module 1002 is further configured to:

[0456] In the image processing unit, dividing the target object image into the plurality of sub-images;

[0457] Performing encoding processing on pixels contained in each sub-image respectively to obtain initial image codes corresponding to the pixels, and using the initial image codes to obtain initial image code sequences corresponding to the sub-images, wherein the initial image codes are used to characterize local features of the corresponding sub-images;

[0458] The image information code is inserted into each initial image code sequence respectively, and the target image code sequence is obtained by using a plurality of initial image code sequences into which the image information code is inserted.

[0459] The image processing device in one or more embodiments of the present specification, during the process of processing the target object image, will encode multiple sub-images of the target object image, so as to obtain a target image coding sequence of the target object image, the target image coding sequence includes an initial image coding sequence corresponding to each sub-image, and an image information code inserted in each initial image coding sequence, each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence can characterize the local features of the corresponding sub-image; thereby, the features of each sub-image of the target object image can be accurately represented from multiple aspects through the initial image coding sequence and the image information coding.

[0460] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, and obtain the target object image coding sequence. Finally, according to the target object image coding sequence that can accurately express the image features of the target object image and takes into account the correlation between the codings, the image processing result corresponding to the target object image is accurately determined; the problem of inaccurate image processing results is avoided, and the accuracy of the image processing results is improved.

[0461] The above is a schematic scheme of an image processing device of this embodiment. It should be noted that the technical scheme of the image processing device and the technical scheme of the above-mentioned image processing method belong to the same concept, and the details not described in detail in the technical scheme of the image processing device can be referred to the description of the technical scheme of the above-mentioned image processing method.

[0462] Corresponding to the above method embodiment, this specification also provides a pathological image processing device embodiment, the device comprising:

[0463] A first encoding processing module is configured to determine a plurality of sub-images of a pathological image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted into each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image;

[0464] A second coding processing module is configured to input the target image coding sequence into a coding processing sub-model of an image processing model, and use the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence to obtain a pathological image coding sequence;

[0465] The result determination module is configured to determine the image processing result corresponding to the pathological image according to the pathological image coding sequence.

[0466] The pathological image processing device in one or more embodiments of the present specification may encode multiple sub-images of the pathological image during the process of processing the pathological image, thereby obtaining a target image coding sequence of the pathological image, wherein the target image coding sequence includes an initial image coding sequence corresponding to each sub-image, and an image information code inserted in each initial image coding sequence, wherein each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence may characterize the local features of the corresponding sub-image; thereby, the features of each sub-image of the pathological image may be accurately represented from multiple aspects through the initial image coding sequence and the image information code.

[0467] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, and obtain the pathological image coding sequence. Finally, based on the pathological image coding sequence that can accurately express the image features of the pathological image and considers the correlation between the codings, the image processing result corresponding to the pathological image is accurately determined; the problem of inaccurate image processing results is avoided, and the accuracy of the image processing results is improved.

[0468] The above is a schematic scheme of a pathological image processing device of this embodiment. It should be noted that the technical scheme of the pathological image processing device and the technical scheme of the pathological image processing method described above belong to the same concept, and the details not described in detail in the technical scheme of the pathological image processing device can be referred to the description of the technical scheme of the pathological image processing method described above.

[0469] Corresponding to the above method embodiment, this specification also provides an image processing model training device embodiment, the device comprising:

[0470] A first encoding processing module is configured to determine a plurality of sub-images of a sample object image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted into each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image;

[0471] A second encoding processing module is configured to input the target image encoding sequence into an encoding processing sub-model of an image processing model, and use the encoding processing sub-model to learn the encoding relationship of each initial image encoding sequence and the image information encoding inserted in each initial image encoding sequence to obtain a sample object image encoding sequence;

[0472] A result determination module, configured to determine a sample image processing result corresponding to the sample object image according to the sample object image coding sequence;

[0473] The model training module is configured to determine the sample label corresponding to the sample object image, and use the sample label and the sample image processing result to perform model training on the encoding processing sub-model of the image processing model to obtain the trained encoding processing sub-model.

[0474] In one or more embodiments of the present specification, the image processing model training device will encode multiple sub-images of the sample object image during the model training process, so as to obtain a target image coding sequence of the sample object image, wherein the target image coding sequence includes an initial image coding sequence corresponding to each sub-image, and an image information code inserted in each initial image coding sequence, wherein each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence can characterize the local features of the corresponding sub-image; thereby, the features of each sub-image of the sample object image can be accurately represented from multiple aspects through the initial image coding sequence and the image information coding.

[0475] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, obtain the sample object image coding sequence that can accurately express the image features of the sample object image and take into account the correlation between the codings, and determine the sample image processing result based on the sample object image coding sequence.

[0476] The sample labels and sample image processing results are used to perform efficient model training on the encoding processing sub-model of the image processing model, thereby obtaining the encoding processing sub-model that can determine accurate image processing results.

[0477] The above is a schematic scheme of an image processing model training device of this embodiment. It should be noted that the technical scheme of the image processing model training device and the technical scheme of the above-mentioned image processing model training method belong to the same concept, and the details not described in detail in the technical scheme of the image processing model training device can be found in the description of the technical scheme of the above-mentioned image processing model training method.

[0478] Corresponding to the above method embodiment, this specification also provides an embodiment of a computer-aided cancer diagnosis device, the device comprising:

[0479] A first encoding processing module is configured to determine a plurality of sub-images of a tumor image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted into each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image;

[0480] A second encoding processing module is configured to input the target image encoding sequence into an encoding processing sub-model of an image processing model, and use the encoding processing sub-model to learn the encoding relationship of each initial image encoding sequence and the image information encoding inserted in each initial image encoding sequence to obtain a tumor image encoding sequence;

[0481] The result determination module is configured to determine the tumor image processing result corresponding to the tumor image according to the tumor image coding sequence.

[0482] The computer-aided cancer diagnosis device in one or more embodiments of the present specification, in the process of implementing computer-aided cancer diagnosis, will encode multiple sub-images of a tumor image to obtain a target image coding sequence of the tumor image, wherein the target image coding sequence includes an initial image coding sequence corresponding to each sub-image, and an image information code inserted in each initial image coding sequence, wherein each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence can characterize the local features of the corresponding sub-image; thereby, the features of each sub-image of the tumor image can be accurately represented from multiple aspects through the initial image coding sequence and the image information coding.

[0483] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, and obtain the tumor image coding sequence. Finally, according to the tumor image coding sequence that can accurately express the image characteristics of the tumor image and considers the correlation between the codings, the image processing result corresponding to the tumor image is accurately determined; the problem of inaccurate image processing results is avoided, and the accuracy of the image processing results is improved.

[0484] The above is a schematic scheme of a cancer computer-aided diagnosis device of this embodiment. It should be noted that the technical scheme of the cancer computer-aided diagnosis device and the technical scheme of the cancer computer-aided diagnosis method described above belong to the same concept, and the details not described in detail in the technical scheme of the cancer computer-aided diagnosis device can be referred to the description of the technical scheme of the cancer computer-aided diagnosis method described above.

[0485] Corresponding to the above method embodiment, this specification also provides an image processing device embodiment, which is applied to a client of a medical system, and includes:

[0486] An image determination module, configured to determine a medical image to be processed in response to a user's selection operation on the display interface of the client;

[0487] The result receiving module is configured to send the medical image to be processed to the server of the medical system, and receive the image processing result corresponding to the medical image to be processed returned by the server, wherein the image processing result is determined according to the medical image coding sequence, and the medical image coding sequence is obtained by learning the coding relationship of the initial image coding sequence corresponding to each sub-image included in the target image coding sequence and the image information coding inserted in each initial image coding sequence using the coding processing sub-model of the image processing model, the image information coding is used to characterize the global features of the corresponding sub-image, and the target image coding sequence is obtained by encoding multiple sub-images of the medical image to be processed.

[0488] In one or more embodiments of the present specification, the image processing device applied to the client of the medical system may use the client to send the medical image to be processed to the server of the medical system during the process of processing the medical image to be processed.

[0489] The server will encode multiple sub-images of the medical image to be processed, so as to obtain a target image coding sequence of the medical image to be processed, wherein the target image coding sequence includes an initial image coding sequence corresponding to each sub-image, and an image information code inserted in each initial image coding sequence, wherein each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence can characterize the local features of the corresponding sub-image; thus, the features of each sub-image of the medical image to be processed can be accurately represented from multiple aspects through the initial image coding sequence and the image information code.

[0490] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, and obtain the coding sequence of the medical image to be processed. Finally, according to the coding sequence of the medical image to be processed that can accurately express the image features of the medical image to be processed and takes into account the correlation between the codings, the image processing result corresponding to the medical image to be processed is accurately determined.

[0491] Finally, the server will send the image processing results to the client of the medical system, thereby avoiding the problem of inaccurate image processing results and improving the accuracy of image processing results.

[0492] The above is a schematic scheme of an image processing device applied to a client of a medical system in this embodiment. It should be noted that the technical scheme of the image processing device applied to a client of a medical system and the technical scheme of the image processing method applied to a client of a medical system belong to the same concept, and the details not described in detail in the technical scheme of the image processing device applied to a client of a medical system can be referred to the description of the technical scheme of the image processing method applied to a client of a medical system.

[0493] Corresponding to the above method embodiment, this specification also provides an image processing device embodiment, which is applied to a cloud-side device, and includes:

[0494] A first encoding processing module is configured to receive multiple sub-images of a target object image sent by a terminal-side device, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted in each initial image encoding sequence, and each image information code is used to represent a global feature of a corresponding sub-image;

[0495] A second encoding processing module is configured to input the target image encoding sequence into an encoding processing sub-model of an image processing model, and use the encoding processing sub-model to learn the encoding relationship of each initial image encoding sequence and the image information encoding inserted in each initial image encoding sequence to obtain a target object image encoding sequence;

[0496] A result determination module is configured to determine an image processing result corresponding to the target object image according to the target object image coding sequence;

[0497] The result sending module is configured to send the image processing result to the terminal device.

[0498] In one or more embodiments of the present specification, an image processing device applied to a cloud-side device will, during the process of processing a target object image, receive multiple sub-images of the target object image sent by an end-side device, and encode the multiple sub-images of the target object image to obtain a target image coding sequence of the target object image. The target image coding sequence includes an initial image coding sequence corresponding to each sub-image, and an image information code inserted in each initial image coding sequence. Each image information code is used to characterize the global features of the corresponding sub-image, and the initial image code in the initial image coding sequence can characterize the local features of the corresponding sub-image; thereby, the initial image coding sequence and the image information code can accurately represent the features of each sub-image of the target object image from multiple aspects.

[0499] Then, the coding processing sub-model of the image processing model is used to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, so as to determine the correlation between the initial image coding and the image information coding in each initial image coding sequence, and obtain the target object image coding sequence. Finally, according to the target object image coding sequence that can accurately express the image features of the target object image and considers the correlation between the codings, the image processing result corresponding to the target object image is accurately determined, and the image processing result is sent to the terminal side device, thereby avoiding the problem of inaccurate image processing results and improving the accuracy of the image processing results.

[0500] The above is a schematic scheme of an image processing device applied to a cloud-side device in this embodiment. It should be noted that the technical scheme of the image processing device applied to the cloud-side device and the technical scheme of the image processing method applied to the cloud-side device belong to the same concept, and the details not described in detail in the technical scheme of the image processing device applied to the cloud-side device can be referred to the description of the technical scheme of the image processing method applied to the cloud-side device.

[0501] Fig.11 The block diagram of a computing device 1100 according to one embodiment of the present specification is shown. The components of the computing device 1100 include but are not limited to a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.

[0502] The computing device 1100 also includes an access device 1140 that enables the computing device 1100 to communicate via one or more networks 1160. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1140 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).

[0503] In one embodiment of the present specification, the above components of the computing device 1100 and Fig.11 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Fig.11 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0504] The computing device 1100 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1100 may also be a mobile or stationary server.

[0505] The processor 1120 is used to execute the following computer executable instructions, which implement the steps of any of the above methods when executed by the processor.

[0506] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computing device embodiment, since it is basically similar to any of the above method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of any of the above method embodiments.

[0507] An embodiment of the present specification further provides a computer-readable storage medium storing a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.

[0508] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computer-readable storage medium embodiment, since it is basically similar to any of the above method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of any of the above method embodiments.

[0509] An embodiment of the present specification further provides a computer program product, including a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.

[0510] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of the computer program product and the technical solution of any of the above methods belong to the same concept, and the details not described in detail in the technical solution of the computer program product can be referred to the description of the technical solution of any of the above methods.

[0511] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0512] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0513] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0514] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0515] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. An image processing method, comprising: Determine multiple sub-images of the target object image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted in each initial image encoding sequence, and each image information code is used to represent a global feature of a corresponding sub-image; Inputting the target image coding sequence into the coding processing sub-model of the image processing model, using the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, to obtain the target object image coding sequence; An image processing result corresponding to the target object image is determined according to the target object image coding sequence.

2. The image processing method according to claim 1, wherein the step of inputting the target image coding sequence into a coding processing sub-model of an image processing model, using the coding processing sub-model to learn coding relationships between the initial image coding sequences and the image information coding inserted into the initial image coding sequences, and obtaining the target object image coding sequence comprises: Inputting the target image coding sequence into the coding processing sub-model of the image processing model, wherein the coding processing sub-model includes a relation learning network layer and a coding fusion network layer; Using the relationship learning network layer, the coding relationship learning is performed on the initial image coding sequences and the image information coding inserted into the initial image coding sequences to obtain a plurality of image coding sequences to be fused; The coding fusion network layer is used to perform coding fusion on the plurality of image coding sequences to be fused, so as to obtain the target object image coding sequence.

3. The image processing method according to claim 2, wherein the encoding processing sub-model further comprises an encoding extension network layer; The step of utilizing the coding fusion network layer to perform coding fusion on the plurality of image coding sequences to be fused to obtain the target object image coding sequence includes: Using the coding fusion network layer, coding and fusing the plurality of image coding sequences to be fused to obtain a fused image coding sequence; Using the coding extension network layer, coding and extending the fused image coding sequence to obtain a plurality of extended image coding sequences; Based on the multiple extended image coding sequences, the target object image coding sequence is determined.

4. The image processing method according to claim 2, wherein the encoding relationship is a long-distance dependency relationship; The method of using the relation learning network layer to perform coding relation learning on the initial image coding sequences and the image information coding inserted into the initial image coding sequences to obtain a plurality of image coding sequences to be fused includes: Inputting the initial image coding sequences into the relational learning network layer, and determining a coding sequence to be processed from the initial image coding sequences in the relational learning network layer, wherein the coding sequence to be processed is any one of the initial image coding sequences; Determine the long-distance dependency relationship between multiple image codes in the code sequence to be processed by performing forward propagation and backward propagation on the code sequence to be processed, wherein the multiple image codes are at least one initial image code and the image information code contained in the code sequence to be processed; Based on the long-distance dependency, encoding adjustment is performed on each of the initial image encoding sequences to obtain the multiple image encoding sequences to be fused.

5. The image processing method according to claim 3, wherein the step of using the coding fusion network layer to perform coding fusion on the plurality of image coding sequences to be fused to obtain a fused image coding sequence comprises: Inputting the plurality of image coding sequences to be fused into the coding fusion network layer; In the coding fusion network layer, the multiple image coding sequences to be fused are divided into a first coding sequence set and a second coding sequence set, wherein the first coding sequence set includes a first image coding sequence to be fused, the second coding sequence set includes a second image coding sequence to be fused, and the number of sequences of the first image coding sequence to be fused is the same as the number of sequences of the second image coding sequence to be fused; Determining a first image information code inserted in the first image coding sequence to be fused, and determining a second image information code inserted in the second image coding sequence to be fused; Determine a coding similarity between a target first image information coding and a target second image information coding, wherein the target first image information coding is any one of the first image information codings, and the target second image information coding is any one of the second image information codings; Based on the coding similarity, the first coding sequence set and the second coding sequence set are fused to obtain the fused image coding sequence.

6. The image processing method according to claim 5, wherein the step of fusing the first coding sequence set and the second coding sequence set based on the coding similarity to obtain the fused image coding sequence comprises: Determine a target first coding sequence from the first coding sequence set, and determine a target second coding sequence from the second coding sequence set, wherein the target first coding sequence is the first image coding sequence to be fused inserted into the target first image information coding, and the target second coding sequence is the second image coding sequence to be fused inserted into the target second image information coding; Determine the target first coding sequence and the target second coding sequence as a coding sequence pair, and determine a similar coding sequence pair from the coding sequence pairs according to the coding similarity; Determine, from the similar coding sequences, a plurality of first image codes to be fused and the target first image information coding in the target first coding sequence, and determine a plurality of second image codes to be fused and the target second image information coding in the target second coding sequence; According to the sequence position information of each first image code to be fused and the sequence position information of each second image code to be fused, determining corresponding associated codes for each first image code to be fused from the second image code to be fused; respectively calculating average values ​​between the first image codes to be fused and the corresponding associated codes, and obtaining multiple fused image codes based on multiple average values; The target first image information code and the target second image information code are encoded to obtain a fused image information code, and the fused image code sequence is constructed based on the fused image information code and the multiple fused image codes.

7. The image processing method according to claim 3, wherein the step of utilizing the coding extension network layer to perform coding extension on the fused image coding sequence to obtain the extended image coding sequence comprises: Inputting the fused image coding sequence into the coding extension network layer, and generating a coding feature map in the coding extension network layer according to a plurality of fused image codes in the fused image coding sequence; Determining a target coding feature from a plurality of coding features included in the coding feature map, and expanding the coding feature map according to the target coding feature to obtain an expanded coding feature map; Performing a shift operation on a plurality of coding features included in the extended coding feature map to obtain a shift coding feature map, and performing coding expansion based on the extended coding feature map and the shift coding feature map to obtain a plurality of extended codes; An extended image coding sequence is generated based on the fused image information coding in the fused image coding sequence and the multiple extended codings.

8. The image processing method according to claim 7, wherein the step of performing coding extension based on the extended coding feature map and the displacement coding feature map to obtain a plurality of extended codes comprises: Determining a plurality of extended coding features in the extended coding feature map, and determining a plurality of displacement coding features in the displacement coding feature map; Determining a corresponding associated displacement coding feature for each of the extended coding features from the plurality of displacement coding features according to the extended coding feature map position of each of the extended coding features and the displacement coding feature map position of each of the displacement coding features; Combining the target extended coding feature and the corresponding associated displacement coding feature to obtain a plurality of coding feature pairs, wherein the target extended coding feature is any one of the plurality of extended coding features; Merging the expanded coding features and the displacement coding features contained in each coding feature pair respectively to obtain a plurality of merged coding features, and determining a merged coding feature map based on the plurality of merged coding features; Two adjacent merged coding features in the merged coding feature map are subjected to feature fusion to obtain a coding extension feature map, and the coding extension feature map is encoded to obtain multiple extension codes.

9. The image processing method according to any one of claims 1 to 8, wherein determining the image processing result corresponding to the target object image according to the target object image coding sequence comprises: The target object image coding sequence is processed by utilizing the task execution unit in the image processing model to obtain an image processing result corresponding to the target object image, wherein the task execution unit is associated with the target image processing task.

10. The image processing method according to claim 9, wherein the task execution unit is a task execution network layer; The step of using the task execution network layer in the image processing model to process the target object image coding sequence to obtain the image processing result corresponding to the target object image includes: Determining a plurality of target image information codes from the target object image code sequence, and performing average pooling processing on the plurality of target image information codes using an average pooling network layer in the image processing model to obtain an average image information code; The task execution network layer in the image processing model is utilized to execute the target image processing task for the target object image according to the average image information encoding to obtain the image processing result.

11. The image processing method according to any one of claims 1 to 8, wherein the step of determining a plurality of sub-images of the target object image, encoding each sub-image, and obtaining a target image encoding sequence comprises: The image processing unit of the image processing model is used to determine the multiple sub-images from the target object image, and perform encoding processing on each sub-image to obtain a target image encoding sequence.

12. The image processing method according to claim 11, wherein determining the plurality of sub-images from the target object image and encoding each sub-image to obtain a target image encoding sequence comprises: In the image processing unit, dividing the target object image into the plurality of sub-images; Performing encoding processing on pixels contained in each sub-image respectively to obtain initial image codes corresponding to the pixels, and using the initial image codes to obtain initial image code sequences corresponding to the sub-images, wherein the initial image codes are used to characterize local features of the corresponding sub-images; The image information code is inserted into each initial image code sequence respectively, and the target image code sequence is obtained by using a plurality of initial image code sequences into which the image information code is inserted.

13. A pathological image processing method, comprising: Determine multiple sub-images of the pathological image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted in each initial image encoding sequence, and each image information code is used to characterize the global features of the corresponding sub-image; Inputting the target image coding sequence into the coding processing sub-model of the image processing model, using the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, to obtain a pathological image coding sequence; An image processing result corresponding to the pathological image is determined according to the pathological image coding sequence.

14. A method for training an image processing model, comprising: Determine a plurality of sub-images of a sample object image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted into each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image; Inputting the target image coding sequence into the coding processing sub-model of the image processing model, using the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, to obtain a sample object image coding sequence; Determining a sample image processing result corresponding to the sample object image according to the sample object image coding sequence; A sample label corresponding to the sample object image is determined, and the sample label and the sample image processing result are used to perform model training on the encoding processing sub-model of the image processing model to obtain a trained encoding processing sub-model.

15. A method for computer-aided diagnosis of cancer, comprising: Determine multiple sub-images of the tumor image, perform encoding processing on each sub-image, and obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information code inserted in each initial image encoding sequence, and each image information code is used to characterize a global feature of a corresponding sub-image; Inputting the target image coding sequence into a coding processing sub-model of an image processing model, and using the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, to obtain a tumor image coding sequence; A tumor image processing result corresponding to the tumor image is determined according to the tumor image coding sequence.

16. An image processing method, applied to a client of a medical system, comprising: In response to a user's clicking operation on a display interface of the client, determining a medical image to be processed; The medical image to be processed is sent to the server of the medical system, and the image processing result corresponding to the medical image to be processed returned by the server is received, wherein the image processing result is determined according to a medical image coding sequence, and the medical image coding sequence is obtained by learning the coding relationship of an initial image coding sequence corresponding to each sub-image included in a target image coding sequence and an image information coding inserted in each initial image coding sequence using a coding processing sub-model of an image processing model, the image information coding is used to characterize the global features of the corresponding sub-image, and the target image coding sequence is obtained by coding multiple sub-images of the medical image to be processed.

17. An image processing method, applied to a cloud-side device, comprising: A plurality of sub-images of a target object image sent by a receiving end-side device are encoded to obtain a target image encoding sequence, wherein the target image encoding sequence includes an initial image encoding sequence corresponding to each sub-image and an image information encoding inserted into each initial image encoding sequence, and each image information encoding is used to represent a global feature of a corresponding sub-image; Inputting the target image coding sequence into the coding processing sub-model of the image processing model, using the coding processing sub-model to learn the coding relationship of each initial image coding sequence and the image information coding inserted in each initial image coding sequence, to obtain the target object image coding sequence; Determining an image processing result corresponding to the target object image according to the target object image coding sequence; The image processing result is sent to the terminal side device.

18. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 17 are implemented.

19. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 17.

20. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 17.