Image Processing, Training Method, Device, Equipment and Medium of Image Processing Model
By using the target image processing model to predict probability of the image to be processed, a first probability map is generated and an image of the target object is acquired, the problem of low efficiency and poor accuracy of manual participation in the prior art is solved, and an automated, efficient and high-precision target object image acquisition is achieved.
Patent Information
- Application Number
- CN202111115221.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-09-23
AI Technical Summary
In the prior art, the process of obtaining the image of the target object requires manual participation, which is inefficient and poorly accurate, and is greatly affected by artificial subjective factors.
By acquiring the to-process image and the target image processing model of the reference object, using the target image processing model to probability prediction on the to-process image, generating a first probability map, and obtaining the image of the target object based on the probability map. The target image processing model is trained based on the sample image and the corresponding label.
It realizes automatic acquisition of the image of the target object, improves efficiency and accuracy, reduces manual intervention, and enhances the reliability of image recognition.
Smart Images

Figure CN114299353B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and particularly to an image processing method, an image processing model training method, device, equipment, and medium for an image processing model. Background Art
[0002] With the development of artificial intelligence technology, there are more and more application scenarios for image processing. Among them, one application scenario is: obtaining an image of a target object according to an image of a reference object, so as to perform tasks such as registration and image annotation using the image of the target object. Among them, the target object is an object that meets the reference conditions corresponding to the reference object. For example, meeting the reference conditions means that there are no foreign objects attached to the contour and there are no holes inside.
[0003] Currently, the process of obtaining an image of a target object requires manual participation, with low efficiency, being greatly affected by subjective factors of humans, and the accuracy of the obtained image of the target object being poor. Summary of the Invention
[0004] Embodiments of the present application provide an image processing method, an image processing model training method, device, equipment, and medium for an image processing model, which can be used to improve the efficiency of obtaining an image of a target object and the accuracy of the obtained image of the target object. The technical solution is as follows:
[0005] On the one hand, embodiments of the present application provide an image processing method, and the method includes:
[0006] Obtaining a to-be-processed image of a reference object and a target image processing model, where the target image processing model is trained based on a sample image of the reference object and a label corresponding to the sample image, and the label corresponding to the sample image is used to indicate a first sample pixel point in the sample image, and the object category corresponding to the first sample pixel point is a target object that meets the reference conditions corresponding to the reference object;
[0007] Invoking the target image processing model to perform probability prediction on a representation image of the to-be-processed image to obtain a first probability map, where the pixel value of a pixel point in the first probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object;
[0008] Based on the first probability map, obtaining an image of the target object.
[0009] There is also provided an image processing model training method, and the method includes:
[0010] Obtain a sample image of a reference object and a label corresponding to the sample image, where the label corresponding to the sample image is used to indicate a first sample pixel point in the sample image, and the object category corresponding to the first sample pixel point is a target object corresponding to the reference object that meets the reference condition;
[0011] Call an initial image processing model to perform probability prediction on the representation image of the sample image to obtain a sample probability map, where the pixel value of a pixel point in the sample probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object;
[0012] Based on the sample probability map and the label corresponding to the sample image, train the initial image processing model to obtain a target image processing model.
[0013] On the other hand, an image processing device is provided, and the device includes:
[0014] A first acquisition unit, configured to acquire a to-be-processed image of a reference object and a target image processing model, where the target image processing model is trained based on a sample image of the reference object and a label corresponding to the sample image, and the label corresponding to the sample image is used to indicate a first sample pixel point in the sample image, and the object category corresponding to the first sample pixel point is a target object corresponding to the reference object that meets the reference condition;
[0015] A first processing unit, configured to call the target image processing model to perform probability prediction on the representation image of the to-be-processed image to obtain a first probability map, where the pixel value of a pixel point in the first probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object;
[0016] A second acquisition unit, configured to acquire an image of the target object based on the first probability map.
[0017] In a possible implementation manner, the device further includes:
[0018] A second processing unit, configured to perform binarization processing on the to-be-processed image according to a contour threshold corresponding to the contour of the reference object to obtain a threshold segmentation image corresponding to the to-be-processed image;
[0019] A third acquisition unit, configured to acquire a representation image of the to-be-processed image based on the threshold segmentation image.
[0020] In a possible implementation manner, the third acquisition unit is configured to acquire an auxiliary image corresponding to the to-be-processed image; splice the threshold segmentation image and the auxiliary image to obtain the representation image.
[0021] In a possible implementation, the auxiliary image includes at least one of a first image and a second image. The first image is obtained by processing the image to be processed according to an imaging value range corresponding to the contour of the reference object, and the second image is obtained by processing the image to be processed according to an imaging value range corresponding to a reference element inside the reference object.
[0022] In a possible implementation, the apparatus further includes:
[0023] A determination unit, configured to determine a reference threshold range corresponding to the contour of the reference object; and determine the contour threshold within the reference threshold range.
[0024] In a possible implementation, the second acquisition unit is configured to obtain a second probability map having the same size as the image to be processed based on the first probability map; and perform binarization processing on the second probability map according to a probability threshold to obtain an image of the target object.
[0025] In a possible implementation, the apparatus further includes:
[0026] A registration unit, configured to construct a virtual model of the target object based on the image of the target object; obtain point cloud data of the reference object, where the point cloud data is obtained by scanning the reference object; register the point cloud data with the virtual model, and display the registration result.
[0027] In a possible implementation, the image to be processed is an image of an interest region in an original image, the image of the target object has the same size as the image to be processed, and the apparatus further includes:
[0028] A labeling unit, configured to label the original image according to the image of the target object, and display the labeled original image.
[0029] In a possible implementation, the image to be processed is a head medical image, the reference object is a head with no foreign object attached to the contour and having a hole inside, or a head with a foreign object attached to the contour and having a hole inside, and the target object meeting the reference condition is a head meeting at least one of the conditions that no foreign object is attached to the contour or no hole exists inside.
[0030] There is also provided a training apparatus for an image processing model, where the apparatus includes:
[0031] An acquisition unit, configured to acquire a sample image of a reference object and a label corresponding to the sample image, where the label corresponding to the sample image is used to indicate a first sample pixel point in the sample image, and the object category corresponding to the first sample pixel point is a target object meeting the reference condition corresponding to the reference object;
[0032] A processing unit, configured to call an initial image processing model to perform probability prediction on a representation image of the sample image, so as to obtain a sample probability map, where the pixel value of a pixel point in the sample probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object;
[0033] A training unit, configured to train the initial image processing model based on the sample probability map and the label corresponding to the sample image, so as to obtain a target image processing model.
[0034] In a possible implementation manner, the obtaining unit is configured to obtain an initial image of a reference object; perform data augmentation on the initial image to obtain the sample image of a reference size.
[0035] On the other hand, a computer device is provided, which includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor, so that the computer device implements any one of the above-mentioned image processing methods or the training method of the image processing model.
[0036] On the other hand, a computer-readable storage medium is further provided. At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor, so that a computer implements any one of the above-mentioned image processing methods or the training method of the image processing model.
[0037] On the other hand, a computer program product is further provided, which includes a computer program or computer instructions. The computer program or the computer instructions are loaded and executed by a processor, so that a computer implements any one of the above-mentioned image processing methods or the training method of the image processing model.
[0038] The technical solution provided by the embodiments of the present application at least brings the following beneficial effects:
[0039] The technical solution provided by the embodiments of the present application calls a target image processing model to automatically obtain an image of a target object. The process of obtaining the target image does not require manual participation, and the efficiency of obtaining the target image is relatively high. In addition, the target image processing model is trained based on the label of the pixel point in the sample image corresponding to the sample image, which is used to indicate that the corresponding object category in the sample image is the target object. It has the function of accurately identifying the pixel point whose corresponding object category in the image is the target object. The reliability of the first probability map obtained by calling the target image processing model is relatively high, so that the accuracy of the finally obtained image of the target object is relatively high. Description of the Drawings
[0040] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0041] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0042] Figure 2 It is a flowchart of an image processing method provided by an embodiment of the present application;
[0043] Figure 3 It is a schematic diagram of an image to be processed provided by an embodiment of the present application;
[0044] Figure 4 It is a schematic diagram of a process of obtaining an image of a target object provided by an embodiment of the present application;
[0045] Figure 5 It is a schematic diagram of a comparison result provided by an embodiment of the present application;
[0046] Figure 6 It is a schematic diagram of a comparison result provided by an embodiment of the present application;
[0047] Figure 7 It is a schematic diagram of a comparison result provided by an embodiment of the present application;
[0048] Figure 8 It is a schematic diagram of a comparison result provided by an embodiment of the present application;
[0049] Figure 9 It is a schematic diagram of a comparison result provided by an embodiment of the present application;
[0050] Figure 10 It is a schematic diagram of a registration process provided by an embodiment of the present application;
[0051] Figure 11 It is a schematic diagram of an image annotation process provided by an embodiment of the present application;
[0052] Figure 12 It is a flowchart of a training method for an image processing model provided by an embodiment of the present application;
[0053] Figure 13 It is a schematic diagram of an image processing device provided by an embodiment of the present application;
[0054] Figure 14 It is a schematic diagram of a training device for an image processing model provided by an embodiment of the present application;
[0055] Figure 15 It is a schematic structural diagram of a terminal provided by an embodiment of the present application;
[0056] Figure 16 It is a schematic structural diagram of a server provided by an embodiment of the present application. Specific implementation manners
[0057] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe in detail the embodiments of the present application with reference to the accompanying drawings.
[0058] In an exemplary embodiment, the image processing method and the training method of the image processing model provided by the embodiments of the present application can be applied to the field of artificial intelligence technology. Next, the artificial intelligence technology will be introduced.
[0059] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0060] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation. The image processing method and the training method of the image processing model provided by the embodiments of the present application involve computer vision technology and machine learning technology.
[0061] Computer Vision (CV) technology is a science that studies how to enable machines to "see". Further speaking, it refers to machine vision that uses cameras and computers to replace human eyes to identify and measure targets, and further perform graphic processing to make the computer process images that are more suitable for human eyes to observe or transmit to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D (Three Dimensional) technology, virtual reality, augmented reality, map construction, autonomous driving, intelligent transportation and other technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0062] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specializes in studying how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0063] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, intelligent healthcare, intelligent customer service, vehicle networking, autonomous driving, intelligent transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0064] In an exemplary embodiment, the image processing method and the training method of the image processing model provided in the embodiments of the present application are implemented in a blockchain system. The images to be processed, images of target objects, etc. involved in the image processing method provided in the embodiments of the present application, as well as the sample images, labels corresponding to the sample images, and the target image processing model involved in the training method of the image processing model are all stored on the blockchain in the blockchain system for each node device in the blockchain system to use, so as to ensure the security and reliability of the data.
[0065] Figure 1The figure shows a schematic diagram of the implementation environment provided by the embodiments of the present application. The implementation environment includes: a terminal 11 and a server 12.
[0066] The image processing method provided by the embodiments of the present application can be executed by the terminal 11, or by the server 12, or jointly by the terminal 11 and the server 12. The embodiments of the present application do not limit this. For the case where the image processing method provided by the embodiments of the present application is jointly executed by the terminal 11 and the server 12, the server 12 undertakes the main computing work, and the terminal 11 undertakes the secondary computing work; or, the server 12 undertakes the secondary computing work, and the terminal 11 undertakes the main computing work; or, the server 12 and the terminal 11 adopt a distributed computing architecture for collaborative computing.
[0067] The training method of the image processing model provided by the embodiments of the present application can be executed by the terminal 11, or by the server 12, or jointly by the terminal 11 and the server 12. The embodiments of the present application do not limit this. For the case where the training method of the image processing model provided by the embodiments of the present application is jointly executed by the terminal 11 and the server 12, the server 12 undertakes the main computing work, and the terminal 11 undertakes the secondary computing work; or, the server 12 undertakes the secondary computing work, and the terminal 11 undertakes the main computing work; or, the server 12 and the terminal 11 adopt a distributed computing architecture for collaborative computing.
[0068] The image processing method and the training method of the image processing model provided by the embodiments of the present application can be executed by the same device, or by different devices. The embodiments of the present application do not limit this.
[0069] In a possible implementation, the terminal 11 can be any electronic product that can perform human-computer interaction with the user in one or more ways such as a keyboard, a touchpad, a touch screen, a remote control, voice interaction, or a handwriting device. For example, a PC (Personal Computer), a mobile phone, a smart phone, a PDA (Personal Digital Assistant), a wearable device, a PPC (Pocket PC), a tablet computer, a smart vehicle console, a smart TV, a smart speaker, a vehicle-mounted terminal, etc. The server 12 can be a single server, or a server cluster composed of multiple servers, or a cloud computing service center. The terminal 11 and the server 12 establish a communication connection through a wired or wireless network.
[0070] Those skilled in the art should understand that the above-mentioned terminal 11 and server 12 are only examples. Other existing or future terminal or server that can be applied to the present application should also be included within the protection scope of the present application and are hereby incorporated by reference.
[0071] Based on the above Figure 1 shown implementation environment, an embodiment of the present application provides an image processing method, which is executed by a computer device. The computer device can be server 12 or terminal 11, and the embodiment of the present application does not limit this. As Figure 2 shown, the image processing method provided by the embodiment of the present application includes the following steps 201 to 203.
[0072] In step 201, an image to be processed of a reference object and a target image processing model are obtained. The target image processing model is trained based on a sample image and a label corresponding to the sample image. The label corresponding to the sample image is used to indicate a first sample pixel point in the sample image, and the object category corresponding to the first sample pixel point is a target object that meets the reference condition corresponding to the reference object.
[0073] The image to be processed of the reference object is an image that needs to be processed corresponding to the reference object, and the reference object refers to an object that needs to be concerned about. Exemplarily, the image to be processed of the reference object includes a sub-image of the reference object. The embodiment of the present application does not limit the type of the image to be processed. Exemplarily, the image to be processed is a medical image, and different modalities of medical images can be obtained by using different medical image acquisition devices, and various modalities of medical images can be used as the image to be processed.
[0074] The type of the image to be processed is related to the type of the reference object, and the embodiment of the present application does not limit this. In an exemplary embodiment, the type of the reference object is the head of a human body, and the image to be processed refers to a head medical image obtained by image acquisition of the head of a human body. For example, a CT (Computed Tomography) image obtained by image acquisition of the head of a human body, an MRI (Magnetic Resource Imaging) image obtained by image acquisition of the head of a human body, etc. In an exemplary embodiment, the type of the reference object is the abdomen of a human body, and the image to be processed refers to an abdominal medical image obtained by image acquisition of the abdomen of a human body. In an exemplary embodiment, the type of the reference object is the chest of a human body, and the image to be processed refers to a chest medical image obtained by image acquisition of the chest of a human body.
[0075] Of course, the image to be processed can also refer to a medical image obtained by image acquisition of other parts of a human body or other images that need to be processed, and the embodiment of the present application does not limit this. It should be noted that the image to be processed can be a two-dimensional (2D) image (also referred to as a planar image) or a three-dimensional (3D) image (also referred to as a stereoscopic image), and the embodiment of the present application does not limit this.
[0076] In an exemplary embodiment, in addition to the sub-image of the reference object, the image to be processed of the reference object may also include the sub-images of objects that are not connected to the contour of the reference object, which is not limited in the embodiments of the present application. During the process of processing the image to be processed, the objects that are not connected to the contour of the reference object are objects that do not need to be concerned about. In an exemplary embodiment, the objects that are not connected to the contour of the reference object include, but are not limited to, at least one of an examination table or a fixing rack. Exemplarily, for the case where the image to be processed is a CT image, the examination table refers to a CT examination table. Due to the CT circular scanning, the CT examination table forms a plate-like structure in the image. Again, for the case where the image to be processed is an MRI image, the examination table refers to an MRI examination table. It should be noted that in some cases, the examination table may be connected to the contour of the reference object. In such a case, the examination table is regarded as a foreign object attached to the contour of the reference object.
[0077] In an exemplary embodiment, the reference object satisfies at least one of having a hole inside or having a foreign object attached to its contour. That is to say, the reference object may be an object with a hole inside, or an object with a foreign object attached to its contour, or an object with a hole inside and a foreign object attached to its contour. In an exemplary embodiment, the image to be processed is a head medical image. Since there are holes inside the head, the reference object is a head with no foreign object attached to its contour and having a hole inside, or a head with a foreign object attached to its contour and having a hole inside.
[0078] For the case where there is a hole inside the reference object, in addition to the hole, the inside of the reference object also includes some solid elements. The embodiments of the present application do not limit the type, quantity, size, etc. of the holes inside the reference object, which are related to the type of the reference object and the actual situation. Exemplarily, in the case where the reference object is an object with a hole inside, the image to be processed is as Figure 3 shown in (a) and (b) of Figure 3 The 3D image of the reference object with a hole inside is shown in (a) of Figure 3 The 2D image of the reference object with a hole inside is shown in (b) of
[0079] For the case where there is a foreign object attached to the contour of the reference object, the embodiments of the present application do not limit the position, type, size, etc. of the foreign object attached to the contour of the reference object, which are related to the type of the reference object and the actual situation. In an exemplary embodiment, the type of the reference object is a certain part of the human body or a certain anatomical structure, and the foreign objects attached to the contour of the reference object include, but are not limited to, at least one of an infusion tube, an air supply tube, and a surgical bandage.
[0080] Exemplarily, in the case where there is a foreign object attached to the contour of the reference object, the image to be processed is asFigure 3 as shown in (c) and (d) therein. In Figure 3 the to-be-processed image shown in (c) therein, the type of the reference object is the head of a human body, and the foreign objects attached to the contour of the reference object include an examination bed; in Figure 3 the to-be-processed image shown in (d) therein, the type of the reference object is the head of a human body, and the foreign objects attached to the contour of the reference object include surgical bandages, etc.
[0081] In an exemplary embodiment, the holes existing inside the reference object and the foreign objects attached to the contour of the reference object will interfere with the execution effect of downstream tasks, and the downstream tasks include but are not limited to registration tasks, image annotation tasks, etc. The execution process of the downstream tasks is based on the target object corresponding to the reference object that meets the reference conditions. The target object corresponding to the reference object can be regarded as a virtual object obtained by optimizing the reference object.
[0082] The reference conditions are determined according to application requirements, and the embodiments of the present application do not limit this. In an exemplary embodiment, meeting the reference conditions includes at least one of having no foreign objects attached to the contour or having no holes inside. That is to say, the target object may refer to an object with no foreign objects attached to the contour corresponding to the reference object, or may refer to an object with no holes inside the reference object, or may refer to an object with no foreign objects attached to the contour and no holes inside the reference object corresponding to the reference object. Exemplarily, taking the to-be-processed image as a head medical image as an example, the target object that meets the reference conditions is a head that meets at least one of having no foreign objects attached to the contour or having no holes inside.
[0083] The embodiments of the present application will be described by taking the reference condition as referring to having no foreign objects attached to the contour and having no holes inside, and the target object being an object with no foreign objects attached to the contour and no holes inside corresponding to the reference object as an example. In this case, for the case where the reference object is an object with no foreign objects attached to the contour and no holes inside, the target object is the reference object itself; for the case where the reference object is an object with no foreign objects attached to the contour and having holes inside, the target object is the object obtained by filling the holes existing inside the reference object; for the case where the reference object is an object with foreign objects attached to the contour and no holes inside, the target object is the object obtained by removing the foreign objects attached to the contour of the reference object; for the case where the reference object is an object with foreign objects attached to the contour and having holes inside, the target object is the object obtained by removing the foreign objects attached to the contour of the reference object and filling the holes existing inside the reference object.
[0084] The embodiments of the present application do not limit the way of obtaining the image to be processed. In an exemplary embodiment, the ways for a computer device to obtain the image to be processed of a reference object include but are not limited to: the computer device extracts the image to be processed of the reference object from an image library; an image acquisition device having a communication connection with the computer device sends the acquired image to be processed of the reference object to the computer device; the computer device obtains the image to be processed of the reference object uploaded manually, etc.
[0085] In an exemplary embodiment, the image to be processed of the reference object is an image of the region of interest of the original image. That is to say, the computer device crops out the image of the region of interest indicated manually or by the computer device itself in the original image to obtain the image to be processed of the reference object. The region of interest refers to the region that is of interest to a human or the computer device, and is usually a local region in the entire region of the original image. In different application scenarios, the size of the region of interest and the position of the region of interest in the entire region of the original image may be different, and the embodiments of the present application do not limit this. In an exemplary embodiment, if the original image is a 3D image, the image of the region of interest is also a 3D image, and in this case, the region of interest can be called VOI (Volume of Interest); if the original image is a 2D image, the image of the region of interest is also a 2D image, and in this case, the region of interest can be called ROI (Region of Interest).
[0086] In an exemplary embodiment, the way for the computer device to obtain the image to be processed of the reference object can also be: the computer device obtains the image to be processed of the reference size based on the candidate image of the reference object. The candidate image of the reference object is extracted from an image library or uploaded manually, etc. This way can ensure that the size of the image to be processed is the reference size. The reference size is set according to experience or flexibly adjusted according to the actual application scenario, and the embodiments of the present application do not limit this. Exemplarily, the image to be processed is a 3D image, and the reference size is 192×192×160 (pixels). Exemplarily, the image to be processed is a 2D image, and the reference size is 192×192 (pixels).
[0087] In an exemplary embodiment, the manner of obtaining a to-be-processed image with a reference size based on a candidate image related to the reference object is related to the size relationship between the size of the candidate image and the reference size. Exemplarily, if the size of the candidate image is smaller than the reference size, the candidate image is upsampled to obtain a to-be-processed image with the reference size; if the size of the candidate image is larger than the reference size, the candidate image is downsampled to obtain a to-be-processed image with the reference size; if the size of the candidate image is the same as the reference size, the candidate image is directly used as the to-be-processed image. Upsampling refers to the operation of enlarging the image size, and downsampling refers to the operation of reducing the image size. The embodiments of the present application do not limit the implementation manners of upsampling and downsampling. Exemplarily, the implementation manners of upsampling include, but are not limited to, nearest neighbor interpolation, bilinear interpolation, mean interpolation, etc.; the implementation manners of downsampling include, but are not limited to, maximum sampling, average sampling, etc.
[0088] The target image processing model is a trained model for performing probability prediction on the representation image of the to-be-processed image to obtain a first probability map. The target image processing model is trained based on the sample image of the reference object and the label corresponding to the sample image. The label corresponding to the sample image is used to indicate the first sample pixel points in the sample image, and the object category corresponding to the first sample pixel points is the target object corresponding to the reference object. The manner of obtaining the target image processing model may refer to obtaining the target image processing model through real-time training, or may refer to extracting the target image processing model that has been pre-trained and stored. The embodiments of the present application do not limit this. For the process of training the target image processing model, refer to Figure 12 the embodiments shown, which will not be elaborated here for the time being.
[0089] The embodiments of the present application do not limit the model structure of the target image processing model. In an exemplary embodiment, the target image processing model is a convolutional neural network model. Exemplarily, the target image processing model is a convolutional neural network model for segmentation, such as a UNet (U-shaped network) model, an FCN (Fully Convolutional Networks) model, a DSN (Deeply-Supervised Networks) model, etc. In an exemplary embodiment, the target image processing model can also be a lightweight model set according to experience to improve the processing efficiency of the model. In an exemplary embodiment, in different application scenarios, target image processing models with different model structures can be used. For example, in the case where the size of the to-be-processed image is small and the task is relatively simple, a target image processing model with a more streamlined model structure is used to achieve the purpose of fast processing.
[0090] In step 202, the target image processing model is called to perform probability prediction on the representative image of the image to be processed, and a first probability map is obtained. The pixel value of a pixel point in the first probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object.
[0091] After obtaining the target image processing model and the image to be processed, the target image processing model is called to perform probability prediction on the representative image of the image to be processed, and a first probability map is obtained. The representative image of the image to be processed is used to represent the image to be processed, and the representative image of the image to be processed is obtained based on the image to be processed. Before performing step 202, it is necessary to first obtain the representative image of the image to be processed.
[0092] In a possible implementation manner, the process of obtaining the representative image of the image to be processed includes the following steps 2021 and 2022:
[0093] Step 2021: Perform binarization processing on the image to be processed according to the contour threshold corresponding to the contour of the reference object, and obtain a threshold segmentation image corresponding to the image to be processed.
[0094] The contour threshold corresponding to the contour of the reference object is the threshold based on which the binarization processing of the image to be processed is performed. The contour threshold corresponding to the contour of the reference object is determined by considering the characteristics of the contour of the reference object, and can separate the reference object from other objects to a large extent, so as to obtain a threshold segmentation image that can carry certain contour information.
[0095] The contour threshold is used to constrain the pixel value of each pixel point in the image to be processed. That is to say, the physical meanings of the contour threshold and the pixel value of each pixel point in the image to be processed are the same. Exemplarily, the physical meanings of the pixel values of pixel points in different types of images to be processed may be different. For example, if the image to be processed is a CT image, the pixel value of a pixel point in the image to be processed refers to the CT value, and the CT value is a measurement unit for measuring the density of a certain local tissue or organ of the human body. In this case, the contour threshold also refers to the CT value; if the image to be processed is an MRI image, the pixel value of a pixel point in the image to be processed refers to the magnetic resonance signal intensity value. In this case, the contour threshold also refers to the magnetic resonance signal intensity value. Exemplarily, pixel points with different pixel values are presented using different grayscales, and the pixel value of a pixel point can also be referred to as the gray value of the pixel point.
[0096] The contour threshold corresponding to the reference object is set empirically or adjusted flexibly according to the type of the reference object and the actual application scenario, and the embodiments of the present application do not limit this. In a possible implementation manner, the method for determining the contour threshold corresponding to the contour of the reference object is: determining the reference threshold range corresponding to the contour of the reference object; determining the contour threshold within the reference threshold range. The reference threshold range corresponding to the contour of the reference object is a range composed of thresholds that can separate the reference object from other objects to a certain extent. The reference threshold range corresponding to the contour of the reference object is set manually according to experience or determined through segmentation experiments, and the embodiments of the present application do not limit this. For example, exemplarily, taking the type of the reference object as a certain part of the human body and the image to be processed as a CT image, the contour of the reference object can be regarded as the skin of a certain part of the human body, and the reference threshold range corresponding to the skin is [-490, 290].
[0097] In an exemplary embodiment, during the process of processing the image to be processed, the contour threshold corresponding to the contour of the reference object is a certain fixed threshold within the reference threshold range, so as to facilitate the stability of the processing process of the image to be processed. That is to say, the same contour threshold is used for binarizing different images to be processed. Which threshold within the reference threshold range is used as the contour threshold is set according to experience, and the embodiments of the present application do not limit this. Exemplarily, the average value of the lower limit value and the upper limit value in the reference threshold range is used as the contour threshold. Exemplarily, the threshold that can achieve relatively accurate contour segmentation specified according to experience in the reference threshold range is used as the contour threshold. In an exemplary embodiment, taking the contour of the reference object as the skin of a certain part of the human body as an example, the reference threshold range corresponding to the skin is [-490, 290], and the contour threshold corresponding to the skin is -390, and -390 is the threshold that can achieve relatively accurate skin segmentation specified according to experience in the reference threshold range.
[0098] In an exemplary embodiment, the process of binarizing the image to be processed according to the contour threshold corresponding to the contour of the reference object is: for the first pixel point, if the pixel value of the first pixel point is not less than the contour threshold, then the pixel value of the first pixel point is transformed into a first value; if the pixel value of the first pixel point is less than the contour threshold, then the pixel value of the first pixel point is transformed into a second value; the image composed of the pixel points with the transformed pixel values is used as the threshold segmentation image. Wherein, the first pixel point is any pixel point among the respective pixel points in the image to be processed.
[0099] The first value and the second value are set empirically or flexibly adjusted according to the actual application scenario, and the embodiments of the present application do not limit this. Exemplarily, the first value is 0 and the second value is 255; or, the first value is 255 and the second value is 0, etc. Exemplarily, for the case where one of the first value and the second value is 0 and the other is 255, the process of binarizing the image to be processed according to the contour threshold is to set the pixel values of the pixels on the image to be processed to 0 or 255 according to the contour threshold, so that the entire image to be processed presents an obvious black-and-white effect. The image presenting the obvious black-and-white effect is the threshold segmentation image.
[0100] In an exemplary embodiment, the image to be processed is a medical image. In the medical image, most anatomical structures and human body parts, etc., have obvious boundary segmentation on the image. According to the contour threshold, a relatively fine boundary can be segmented. By obtaining the threshold segmentation image, in the subsequent processing process, on the basis of retaining the relatively fine boundary, false positive segmentation regions and the like irrelevant to the task can be further removed to further improve the fineness of the boundary. Exemplarily, the false positive segmentation regions include, but are not limited to, the regions where foreign objects are attached to the contour and the regions where other unconnected objects are located, etc.
[0101] Step 2022: Based on the threshold segmentation image, obtain the characterization image of the image to be processed.
[0102] After obtaining the threshold segmentation image, the characterization image of the image to be processed is obtained based on the threshold segmentation image. Since the threshold segmentation image is obtained by binarizing the image to be processed according to the contour threshold and can carry certain contour information, the characterization image of the image to be processed obtained based on the threshold segmentation image can at least characterize the features of the image to be processed in terms of the contour, which is beneficial to improving the recognition accuracy of the model for the contour and improving the reliability of the first prediction map.
[0103] In a possible implementation manner, the implementation process of obtaining the characterization image of the image to be processed based on the threshold segmentation image is: directly using the threshold segmentation image as the characterization image of the image to be processed. In this way, the characterization image of the image to be processed can be obtained without obtaining other images, and the efficiency of obtaining the characterization image of the image to be processed is relatively high. The threshold segmentation image can enable the model to better learn to fill holes and remove foreign objects.
[0104] In another possible implementation manner, the implementation process of obtaining the characterization image of the image to be processed based on the threshold segmentation image is: obtaining the auxiliary image corresponding to the image to be processed; splicing the threshold segmentation image and the auxiliary image to obtain the characterization image. The auxiliary image is used to provide more information on the basis of the threshold segmentation image. The embodiments of the present application do not limit the type and quantity of the auxiliary image, and can be flexibly adjusted according to the actual application scenario.
[0105] In an exemplary embodiment, the auxiliary image includes at least one of a first image and a second image. The first image is obtained by processing the image to be processed according to the imaging value range corresponding to the contour of the reference object, and the second image is obtained by processing the image to be processed according to the imaging value range corresponding to the reference element inside the reference object. The first image can provide information about the contour. Since the first image is obtained by processing the image to be processed according to the imaging value range of the contour, it can provide more comprehensive information about the contour than the threshold segmentation image.
[0106] The second image can provide information about the reference element inside the reference object. The reference element refers to the element that needs to be concerned inside the reference object. The reference element is related to the type of the reference object and the application scenario, etc., and the embodiments of the present application do not limit this. Exemplarily, when the type of the reference object is the head of a human body, the reference element refers to the brain tissue in the head of the human body. Exemplarily, when the type of the reference object is the abdomen of a human body, the reference element refers to the abdominal tissue in the abdomen of the human body. In an exemplary embodiment, in addition to the reference element, the inside of the reference object may also include other elements (such as holes, etc.), and the embodiments of the present application do not limit this.
[0107] The imaging value range corresponding to the contour of the reference object refers to the range of the pixel values of the pixel points involved in the contour of the reference object, and is also the range of the pixel values that can more reliably image the contour of the reference object. The imaging value range corresponding to the contour of the reference object is related to the imaging method of the image to be processed. The imaging value range corresponding to the contour of the reference object is composed of a first lower limit value and a first upper limit value. In an exemplary embodiment, for the case where the image to be processed is a CT image, the imaging value range corresponding to the contour of the reference object refers to the imaging window corresponding to the contour of the reference object. Exemplarily, the contour of the reference object refers to the skin of a human body, and the imaging window corresponding to the skin of the human body is [-390, 2000], that is, the lower limit CT value corresponding to the skin of the human body is -390, and the upper limit CT value corresponding to the skin of the human body is 2000. This imaging window [-390, 2000] is the CT value range that can more accurately display the skin of the human body.
[0108] In a possible implementation manner, the method for obtaining the first image by processing the image to be processed according to the imaging value range corresponding to the contour of the reference object is as follows: based on the pixel values of each pixel point in the image to be processed in the image to be processed, obtain the updated pixel values of each pixel point; based on the first lower limit value, the first upper limit value and the updated pixel values of each pixel point, determine the processed pixel values of each pixel point, and use the image composed of each pixel point with the processed pixel values as the first image.
[0109] In a possible implementation manner, the process of obtaining the updated pixel value of each pixel point based on the pixel value that each pixel point in the image to be processed has in the image to be processed includes: for the first pixel point, in response to the pixel value that the first pixel point has in the image to be processed being greater than the first upper limit value, taking the first upper limit value as the updated pixel value of the first pixel point; in response to the pixel value that the first pixel point has in the image to be processed being less than the first lower limit value, taking the first lower limit value as the updated pixel value of the first pixel point; in response to the pixel value that the first pixel point has in the image to be processed being not less than the first lower limit value and not greater than the first upper limit value, taking the pixel value that the first pixel point has in the image to be processed as the updated pixel value of the any pixel point. Wherein, the first pixel point is any pixel point among each pixel point in the image to be processed.
[0110] In an exemplary embodiment, the process of determining the processed pixel value of each pixel point based on the first lower limit value, the first upper limit value, and the updated pixel value of each pixel point is a process of constraining the updated pixel value of each pixel point by using the first lower limit value and the first upper limit value. Taking the first pixel point as an example, in an exemplary embodiment, the process of determining the processed pixel value of the first pixel point based on the first lower limit value, the first upper limit value, and the updated pixel value of the first pixel point is: calculating the difference between the updated pixel value of the first pixel point and the first lower limit value; taking the ratio of the difference to the first upper limit value as the processed pixel value of the first pixel point.
[0111] In an exemplary embodiment, the process of obtaining the processed pixel value of each pixel point is implemented based on Formula 1:
[0112]
[0113] where w min represents the first lower limit value; w max represents the first upper limit value; I in represents the pixel value that the pixel point has in the image to be processed; I out represents the processed pixel value of the pixel point; the Clip function is used to set the pixel values of all pixel points whose pixel values in the image to be processed are less than w min to w min , and set the pixel values of all pixel points whose pixel values in the image to be processed are greater than w max to w max . Taking the image composed of each pixel point with I out as the first image.
[0114] The imaging value range corresponding to the reference element inside the reference object refers to the range of pixel values of the pixel points involved in the reference element inside the reference object, and is also the range of pixel values that can more reliably image the reference element inside the reference object. The imaging value range corresponding to the reference element inside the reference object is related to the imaging method of the image to be processed. In an exemplary embodiment, for the case where the image to be processed is a CT image, the imaging value range corresponding to the reference element inside the reference object refers to the imaging window corresponding to the reference element inside the reference object. Exemplarily, the reference element inside the reference object refers to the brain tissue in the human head, and the imaging window corresponding to the brain tissue in the human head is [0, 80], that is, the lower limit CT value corresponding to the brain tissue in the human head is 0, and the upper limit CT value corresponding to the human skin is 80. This imaging window [0, 80] is the CT value range that can more accurately display the brain tissue in the human head. The principle of processing the image to be processed according to the imaging value range corresponding to the reference element inside the reference object to obtain the second image is the same as the principle of processing the image to be processed according to the imaging value range corresponding to the contour of the reference object to obtain the first image, which will not be elaborated here.
[0115] In a possible implementation manner, the auxiliary image includes at least one of the first image and the second image. That is to say, the auxiliary image may only include the first image, may only include the second image, or may include both the first image and the second image. After determining the specific situation of the auxiliary image, the auxiliary image can be obtained according to the acquisition methods of the first image and the second image introduced above. It should be noted that the auxiliary image may also include other images, and the embodiments of the present application are not limited thereto.
[0116] After obtaining the auxiliary image, the threshold segmentation image and the auxiliary image are spliced to obtain a characterization feature. In an exemplary embodiment, if the auxiliary image only includes the first image, then the threshold segmentation image and the first image are spliced to obtain a characterization image of the image to be processed. In an exemplary embodiment, if the auxiliary image only includes the second image, then the threshold segmentation image and the second image are spliced to obtain a characterization image of the image to be processed. In an exemplary embodiment, if the auxiliary image includes the first image and the second image, then the threshold segmentation image, the first image, and the second image are spliced to obtain a characterization image of the image to be processed.
[0117] In an exemplary embodiment, the threshold-segmented image and the auxiliary image are images of the same size and the same dimension. The splicing of the threshold-segmented image and the auxiliary image refers to splicing them in the channel dimension. In an exemplary embodiment, the process of splicing the threshold-segmented image and the auxiliary image can be implemented by writing a program or code before inputting the target image processing model, or can be implemented by the target image processing model after inputting the target image processing model. The embodiments of the present application do not limit this.
[0118] It should be noted that the above steps 2021 and 2022 are only an exemplary implementation for obtaining the representative image of the image to be processed, and the embodiments of the present application are not limited thereto. In an exemplary embodiment, the image to be processed itself can also be directly used as the representative image of the image to be processed.
[0119] In any case, after obtaining the representative image of the image to be processed, the target image processing model is called to perform probability prediction on the representative image of the image to be processed to obtain the first probability map. In an exemplary embodiment, for the case where the representative image refers to the image to be processed, the image to be processed is directly input into the target image processing model; for the case where the representative image refers to the threshold-segmented image corresponding to the image to be processed, the threshold-segmented image is directly input into the target image processing model; for the case where the representative image refers to the spliced image corresponding to the image to be processed, the respective images constituting the spliced image are input into the target image processing model.
[0120] Exemplarily, the respective images constituting the spliced image may refer to the two images of the threshold-segmented image and the first image, or may refer to the two images of the threshold-segmented image and the second image, or may also refer to the three images of the threshold-segmented image, the first image, and the second image. Exemplarily, for the case where the respective images constituting the spliced image refer to the three images of the threshold-segmented image, the first image, and the second image, the image input into the target image processing model can be regarded as a three-channel image, so that the target image processing model can obtain the first probability map from the information provided by the three-channel image. This method can enable the target image processing model to better adapt to the changes of the image to be processed and the threshold-segmented image.
[0121] The process of using the target image processing model to perform probability prediction on the representation image of the image to be processed is the internal processing process of the target image processing model. The specific implementation method is related to the model structure of the target image processing model. The internal processing processes of target image processing models with different structures may be different, and the embodiments of the present application do not limit this. In an exemplary embodiment, the last layer of the target image processing model is an activation layer, and the first probability map is the probability map output after being activated by the activation layer. The embodiments of the present application do not limit the type of the activation layer. For example, the activation layer is a Sigmoid (S-shaped) activation layer, a ReLU (Rectified Linear Unit) activation layer, etc.
[0122] In an exemplary embodiment, if the image to be processed is a 3D image, then the threshold segmentation image and the auxiliary image corresponding to the image to be processed are both 3D images, and the target image processing model is a 3D processing model for processing 3D images; if the image to be processed is a 2D image, then the threshold segmentation image and the auxiliary image corresponding to the image to be processed are both 2D images, and the target image processing model is a 2D processing model for processing 2D images.
[0123] The pixel value of the pixel point in the first probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object corresponding to the reference object. That is to say, the first probability map is an image presented according to the probability that the object category corresponding to each pixel point is the target object corresponding to the reference object. From the first probability map, it is possible to know the probability that the object category corresponding to each pixel point in the first probability map is the target object corresponding to the reference object. The dimension of the first probability map is the same as that of the image to be processed. In an exemplary embodiment, the size of the first probability map may be the same as the size of the image to be processed, or may be different from the size of the image to be processed. The embodiments of the present application do not limit this, which is related to the model structure of the target image processing model.
[0124] In step 203, based on the first probability map, an image of the target object is obtained.
[0125] The pixel value of the pixel point of the first probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object corresponding to the reference object. Based on the first probability map, an image of the target object can be obtained. The image of the target object refers to an image that removes other objects and only retains the target object. Exemplarily, the image of the target object has the same size as the image to be processed, so as to facilitate comparing and viewing the image of the target object with the image to be processed.
[0126] In a possible implementation manner, the method for obtaining an image of a target object based on a first probability map is as follows: based on the first probability map, obtain a second probability map with the same size as the image to be processed; perform binarization processing on the second probability map according to a probability threshold to obtain the image of the target object. In an exemplary embodiment, the method for obtaining a second probability map with the same size as the image to be processed based on the first probability map is as follows: in response to the first probability map having the same size as the image to be processed, use the first probability map as the second probability map; in response to the first probability map having a different size from the image to be processed, sample the first probability map to obtain the second probability map. The way of sampling the first probability map is related to the size relationship between the size of the first probability map and the size of the image to be processed. Exemplarily, if the size of the first probability map is smaller than the size of the image to be processed, then perform upsampling on the first probability map; if the size of the first probability map is larger than the size of the image to be processed, then perform downsampling on the first probability map. Since the second probability map is obtained based on the first probability map, the pixel values of the pixel points in the second probability map are also used to indicate the probability that the object category corresponding to the pixel point is the target object.
[0127] The probability threshold is used to constrain the pixel values of the pixel points in the second probability map. The probability threshold is set according to experience or flexibly adjusted according to the application scenario. The embodiments of the present application do not limit this. For example, the probability threshold is 0.5. Exemplarily, the method for performing binarization processing on the second probability map according to the probability threshold is as follows: in response to the pixel value of the second pixel point being less than the probability threshold, convert the pixel value of the second pixel point to a third value; in response to the pixel value of the second pixel point being not less than the probability threshold, convert the pixel value of the second pixel point to a fourth value, and use the image composed of the pixel points with the converted pixel values as the image of the target object. Wherein, the second pixel point is any one of the pixel points in the second probability map.
[0128] In an exemplary embodiment, the process of obtaining the image of the target object is as Figure 4As shown. Obtain a candidate image 401, sample the candidate image 401 to a reference size (e.g., 192×192×160 (pixels)) to obtain an image to be processed. Process the image to be processed according to the imaging value range corresponding to the contour of the image to be processed (e.g., [-390, 2000]) to obtain a first image; process the image to be processed according to the imaging value range corresponding to the reference elements inside the image to be processed (e.g., [0, 80]) to obtain a second image; perform binarization processing on the image to be processed according to the contour threshold determined from the reference threshold range (e.g., [-490, 290]) to obtain a threshold segmentation image. Input these three images, namely the first image, the second image, and the threshold segmentation image, into the target image processing model 402 to obtain a first probability map output by the target image processing model 402. Adjust the first probability map into a second probability map with the same size as the image to be processed, and perform binarization processing on the second probability map using a probability threshold to obtain an image 403 of the target object. This method of first sampling the probability map and then using threshold binary segmentation is beneficial to improving the smoothness of the image of the target object.
[0129] Exemplarily, compare the image of the target object obtained by using the method provided in the embodiments of the present application with the image of the target object obtained by using the method provided by the related art to verify the effectiveness of the method provided in the embodiments of the present application. The comparison results are as Figure 5 shown. Figure 5 In it, the type of the reference object is the head of a human body, and the target object is a virtual head with no foreign objects attached to the contour corresponding to the head of the human body and no holes inside. It should be noted that Figure 5 In each of the figures, there are some auxiliary lines, which are used to position the orientation of the head of the human body and will not affect the accuracy of the image of the target object. Figure 5 In the first column of images in it, the images of the target object are obtained after manually filling the holes in the threshold segmentation image, Figure 5 In the second column of images in it, the images of the target object are obtained by using the method provided in the embodiments of the present application, Figure 5 In the third column of images in it, the images of the target object are obtained after filling the holes in the threshold segmentation image by using a filling tool (e.g., 3D Slicer, an open-source medical image processing software).
[0130] According to Figure 5It can be seen that the details of the image of the target object obtained by manual filling are clear, but there is a problem of long time consumption. The details of the image of the target object obtained by using the method provided by the embodiments of the present application are also relatively clear, and it can be executed with one key, making full use of the GPU (Graphics Processing Unit), and the speed is relatively fast. There are many parameters in the filling tool and it is difficult to control. In the image of the target object obtained by filling with the filling tool, there is a situation where the contour segmentation is damaged.
[0131] Exemplarily, taking the type of the reference object as the human head as an example, the embodiments of the present application can be applied to various situations. For example, the sub-image of the examination bed is included in the image to be processed, the ear of the human head is connected to the examination bed, and there is an infusion tube at the nose position of the human head, etc. It is difficult to obtain the image of the target object through simple analysis (such as, the largest connected domain analysis), and the method provided by the embodiments of the present application can obtain a relatively accurate image of the target object in these relatively complex situations.
[0132] Exemplarily, the comparison result between the image obtained through simple analysis and the image of the target object obtained by using the method provided by the embodiments of the present application is as Figures 6 to 9 shown. Figure 6 The left figure in is the threshold segmentation image, and the sub-image of the examination bed exists in the threshold segmentation image; Figure 6 The right figure in is the image of the target object obtained by using the method provided by the embodiments of the present application. In the image of the target object, the sub-image of the examination bed is removed.
[0133] Figure 7 and Figure 8 The left figure in is the threshold segmentation image or the image after manually removing the sub-image of the examination bed in the threshold segmentation image. In the Figure 7 and Figure 8 left figure in, there are many foreign objects attached to the contour of the human head; Figure 7 and Figure 8 The right figure in is the image of the target object obtained by using the method provided by the embodiments of the present application. In the image of the target object, the foreign objects attached to the contour of the human head are removed. Figure 9 is the result of processing the image to be processed collected by thick-layer CT, Figure 9 The left figure in is the threshold segmentation image or the image after manually partially removing the sub-image of the examination bed that is not connected to the human head in the threshold segmentation image. In the Figure 9 left figure in, there is a sub-image of the examination bed connected to the human head; Figure 9 The right figure in is the image of the target object obtained by using the method provided by the embodiments of the present application. In the image of the target object, the sub-image of the examination bed connected to the human head is removed.
[0134] The image of the target object is the image based on which subsequent tasks are performed. The types of subsequent tasks performed using the image of the target object are not limited in the embodiments of the present application, and can be set according to experience or flexibly adjusted according to the actual application scenario. The embodiments of the present application will be described by taking the subsequent tasks of performing registration tasks and image annotation tasks using the image of the target object as examples.
[0135] In the application of head surgical navigation based on images such as CT and MRI, it is necessary to pre-segment the head and reconstruct the virtual model of the head, and register the point cloud data of the head obtained in real time by the surgical navigation with the reconstructed virtual model, so as to combine the preoperative image and the real-time picture to achieve the navigation function. The embodiments of the present application provide a fully automatic image processing method for quickly, highly accurately, and automatically filling the holes inside the head, which can effectively remove the sub-image of the examination bed in the CT image and the attached foreign objects on the skin of the head, and can also automatically fill the holes existing inside the head, thereby improving the accuracy of the reconstructed virtual model, the efficiency and registration accuracy of the registration based on the point cloud data.
[0136] In the application scenario of performing a registration task using the image of the target object, after obtaining the image of the target object, it further includes: constructing a virtual model of the target object based on the image of the target object; obtaining the point cloud data of the reference object, where the point cloud data is obtained by scanning the reference object; registering the point cloud data with the virtual model, and displaying the registration result.
[0137] The image of the target object may be a 2D image or a 3D image. In either case, after obtaining the image of the target object, a virtual model of the target object can be constructed based on the image of the target object. The virtual model of the target object is a three-dimensional model of the target object. The embodiments of the present application do not limit the manner of constructing the virtual model of the target object based on the image of the target object. Exemplarily, a virtual model reconstruction tool can be used to implement the process of constructing the virtual model of the target object based on the image of the target object. Exemplarily, based on the image of the target object, the virtual model of the target object is constructed using the mesh generation technology. It should be noted that the virtual model of the target object is the virtual model of the optimized object corresponding to the reference object, and is used to register with the actual point cloud data of the reference object.
[0138] The point cloud data of the reference object is obtained by scanning the reference object and is the actual data of the reference object. It should be noted that the point cloud data of the reference object obtained here is the data that can be used for registration with the virtual model of the target object. The point cloud data of the reference object and the image to be processed are for the same reference object, which can be for the same reference object at different times to ensure the reliability of registration. For example, the image to be processed is an image obtained by collecting an image of the reference object at the first moment (such as, the preoperative moment), and the point cloud data is the point cloud data obtained by scanning the same reference object at the second moment (such as, the intraoperative moment).
[0139] The device for scanning the reference object is selected according to experience or flexibly adjusted according to the actual application scenario, and the embodiments of the present application do not limit this. In an exemplary embodiment, the device for scanning the reference object includes, but is not limited to, a depth camera, a point cloud scanner, etc.
[0140] After obtaining the point cloud data and the virtual model, the point cloud data is registered with the virtual model to obtain a registration result, and the configuration result is used to indicate the matching relationship between the point cloud data and the virtual model. In an exemplary embodiment, the computer device has a display function, and after obtaining the registration result, the registration result is displayed, and this registration result can provide navigation for the operation. The embodiments of the present application do not limit the method of registering the point cloud data with the virtual model, and any registration method can be used, for example, the registration methods include, but are not limited to, the ICP (Iterative Closest Point) registration algorithm, the NDT (Normal Distribution Transform) registration algorithm, etc.
[0141] In an exemplary embodiment, registering the point cloud data with the virtual model may refer to directly registering the point cloud data with the virtual model. In an exemplary embodiment, the point cloud data may be obtained by scanning the reference object in a certain orientation. In this case, before registering the point cloud data with the virtual model, the virtual model can be trimmed first to obtain a sub-virtual model corresponding to a certain orientation, and then the point cloud data is registered with the sub-virtual model. This method can improve the registration efficiency.
[0142] In the application of surgical navigation, the point cloud data of the face, abdomen, and chest is scanned using a depth camera, and then the obtained point cloud data is registered with the virtual model of the target object corresponding to the reference object obtained by the embodiments of the present application using registration algorithms such as ICP, which can combine the preoperative scan images and tumor targets, etc. with the real scene to achieve the application of AR (Augmented Reality) surgical navigation.
[0143] Exemplarily, the registration process is as follows Figure 10 shown. Obtain the image 1001 to be processed, and then use the image processing method provided in the embodiments of the present application to obtain the image 1002 of the target object; based on the image 1002 of the target object, use the grid generation technology to construct the virtual model 1003 of the target object. Use a point cloud scanner / depth camera, etc. to scan the reference object from a certain orientation to obtain the point cloud data 1004. Use the registration algorithm to register the point cloud data 1004 with the sub-virtual model obtained by trimming the virtual model of the target object.
[0144] Image annotation means that first, an interested region is specified in the image manually, and then the image of the interested region is annotated. There are many natural images and medical images in the image annotation scenario. It is often necessary to select an interested region and then perform local threshold segmentation to obtain an initial segmentation. Since there are holes inside the initial segmentation, it is often necessary to fill the holes. The image processing method provided in the embodiments of the present application can also be applied in the image annotation project as an intelligent solution for hole filling.
[0145] In an exemplary embodiment, in the application scenario of image annotation, the image to be processed is the image of the interested region in the original image, and the image of the target object has the same size as the image to be processed. It should be noted that the image of the interested region in the original image is an image including a sub-image of the reference object. In this case, after obtaining the image of the target object, it further includes: annotating the original image according to the image of the target object. By annotating the original image according to the image of the target object, the region where the target object corresponding to the reference object is located can be annotated in the interested region of the original image. In an exemplary embodiment, the method of annotating the original image according to the image of the target object is: copying the image of the target object to the image of the interested region in the original image, so as to realize the annotation of the original image using the image of the target object.
[0146] In an exemplary embodiment, the computer device has a display function. After annotating the original image according to the image of the target object, the annotated original image is displayed to facilitate visually viewing the region where the target object is annotated in the interested region of the original image.
[0147] Exemplarily, the image annotation process is as follows Figure 11As shown in the figure, obtain the original image, manually select the region of interest in the original image; crop out the image of the region of interest from the original image, and use the image of the region of interest as the image to be processed 1101; perform binary processing on the image to be processed 1101 according to the contour threshold to obtain the threshold segmentation image 1102; according to the image processing method provided by the embodiments of the present application, obtain the image 1103 of the target object; copy the image 1103 of the target object into the original image to implement the annotation of the original image with the image of the target object.
[0148] Based on the image processing method provided by the embodiments of the present application, it is possible to use a convolutional neural network to finely segment the details of reference objects in medical images such as CT and MRI, and automatically remove the sub-images of the examination table in the image, as well as foreign objects such as infusion tubes and bandages attached to the contours of the reference objects, to form a smooth contour structure for improving the accuracy of registration algorithms based on various human body parts such as the face, head, and chest, thereby helping to improve the surgical navigation positioning accuracy.
[0149] The methods provided by the embodiments of the present application include but are not limited to the following applications: 1. Remove the sub-images of the examination table in CT images and improve the accuracy of rigid and non-rigid registration based on images; 2. Automatically segment fine faces and reconstruct virtual models of faces; 3. Automatically extract effective regions in images such as CT and MRI; 4. Remove foreign objects attached to the human face in the image; 5. Improve the registration accuracy of virtual models and point cloud data; 6. In the surgical navigation preoperative planning system, it is used to extract target objects and registration and positioning algorithms; 7. Automatically fill holes in the threshold segmentation image, replacing traditional hole filling algorithms, and can be used in some image annotation software such as ITK-SNAP, etc.
[0150] The technical solution provided by the embodiments of the present application calls the target image processing model to automatically obtain the image of the target object. The process of obtaining the target image does not require manual participation, and the efficiency of obtaining the target image is relatively high. In addition, the target image processing model is trained based on the labels of the pixel points corresponding to the sample images that indicate that the corresponding object category in the sample images is the target object, and has the function of accurately identifying the pixel points whose corresponding object category in the image is the target object. The reliability of the first probability map obtained by calling the target image processing model is relatively high, so that the accuracy of the finally obtained image of the target object is relatively high.
[0151] Based on the above Figure 1 shown implementation environment, the embodiments of the present application provide a training method for an image processing model. The training method of the image processing model is executed by a computer device, which can be a server 12 or a terminal 11. The embodiments of the present application do not limit this. As Figure 12As shown, the method for training an image processing model provided by an embodiment of the present application includes the following steps 1201 to 1203.
[0152] In step 1201, a sample image of a reference object and a label corresponding to the sample image are obtained. The label corresponding to the sample image is used to indicate a first sample pixel point in the sample image, and the object category corresponding to the first sample pixel point is a target object that meets the reference condition corresponding to the reference object.
[0153] The sample image of the reference object refers to the image required for training the initial image processing model. Exemplarily, the sample image of the reference object includes a sub-image of the reference object. Exemplarily, the sample image and Figure 2 the to-be-processed image in the shown embodiment are images of the same type and the same size, so as to ensure the processing effect of the trained image processing model on the to-be-processed image. It should be noted that the sample image mentioned in the embodiment of the present application refers to the sample image based on which the initial image processing model is trained once. The number of sample images can be one or multiple, and the embodiment of the present application does not limit this. Exemplarily, the number of sample images is multiple to ensure the model training effect.
[0154] In an exemplary embodiment, the manner in which the computer device obtains the sample image of the reference object includes, but is not limited to: the computer device extracts the sample image of the reference object from an image library; the computer device uses the training image in a certain public dataset as the sample image of the reference object.
[0155] In an exemplary embodiment, the manner in which the computer device obtains the sample image of the reference object is: obtaining an initial image of the reference object; performing data augmentation on the initial image to obtain a sample image of a reference size. The initial image is extracted from an image library or uploaded manually, etc., and the embodiment of the present application does not limit this. The manner of performing data augmentation on the initial image includes at least one of, but is not limited to, random flipping, anisotropic random scaling of the image along different axes, random cropping and padding, etc. No matter which data augmentation method is used, as long as it can ensure that the image obtained after data augmentation is an image of the reference size, the image of the reference size is the sample image. The reference size is used to limit the size of the sample image, and the reference size is set according to experience or flexibly adjusted according to the application scenario. The embodiment of the present application does not limit this. Exemplarily, for the case where the sample image is a 3D image, the reference size is 192×192×160 (pixels); for the case where the sample image is a 2D image, the reference size is 192×192 (pixels).
[0156] In an exemplary embodiment, the preparation strategy of the sample image varies in different application scenarios. Exemplarily, in the application scenario of the registration task, sample images involving fewer scenarios and modalities can be obtained; in the application scenario of the image annotation task, since more complex scenarios need to be dealt with, the sample images usually include different local images of medical images in multiple common scenarios and multiple modalities. For the case where the sample image is a local image, the sample image can also correspond to an image indicating the source of the sample image and annotation information about the position in the source image.
[0157] The label corresponding to the sample image is used to provide supervision information for the training process of the model. The label corresponding to the sample image is used to indicate that the corresponding object category in the sample image is the first sample pixel point of the target object corresponding to the reference object. That is to say, according to the label corresponding to the sample image, it can be known which sample pixel point or points in the sample image correspond to the object category of the target object. Exemplarily, the object category corresponding to a sample pixel point in the sample image may include the target object, the reference object, and other objects, etc.
[0158] In an exemplary embodiment, the way to obtain the label corresponding to the sample image is as follows: perform binary processing on the sample image according to the contour threshold to obtain a sample threshold segmentation image; manually refine the sample threshold segmentation image to obtain a refined image, in which the sample pixel points corresponding to the object category of the target object and other pixel points corresponding to the object category not being the target object are displayed in different display methods; use the refined image as the label corresponding to the sample image. Taking the type of the reference object as a certain part of the human body as an example, since there may be foreign objects such as bandages, infusion tubes, and fixing frames attached to the contour of a certain part of the human body, the boundary of the threshold segmentation image obtained by binary processing according to the contour threshold may not be very clean. Manually use annotation software to remove foreign objects and fill internal holes to obtain a refined image, which is used as the gold standard for training the model.
[0159] In step 1202, call the initial image processing model to perform probability prediction on the representation image of the sample image to obtain a sample probability map, and the pixel value of the pixel point in the sample probability map is used to indicate the probability that the pixel point corresponds to the object category of the target object.
[0160] After obtaining the sample image, further obtain the representation image of the sample image, so as to call the initial image processing model to perform probability prediction on the representation image of the sample image to obtain a sample probability map, and the pixel value of the pixel point in the sample probability map is used to indicate the probability that the pixel point corresponds to the object category of the target object.
[0161] The initial image processing model is an image processing model to be trained. For the introduction of the model structure of the initial image processing model, seeFigure 2 The introduction of the model structure of the target image processing model in the illustrated embodiment will not be elaborated here. The process of obtaining the representative image of the sample image and calling the initial image processing model to perform probability prediction on the representative image of the sample image to obtain the sample probability map is the same as that of Figure 2 step 202 in the illustrated embodiment, which will not be elaborated here.
[0162] In an exemplary embodiment, different from Figure 2 step 202 in the illustrated embodiment, the contour threshold for binarizing the sample image is a threshold randomly selected within the reference threshold range corresponding to the contour, so as to improve the diversity of the input information of the model. For example, the reference threshold range is [-490, 290], the contour threshold for binarizing the image to be processed is a fixed threshold within this reference threshold range (e.g., -390), but the contour threshold for binarizing the sample image is a random threshold within this reference threshold range. For the case where the number of sample images is multiple, the contour thresholds for binarizing different sample images may be the same or different.
[0163] In step 1203, based on the sample probability map and the label corresponding to the sample image, the initial image processing model is trained to obtain the target image processing model.
[0164] After obtaining the sample probability map, the initial image processing model is trained based on the sample probability map and the label corresponding to the sample image to obtain the trained target image processing model. The pixel value of the pixel point in the sample probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object, and the label corresponding to the sample image is used to indicate the pixel points in the sample image whose corresponding object category is the target object. Training the initial image processing model based on the sample probability map and the label corresponding to the sample image helps the model improve its processing performance by learning the difference between the sample probability map and the label corresponding to the sample image.
[0165] In a possible implementation manner, the process of training the initial image processing model based on the sample probability map and the label corresponding to the sample image is: based on the sample probability map and the label corresponding to the sample image, obtain the loss function; use the loss function to train the initial image processing model.
[0166] The number of sample images is one or more. For the case where the number of sample images is one, the number of sample probability maps is one, and the loss function is directly obtained based on the sample probability map and the label corresponding to the sample image. For the case where the number of sample images is multiple, the number of sample probability maps is also multiple. In this case, the way to obtain the loss function is as follows: for each sample probability map and the label corresponding to the corresponding sample image, a sub-loss function is obtained; the average value of all the obtained sub-loss functions is used as the target loss function. The way to obtain a sub-loss function when the number of sample images is multiple is the same as the way to obtain the loss function when the number of sample images is one. In the embodiments of the present application, the case where the number of sample images is one is taken as an example for illustration.
[0167] In an exemplary embodiment, the size of the sample probability map may be the same as the size of the sample image, or may be different from the size of the sample image. If the size of the sample probability map is the same as the size of the sample image, the loss function is directly obtained based on the difference between the sample probability map and the label corresponding to the sample image; if the size of the sample probability map is different from the size of the sample image, the sample probability map needs to be sampled into a probability map with the same size as the sample image first, and then the loss function is obtained based on the difference between the sampled probability map and the label corresponding to the sample image.
[0168] The embodiments of the present application do not limit the way to obtain the loss function based on the difference between the probability map with the same size as the sample image and the label corresponding to the sample image. Exemplarily, based on the difference between the probability map with the same size as the sample image and the label corresponding to the sample image, a cross-entropy loss function is obtained; Exemplarily, based on the difference between the probability map with the same size as the sample image and the label corresponding to the sample image, a dice loss function is obtained.
[0169] Exemplarily, the process of obtaining the dice loss function is implemented based on Formula 2:
[0170]
[0171] where, L dice represents the dice loss function; X represents the probability map with the same size as the sample image; Y represents the label corresponding to the sample image.
[0172] After obtaining the loss function, the initial image processing model is trained using the target loss function to obtain the target image processing model. In an exemplary embodiment, the process of training the initial image processing model using the loss function is an iterative process: the model parameters of the initial image processing model are updated backward using the loss function; each time the parameters are updated, it is determined whether the training process meets the training termination condition; if the training process meets the training termination condition, the iterative process is stopped, and the trained model is used as the trained target image processing model.
[0173] If the training process does not meet the training termination condition, a new loss function is obtained according to the methods of steps 1201 to 1203, and the parameters of the image processing model are updated backward using the new loss function. And so on, until the training process meets the training termination condition, and the trained target image processing model is obtained. It should be noted that in the process of obtaining the new loss function according to the methods of steps 1201 to 1203, the sample images used may or may not change, and this application embodiment does not limit this.
[0174] The embodiment of the present application uses a convolutional neural network to automatically segment the target object that completely fills the holes. The core idea is that in medical images, most anatomical structures have obvious boundaries in the image segmentation. The boundary of the threshold segmentation image according to the contour threshold is very fine. It's just that there are a large number of holes in the interior of the threshold segmentation image that need to be filled. The threshold segmentation image is used as part of the input of the image processing model, and the image processing model is allowed to learn the following objectives: retain the fine boundary; fill the holes in the threshold segmentation image; remove the false positive segmentation regions irrelevant to the task. Thus, an image of the target object can be obtained, where there are no foreign objects attached to the contour corresponding to the reference object and no holes inside.
[0175] The technical solution provided by the embodiment of the present application trains the initial image processing model based on the label corresponding to the sample image, where the label corresponding to the sample image is used to indicate the pixel points in the sample image whose corresponding object category is the target object. Such a training process can improve the recognition accuracy of the image processing model for the pixel points in the image whose corresponding object category is the target object, and thus train a target image processing model with the function of accurately recognizing the pixel points in the image whose corresponding object category is the target object, laying a foundation for the process of realizing the automatic acquisition of the image of the target object by calling the target image processing model to improve the efficiency of acquiring the image of the target object and the accuracy of the acquired image of the target object.
[0176] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be introduced. In an exemplary embodiment, the image to be processed is a head medical image, the type of the reference object is the head of a human body, the reference object is a head with no foreign objects attached to its contour and having holes inside, or a head with foreign objects attached to its contour and having holes inside, and the target object is a virtual head with no foreign objects attached to its contour and no holes inside. In this case, the process of image processing includes the following steps 1 to 3:
[0177] Step 1: Obtain the head medical image and the target image processing model.
[0178] Step 2: Invoke the target image processing model to perform probability prediction on the representation image of the head medical image, and obtain a first probability map. The pixel value of the pixel point in the first probability map is used to indicate the probability that the object category corresponding to the pixel point is a virtual head with no foreign objects attached to its contour and no holes inside.
[0179] Step 3: Based on the first probability map, obtain the image of the virtual head with no foreign objects attached to its contour and no holes inside.
[0180] The implementation principles of the above steps 1 to 3 are the same as those of steps 201 to 203 in the embodiment shown in Figure 2 and will not be elaborated here.
[0181] Refer to Figure 13 , the embodiments of the present application provide an image processing device, and the device includes:
[0182] A first acquisition unit 1301, configured to acquire the image to be processed of the reference object and the target image processing model. The target image processing model is trained based on the sample image of the reference object and the label corresponding to the sample image. The label corresponding to the sample image is used to indicate the first sample pixel point in the sample image, and the object category corresponding to the first sample pixel point is the target object that meets the reference conditions corresponding to the reference object;
[0183] A first processing unit 1302, configured to invoke the target image processing model to perform probability prediction on the representation image of the image to be processed, and obtain a first probability map. The pixel value of the pixel point in the first probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object;
[0184] A second acquisition unit 1303, configured to acquire the image of the target object based on the first probability map.
[0185] In a possible implementation manner, the device further includes:
[0186] A second processing unit, configured to perform binaryzation processing on the image to be processed according to the contour threshold corresponding to the contour of the reference object, and obtain a threshold segmentation image corresponding to the image to be processed;
[0187] A third acquisition unit, configured to segment an image based on a threshold to obtain a representation image of the image to be processed.
[0188] In a possible implementation, the third acquisition unit is configured to obtain an auxiliary image corresponding to the image to be processed; splice the threshold-segmented image and the auxiliary image to obtain a representation image.
[0189] In a possible implementation, the auxiliary image includes at least one of a first image and a second image. The first image is obtained by processing the image to be processed according to an imaging value range corresponding to the contour of a reference object, and the second image is obtained by processing the image to be processed according to an imaging value range corresponding to a reference element inside the reference object.
[0190] In a possible implementation, the apparatus further includes:
[0191] A determination unit, configured to determine a reference threshold range corresponding to the contour of the reference object; determine a contour threshold within the reference threshold range.
[0192] In a possible implementation, the second acquisition unit 1303 is configured to obtain a second probability map having the same size as the image to be processed based on the first probability map; perform binarization processing on the second probability map according to a probability threshold to obtain an image of the target object.
[0193] In a possible implementation, the apparatus further includes:
[0194] A registration unit, configured to construct a virtual model of the target object based on the image of the target object; obtain point cloud data of the reference object, where the point cloud data is obtained by scanning the reference object; register the point cloud data with the virtual model and display the registration result.
[0195] In a possible implementation, the image to be processed is an image of an interest region in an original image, and the image of the target object has the same size as the image to be processed. The apparatus further includes:
[0196] A labeling unit, configured to label the original image according to the image of the target object and display the labeled original image.
[0197] In a possible implementation, the image to be processed is a head medical image, the reference object is a head with no foreign object attached to the contour and having a hole inside, or a head with a foreign object attached to the contour and having a hole inside, and the target object meeting the reference condition is a head meeting at least one of the conditions of having no foreign object attached to the contour or having no hole inside.
[0198] The technical solution provided by the embodiment of the present application calls the target image processing model to automatically obtain the image of the target object. The process of obtaining the target image does not require manual intervention, and the efficiency of obtaining the target image is relatively high. In addition, the target image processing model is trained based on the labels of the pixel points corresponding to the sample images that indicate that the corresponding object category in the sample images is the target object, and has the function of accurately identifying the pixel points whose corresponding object category in the image is the target object. The reliability of the first probability map obtained by calling the target image processing model is relatively high, so that the accuracy of the finally obtained image of the target object is relatively high.
[0199] See Figure 14 , the embodiment of the present application provides a training device for an image processing model, and the device includes:
[0200] An acquisition unit 1401, configured to acquire a sample image of a reference object and a label corresponding to the sample image, where the label corresponding to the sample image is used to indicate a first sample pixel point in the sample image, and the object category corresponding to the first sample pixel point is the target object that satisfies the reference condition corresponding to the reference object;
[0201] A processing unit 1402, configured to call an initial image processing model to perform probability prediction on the representation image of the sample image to obtain a sample probability map, where the pixel value of the pixel point in the sample probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object;
[0202] A training unit 1403, configured to train the initial image processing model based on the sample probability map and the label corresponding to the sample image to obtain a target image processing model.
[0203] In a possible implementation manner, the acquisition unit 1401 is configured to acquire an initial image of a reference object; perform data augmentation on the initial image to obtain a sample image of a reference size.
[0204] The technical solution provided by the embodiment of the present application trains the initial image processing model based on the label corresponding to the sample image, where the label corresponding to the sample image is used to indicate the pixel points whose corresponding object category in the sample image is the target object. Such a training process can improve the recognition accuracy of the image processing model for the pixel points whose corresponding object category in the image is the target object, so as to train a target image processing model with the function of accurately identifying the pixel points whose corresponding object category in the image is the target object, laying a foundation for the process of realizing calling the target image processing model to automatically obtain the image of the target object to improve the efficiency of obtaining the image of the target object and the accuracy of the obtained image of the target object.
[0205] It should be noted that when the device provided in the above embodiment realizes its functions, only the division of the above functional units is used for illustration. In actual applications, the above functions can be allocated to different functional units according to needs, that is, the internal structure of the device is divided into different functional units to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.
[0206] In an exemplary embodiment, a computer device is also provided. The computer device includes a processor and a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by one or more processors so that the computer device realizes any one of the above image processing methods or the training method of the image processing model. The computer device can be a terminal or a server, and the embodiments of the present application do not limit this. Next, the structures of the terminal and the server will be introduced respectively.
[0207] Figure 15 This is a schematic structural diagram of a terminal provided by an embodiment of the present application. Exemplarily, the terminal can be: a PC, a mobile phone, a smart phone, a PDA, a wearable device, a PPC, a tablet computer, a smart car machine, a smart TV, a smart speaker, a vehicle-mounted terminal, etc. The terminal may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0208] Generally, the terminal includes: a processor 1501 and a memory 1502.
[0209] The processor 1501 may include one or more processing cores. The processor 1501 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit, central processor); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 1501 may be integrated with a GPU, and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1501 may also include an AI (Artificial Intelligence, artificial intelligence) processor, and the AI processor is used to process computing operations related to machine learning.
[0210] The memory 1502 may include one or more computer-readable storage media, which may be non-transitory. The memory 1502 may further include high-speed random access memory, as well as non-volatile memory, such as one or more magnetic disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1502 is used to store at least one instruction for being executed by the processor 1501, so that the terminal implements the image processing method or the training method of the image processing model provided in the method embodiments of the present application.
[0211] In some embodiments, the terminal may further optionally include: a peripheral device interface 1503 and at least one peripheral device. The processor 1501, the memory 1502, and the peripheral device interface 1503 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1503 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1504, a display screen 1505, a camera assembly 1506, an audio circuit 1507, and a power supply 1509.
[0212] The peripheral device interface 1503 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1501 and the memory 1502. The radio frequency circuit 1504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1504 communicates with a communication network and other communication devices through electromagnetic signals. The display screen 1505 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. The camera assembly 1506 is used to collect images or videos.
[0213] The audio circuit 1507 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1501 for processing, or input to the radio frequency circuit 1504 to implement voice communication. The speaker is used to convert electrical signals from the processor 1501 or the radio frequency circuit 1504 into sound waves. The power supply 1509 is used to supply power to each component in the terminal. The power supply 1509 may be alternating current, direct current, a disposable battery, or a rechargeable battery.
[0214] In some embodiments, the terminal further includes one or more sensors 1510. The one or more sensors 1510 include, but are not limited to: an acceleration sensor 1511, a gyroscope sensor 1512, a pressure sensor 1513, an optical sensor 1515, and a proximity sensor 1516.
[0215] The acceleration sensor 1511 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established by the terminal. The gyroscope sensor 1512 can detect the body direction and rotation angle of the terminal. The gyroscope sensor 1512 can cooperate with the acceleration sensor 1511 to collect the 3D actions of the user on the terminal. The pressure sensor 1513 can be disposed on the side frame of the terminal and / or the lower layer of the display screen 1505. When the pressure sensor 1513 is disposed on the side frame of the terminal, it can detect the holding signal of the user on the terminal, and the processor 1501 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1513. When the pressure sensor 1513 is disposed on the lower layer of the display screen 1505, the processor 1501 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1505.
[0216] The optical sensor 1515 is used to collect the ambient light intensity. The proximity sensor 1516, also known as the distance sensor, is usually disposed on the front panel of the terminal. The proximity sensor 1516 is used to collect the distance between the user and the front of the terminal.
[0217] Those skilled in the art can understand that Figure 15 the structure shown in
[0218] Figure 16 is a schematic structural diagram of a server provided by an embodiment of the present application. The server may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 1601 and one or more memories 1602. Among them, at least one computer program is stored in the one or more memories 1602, and the at least one computer program is loaded and executed by the one or more processors 1601 so that the server implements the image processing method or the training method of the image processing model provided by each of the above method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server may also include other components for implementing the functions of the device, which will not be elaborated here.
[0219] In an exemplary embodiment, a computer-readable storage medium is further provided. At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor of a computer device so that the computer implements any one of the above image processing methods or the training method of the image processing model.
[0220] In one possible implementation, the above computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, optical data storage device, etc.
[0221] In an exemplary embodiment, there is also provided a computer program product, which includes a computer program or computer instructions. The computer program or computer instructions are loaded and executed by a processor to enable a computer to implement any of the above image processing methods or the training method of an image processing model.
[0222] It should be noted that the terms "first", "second", etc. in this application are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. The embodiments described in the above exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0223] It should be understood that the term "plurality" as mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0224] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An image processing method, characterized in that, The method includes: Obtaining a to-be-processed image of a reference object and a target image processing model, where the target image processing model is trained based on a sample image of the reference object and a label corresponding to the sample image, the label corresponding to the sample image is used to indicate a first sample pixel point in the sample image, the object category corresponding to the first sample pixel point is a target object, the reference object satisfies at least one of having a foreign object attached to its contour or having a hole inside, and the target object is an object with no foreign object attached to its corresponding contour and no hole inside the reference object; Invoking the target image processing model to perform probability prediction on the representation image of the to-be-processed image, obtaining a first probability map, where the pixel value of a pixel point in the first probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object; Based on the first probability map, obtaining an image of the target object.
2. The method according to claim 1, characterized in that, Before invoking the target image processing model to perform probability prediction on the representation image of the to-be-processed image, the method further includes: Performing binarization processing on the to-be-processed image according to a contour threshold corresponding to the contour of the reference object, obtaining a threshold segmentation image corresponding to the to-be-processed image; Based on the threshold segmentation image, obtaining a representation image of the to-be-processed image.
3. The method according to claim 2, wherein The obtaining a representation image of the to-be-processed image based on the threshold segmentation image includes: Obtaining an auxiliary image corresponding to the to-be-processed image; Stitching the threshold segmentation image and the auxiliary image to obtain the representation image.
4. The method according to claim 3, wherein The auxiliary image includes at least one of a first image and a second image, the first image is obtained by processing the to-be-processed image according to an imaging value range corresponding to the contour of the reference object, and the second image is obtained by processing the to-be-processed image according to an imaging value range corresponding to a reference element inside the reference object.
5. The method according to any one of claims 2-4, characterized in that Before performing binarization processing on the to-be-processed image according to a contour threshold corresponding to the contour of the reference object to obtain a threshold segmentation image corresponding to the to-be-processed image, the method further includes: Determining a reference threshold range corresponding to the contour of the reference object; Determining the contour threshold within the reference threshold range.
6. The method according to any one of claims 1-4, characterized in that, The obtaining an image of the target object based on the first probability map includes: Based on the first probability map, obtaining a second probability map having the same size as the to-be-processed image; Performing binarization processing on the second probability map according to a probability threshold to obtain an image of the target object.
7. The method according to any one of claims 1-4, characterized in that After obtaining an image of the target object based on the first probability map, the method further includes: Based on the image of the target object, constructing a virtual model of the target object; Obtaining point cloud data of the reference object, where the point cloud data is obtained by scanning the reference object; Registering the point cloud data with the virtual model and displaying the registration result.
8. The method according to any one of claims 1-4, characterized in that, The image to be processed is an image of the region of interest in the original image. The image of the target object has the same size as the image to be processed. After obtaining the image of the target object based on the first probability map, the method further includes: Annotating the original image according to the image of the target object, and displaying the annotated original image.
9. The method according to any one of claims 1-4, characterized in that, The image to be processed is a head medical image. The reference object is a head with no foreign object attached to the contour and having holes inside, or a head with a foreign object attached to the contour and having holes inside. The target object is a head with no foreign object attached to the contour and no holes inside.
10. A method for training an image processing model, characterized in that, The method includes: Obtaining a sample image of a reference object and a label corresponding to the sample image. The label corresponding to the sample image is used to indicate a first sample pixel point in the sample image. The object category corresponding to the first sample pixel point is the target object. The reference object satisfies at least one of having a foreign object attached to the contour or having holes inside. The target object is an object corresponding to the reference object with no foreign object attached to the contour and no holes inside. Invoking an initial image processing model to perform probability prediction on the representation image of the sample image, obtaining a sample probability map. The pixel value of a pixel point in the sample probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object. Training the initial image processing model based on the sample probability map and the label corresponding to the sample image to obtain a target image processing model.
11. The method according to claim 10, characterized in that, The obtaining of the sample image of the reference object includes: Obtaining an initial image of the reference object; Performing data augmentation on the initial image to obtain the sample image of the reference size.
12. An image processing apparatus, characterized in that, The apparatus includes: A first obtaining unit, configured to obtain an image to be processed of a reference object and a target image processing model. The target image processing model is trained based on a sample image of the reference object and a label corresponding to the sample image. The label corresponding to the sample image is used to indicate a first sample pixel point in the sample image. The object category corresponding to the first sample pixel point is the target object. The reference object satisfies at least one of having a foreign object attached to the contour or having holes inside. The target object is an object corresponding to the reference object with no foreign object attached to the contour and no holes inside. A first processing unit, configured to invoke the target image processing model to perform probability prediction on the representation image of the image to be processed, obtaining a first probability map. The pixel value of a pixel point in the first probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object. A second obtaining unit, configured to obtain an image of the target object based on the first probability map.
13. The device according to claim 12, wherein, The apparatus further includes: A second processing unit, configured to perform binarization processing on the image to be processed according to a contour threshold corresponding to the contour of the reference object, obtaining a threshold segmentation image corresponding to the image to be processed. A third obtaining unit, configured to obtain a representation image of the image to be processed based on the threshold segmentation image.
14. The device according to claim 13, characterized in that, The third acquisition unit is configured to acquire an auxiliary image corresponding to the image to be processed; splice the threshold segmentation image and the auxiliary image to obtain the characterization image.
15. The device according to claim 14, wherein The auxiliary image includes at least one of a first image and a second image. The first image is obtained by processing the image to be processed according to an imaging value range corresponding to the contour of the reference object, and the second image is obtained by processing the image to be processed according to an imaging value range corresponding to a reference element inside the reference object.
16. The device according to any one of claims 13-15, characterized in that, The apparatus further includes: A determination unit, configured to determine a reference threshold range corresponding to the contour of the reference object; determine the contour threshold within the reference threshold range.
17. The device according to any one of claims 12 - 15, characterized in that, The second acquisition unit is configured to obtain a second probability map having the same size as the image to be processed based on the first probability map; perform binarization processing on the second probability map according to a probability threshold to obtain an image of the target object.
18. The device according to any one of claims 12 - 15, characterized in that, The apparatus further includes: A registration unit, configured to construct a virtual model of the target object based on the image of the target object; acquire point cloud data of the reference object, where the point cloud data is obtained by scanning the reference object; register the point cloud data with the virtual model and display the registration result.
19. The device according to any one of claims 12 - 15, characterized in that The image to be processed is an image of an interest region in an original image. The image of the target object has the same size as the image to be processed. The apparatus further includes: A labeling unit, configured to label the original image according to the image of the target object and display the labeled original image.
20. The device according to any one of claims 12-15, characterized in that, The image to be processed is a head medical image. The reference object is a head with no foreign object attached to its contour and having a hole inside, or a head with a foreign object attached to its contour and having a hole inside. The target object is a head with no foreign object attached to its contour and no hole inside.
21. A training device for an image processing model, characterized in that, The apparatus includes: An acquisition unit, configured to acquire a sample image of a reference object and a label corresponding to the sample image. The label corresponding to the sample image is used to indicate a first sample pixel point in the sample image, and the object category corresponding to the first sample pixel point is a target object. The reference object satisfies at least one of having a foreign object attached to its contour or having a hole inside. The target object is an object with no foreign object attached to the corresponding contour of the reference object and no hole inside. A processing unit, configured to call an initial image processing model to perform probability prediction on the characterization image of the sample image to obtain a sample probability map. The pixel value of a pixel point in the sample probability map is used to indicate the probability that the object category corresponding to the pixel point is the target object. A training unit, configured to train the initial image processing model based on the sample probability map and the label corresponding to the sample image to obtain a target image processing model.
22. The device according to claim 21, wherein The acquisition unit is configured to acquire an initial image of the reference object; perform data augmentation on the initial image to obtain the sample image of the reference size.
23. A computer device, characterized in that, The computer device includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor so that the computer device implements the image processing method according to any one of claims 1 to 9, or the training method of the image processing model according to any one of claims 10 to 11.
24. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor so that a computer implements the image processing method according to any one of claims 1 to 9, or the training method of the image processing model according to any one of claims 10 to 11.
25. A computer program product, characterized in that, The computer program product includes a computer program or computer instructions, and the computer program or the computer instructions are loaded and executed by a processor so that a computer implements the image processing method according to any one of claims 1 to 9, or the training method of the image processing model according to any one of claims 10 to 11.
Citation Information
Patent Citations
Medical image processing system and method
CN108010021A
Neural network-based improved RSG liver CT image interactive segmentation algorithm
CN111986216A