Image-based structure segmentation method and device, computer device and storage medium

CN116958155BActive Publication Date: 2026-09-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310255955.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-07
Publication Date
2026-09-18
Estimated Expiration
2043-03-07

AI Technical Summary

Technical Problem

[0004]但是,上述技术方案中的图像分割模型是通过特定的器官分割需求,在该特定的器官对应的数据集上进行训练的,也即是该图像分割模型仅能分割出一种器官,当面临多种器官的分割任务时,需要维护多种图像分割模型,消耗极大的硬件资源,并且分割效率低下

Benefits of technology

[0031] On the other hand, a computer program product is provided, including a computer program stored in a computer-readable storage medium, a processor of a computer device reading the computer program from the computer-readable storage medium, and the processor executing the computer program, causing the computer device to perform the image-based structural segmentation method provided in the above-described aspects or various alternative implementations of the aspects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958155B_ABST
    Figure CN116958155B_ABST
Patent Text Reader

Abstract

The application provides an image-based structure segmentation method and device, computer equipment and storage medium, and belongs to the technical field of image processing. The method comprises the following steps: acquiring a structure image and segmentation requirements; processing the segmentation requirements through a gating network of an image segmentation model to obtain a gating vector; calling multiple target sub-networks in the multiple sub-networks based on the gating vector; processing the structure image through the multiple target sub-networks of the image segmentation model to obtain outputs of the multiple target sub-networks; and predicting positions of the multiple target feature parts in the structure image based on the outputs of the multiple target sub-networks. The above method only needs to load the codes corresponding to the multiple target sub-networks into the memory for running, and can achieve the purpose of segmenting multiple feature parts at one time, thereby reducing the consumption of hardware resources and improving the segmentation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image-based structural segmentation method, apparatus, computer device, and storage medium. Background Technology

[0002] In everyday life, the structure of a target object contains multiple organs. To further study each organ and detect any abnormalities, the target object's structure needs to be segmented beforehand. This target object can be a human or an animal, etc.

[0003] In related technologies, the segmentation of any organ in a target object usually involves inputting the structural image of the target object into the image segmentation model corresponding to the organ, extracting the image features of the structural image through the image segmentation model, and then identifying the image features to determine the region where the organ is located in the structural image, so as to segment the organ from the structural image.

[0004] However, the image segmentation model in the above technical solution is trained on the dataset corresponding to a specific organ based on specific organ segmentation requirements. That is, the image segmentation model can only segment one type of organ. When faced with segmentation tasks of multiple organs, multiple image segmentation models need to be maintained, which consumes a lot of hardware resources and has low segmentation efficiency. Summary of the Invention

[0005] This application provides an image-based structural segmentation method, apparatus, computer device, and storage medium. During image segmentation, only the code corresponding to the required multiple target sub-networks needs to be loaded into memory and run to achieve the goal of segmenting multiple feature regions at once. This not only reduces hardware resource consumption and eliminates the need to maintain multiple image segmentation models, but also improves segmentation efficiency because it can segment multiple feature regions simultaneously. The technical solution is as follows:

[0006] On the one hand, an image-based structural segmentation method is provided, the method comprising:

[0007] Obtain a structural image and segmentation requirements, wherein the structural image contains multiple feature regions and the segmentation requirements are used to indicate the multiple target feature regions that need to be segmented;

[0008] The segmentation requirement is processed by the gating network of the image segmentation model to obtain the gating vector. The image segmentation model includes multiple sub-networks, each of which is used to identify its corresponding feature region. The gating vector is used to indicate which sub-network to call.

[0009] Based on the gating vector, multiple target sub-networks among the multiple sub-networks are invoked, and the multiple target sub-networks are used to identify the corresponding target feature parts;

[0010] The structured image is processed through the multiple target sub-networks of the image segmentation model to obtain the outputs of the multiple target sub-networks;

[0011] Based on the outputs of the multiple target sub-networks, the positions of the multiple target feature regions in the structured image are predicted, and the positions of the multiple target feature regions are used to represent the segmentation result.

[0012] On the other hand, an image-based structural segmentation apparatus is provided, the apparatus comprising:

[0013] The acquisition module is used to acquire a structural image and segmentation requirements. The structural image contains multiple feature regions, and the segmentation requirements are used to indicate the multiple target feature regions that need to be segmented.

[0014] The first processing module is used to process the segmentation requirement through the gating network of the image segmentation model to obtain a gating vector. The image segmentation model includes multiple sub-networks, each of which is used to identify its corresponding feature region. The gating vector is used to indicate the sub-network to be called.

[0015] The calling module is used to call multiple target sub-networks among the multiple sub-networks based on the gating vector, and the multiple target sub-networks are used to identify the corresponding target feature parts;

[0016] The second processing module is used to process the structured image through the multiple target sub-networks of the image segmentation model to obtain the output of the multiple target sub-networks;

[0017] The prediction module is used to predict the positions of the multiple target feature regions in the structured image based on the output of the multiple target sub-networks, and the positions of the multiple target feature regions are used to represent the segmentation result.

[0018] In some embodiments, the first processing module includes:

[0019] The first determining unit is used to determine the confidence scores of the plurality of sub-networks based on the structure image and the segmentation requirements. The confidence scores are used to represent the highest recognition accuracy of the corresponding sub-network for each target feature part.

[0020] The second determining unit is used to determine the gate vector in the image segmentation model based on the confidence scores of the plurality of sub-networks.

[0021] In some embodiments, the first determining unit is configured to select a predetermined number of target scores from the confidence scores of the plurality of sub-networks sorted from high to low; set the weights of the sub-networks corresponding to the predetermined number of target scores to a first value, and set the weights of the remaining sub-networks in the image segmentation model to a second value, wherein the first value indicates that the corresponding sub-network contributes to the recognition of the plurality of target feature regions, and the second value indicates that the corresponding sub-network does not contribute to the recognition of the plurality of target feature regions; and determine the gating vector based on the weights corresponding to the plurality of sub-networks in the image segmentation model.

[0022] In some embodiments, the prediction module includes:

[0023] The generation unit is used to generate multiple prediction kernels based on the structure image and the segmentation requirements. The multiple prediction kernels are used to identify their respective target feature parts, and the element values ​​in the multiple prediction kernels are different.

[0024] The third determining unit is used to determine the positions of the multiple target feature regions in the structure image based on the outputs of the multiple prediction kernels and the multiple target sub-networks.

[0025] In some embodiments, the generation unit is configured to determine the plurality of target feature regions to be segmented based on the structural image and the segmentation requirements; for any target feature region, determine the associated feature region of the target feature region, wherein there is an attribute association between the associated feature region and the target feature region; and determine the element value in the prediction kernel corresponding to the target feature region based on the target feature region and the associated feature region.

[0026] In some embodiments, the third determining unit is used to fuse the features output by the plurality of target sub-networks to obtain target features; and to convolve the plurality of prediction kernels on the target features to obtain the positions of the plurality of target feature parts in the structure image.

[0027] In some embodiments, the third determining unit is used to identify any target feature region by convolving the prediction kernel corresponding to the target feature region with the features output by the target sub-network corresponding to the target feature region to obtain the position of the target feature region in the structure image.

[0028] In some embodiments, the third determining unit is configured to process the structural image through a convolutional network to obtain overall features; fuse the features output by the multiple target sub-networks in the image segmentation model to obtain target features; fuse the overall features and the target features to obtain fused features; and determine the positions of the multiple target feature regions in the structural image based on the multiple prediction kernels and the fused features.

[0029] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded and executed by the processor to implement the image-based structural segmentation method in the embodiments of this application.

[0030] On the other hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to implement the image-based structural segmentation method as described in the embodiments of this application.

[0031] On the other hand, a computer program product is provided, including a computer program stored in a computer-readable storage medium, a processor of a computer device reading the computer program from the computer-readable storage medium, and the processor executing the computer program, causing the computer device to perform the image-based structural segmentation method provided in the above-described aspects or various alternative implementations of the aspects.

[0032] This application provides an image-based structural segmentation method. By inputting the structural image and segmentation requirements into an image segmentation model, multiple target sub-networks in the image segmentation model can be invoked according to the segmentation requirements. The structural image is then processed by the invoked target sub-networks to identify multiple target feature regions to be segmented. During the image segmentation process, only the code corresponding to the required multiple target sub-networks needs to be loaded into memory and run to achieve the goal of segmenting multiple feature regions at once. This not only reduces the consumption of hardware resources and eliminates the need to maintain multiple image segmentation models, but also improves segmentation efficiency because it can segment multiple feature regions at once. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1This is a schematic diagram of the implementation environment of an image-based structural segmentation method provided in an embodiment of this application;

[0035] Figure 2 This is a flowchart of an image-based structural segmentation method provided according to an embodiment of this application;

[0036] Figure 3 This is a flowchart of another image-based structural segmentation method provided according to an embodiment of this application;

[0037] Figure 4 This is a schematic diagram of a dynamic backbone network in an image segmentation model provided according to an embodiment of this application;

[0038] Figure 5 This is a block diagram of an image-based structure segmentation device according to an embodiment of this application;

[0039] Figure 6 This is a block diagram of another image-based structural segmentation device provided according to an embodiment of this application;

[0040] Figure 7 This is a structural block diagram of a terminal provided according to an embodiment of this application;

[0041] Figure 8 This is a schematic diagram of the structure of a server according to an embodiment of this application. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0043] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.

[0044] In this application, the term "at least one" means one or more, and "multiple" means two or more.

[0045] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the structural images of the target objects involved in this application were obtained with full authorization.

[0046] For ease of understanding, the terms used in this application are explained below.

[0047] Artificial Intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making capabilities. AI technology is a comprehensive discipline involving a wide range of fields, encompassing both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0048] With the research and advancement of artificial intelligence (AI) technology, AI is being researched and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with further technological development, AI will be applied in even more fields and play an increasingly important role. The image-based structural segmentation method provided in this application relates to the field of smart healthcare within AI technology.

[0049] Computer-aided Diagnosis (CAD) refers to the use of imaging, medical image processing technology, and other possible physiological and biochemical methods, combined with computer analysis and calculation, to assist in the detection of lesions and improve diagnostic accuracy. Currently, CAD technology primarily refers to computer-aided techniques based on medical imaging. Computer-aided detection is the foundation and essential stage of computer-aided diagnosis. CAD technology is often referred to as the doctor's "third eye," and the widespread application of CAD systems helps improve the sensitivity and specificity of doctors' diagnoses.

[0050] Human anatomical segmentation refers to the science of studying the morphology and structure of the human body, belonging to the morphology category of biological sciences. In the medical field, it is an important foundational course. Its task is to reveal the morphological and structural characteristics of various systems and organs in the human body, as well as the adjacency and connections between organs and structures. This lays the foundation for further study in subsequent basic and clinical medical courses. Automated segmentation of the human body accurately obtains spatial and size information of different organs, structures, and tissues, as well as their adjacency and connections, providing necessary physiological structural information for many computer-aided diagnostic applications.

[0051] The image-based structural segmentation method provided in this application can be executed by a computer device. In some embodiments, the computer device is a terminal or a server. The following describes the implementation environment of the image-based structural segmentation method provided in this application, using a computer device as a server as an example. Figure 1 This is a schematic diagram illustrating the implementation environment of an image-based structural segmentation method according to an embodiment of this application. See also... Figure 1 The implementation environment includes terminal 101 and server 102. Terminal 101 and server 102 can be connected directly or indirectly via wired or wireless communication, which is not limited herein.

[0052] In some embodiments, terminal 101 may be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, smart voice interaction device, smart home appliance, vehicle terminal, etc., but is not limited thereto. Terminal 101 has an application that supports image acquisition installed. The application may be a medical application, a communication application, or a multimedia application, and this application embodiment does not limit this. For example, taking a medical application as an example, terminal 101 can acquire a structural image of a target object, and then send the structural image of the target object to server 102. Server 102 uses an image segmentation model to predict the position of feature parts in the structural image of the target object, thereby achieving the purpose of segmenting the structure of the target object.

[0053] Those skilled in the art will understand that the number of terminals described above can be more or less. For example, there may be only one terminal, or there may be dozens or hundreds of terminals, or even more. This application does not limit the number of terminals or the type of device.

[0054] In some embodiments, server 102 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), big data, and artificial intelligence platforms. Server 102 is used to provide background services for applications that support image acquisition. In some embodiments, server 102 undertakes the main computing work, and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work, and terminal 101 undertakes the main computing work; or, server 102 and terminal 101 collaborate on computing using a distributed computing architecture.

[0055] Figure 2 This is a flowchart of an image-based structural segmentation method provided according to an embodiment of this application. See also... Figure 2 In this embodiment, the method is described using a server as an example. The image-based structural segmentation method includes the following steps:

[0056] 201. The server obtains the structural image and segmentation requirements. The structural image contains multiple feature regions, and the segmentation requirements are used to indicate the multiple target feature regions that need to be segmented.

[0057] In this embodiment, the structural image is an image used to present the structure of a target object. The target object can be a human or an animal, and this embodiment does not limit this. Taking a human as an example, the structural image can be a CT (Computed Tomography) image, an MRI (Magnetic Resonance Imaging) image, or a WSI (Whole Slide Image), and this embodiment does not limit this. The structural image of the target object contains multiple feature parts of the target object. Feature parts can be organs or bones of the target object, and this embodiment does not limit this. Segmentation requirements can include identifiers (such as names or numbers) of the multiple target feature parts to be segmented, and can also include identifiers of subnetworks used to identify the multiple target feature parts, and this embodiment does not limit this.

[0058] 202. The server processes the segmentation requirements through the gating network of the image segmentation model to obtain the gating vector. The image segmentation model includes multiple sub-networks, each of which is used to identify its corresponding feature regions. The gating vector is used to indicate which sub-network to call.

[0059] In this embodiment, the image segmentation model includes a gating network and multiple sub-networks. The gating network can select some sub-networks for prediction based on the structural image of the target object to be segmented. The server processes the segmentation requirements through the gating network of the image segmentation model and outputs a gating vector, with the value in the gating vector indicating the multiple target sub-networks to be invoked.

[0060] 203. The server, based on the gating vector, calls multiple target subnetworks from multiple subnetworks of the image segmentation model. These multiple target subnetworks are used to identify the corresponding target feature parts.

[0061] In this embodiment, different sub-networks in the image segmentation model are used to identify different feature regions. During the image segmentation process, the server can call some sub-networks in the image segmentation model according to the segmentation requirements related to the image. That is, the server calls the target sub-networks corresponding to the multiple target feature regions to be segmented according to the segmentation requirements of the structural image, so that any target sub-network after being called can identify the corresponding target feature regions in the structural image. Here, calling can also be referred to as activation.

[0062] 204. The server processes the structured image through multiple target sub-networks of the image segmentation model to obtain the outputs of multiple target sub-networks.

[0063] In this embodiment, the server can input a structural image into multiple target sub-networks of an image segmentation model, and extract features from the structural image through each of the multiple target sub-networks. Specifically, for any target sub-network, the server extracts features from the target feature region corresponding to that target sub-network. The features of any target feature region can reflect information such as the structure and location of that target feature region; this embodiment does not impose any limitations on this.

[0064] 205. Based on the output of multiple target sub-networks, the server predicts the positions of multiple target feature regions in the structured image. The positions of the multiple target feature regions are used to represent the segmentation results.

[0065] In this embodiment, the server identifies multiple target feature regions based on the features of the structural images output by multiple target sub-networks, thereby obtaining the positions of the multiple target feature regions in the structural images. The identification results of different target feature regions can be displayed using different colors or different contours to indicate the positions of different target feature regions.

[0066] This application provides an image-based structural segmentation method. By inputting the structural image of the target object and the segmentation requirements into an image segmentation model, multiple target sub-networks in the image segmentation model can be called according to the segmentation requirements. The structural image of the target object is then processed by the called target sub-networks to identify multiple target feature regions to be segmented. During the image segmentation process, only the code corresponding to the required multiple target sub-networks needs to be loaded into memory and run to achieve the goal of segmenting multiple feature regions at once. This not only reduces the consumption of hardware resources and eliminates the need to maintain multiple image segmentation models, but also improves the segmentation efficiency because it can segment multiple feature regions at once.

[0067] Figure 3 This is a flowchart of another image-based structural segmentation method provided according to an embodiment of this application. See also... Figure 3 In this embodiment, the method is described using a server as an example. The image-based structural segmentation method includes the following steps:

[0068] 301. The server obtains the structural image and segmentation requirements. The structural image contains multiple feature regions, and the segmentation requirements are used to indicate the multiple target feature regions that need to be segmented.

[0069] In this embodiment, the structural image may include multiple feature regions of the target object, such as the heart, lungs, kidneys, pancreas, and stomach; this embodiment does not impose any limitations on this. The segmentation requirement can indicate multiple target feature regions to be segmented from the multiple feature regions of the target object. This embodiment does not impose any limitations on the number of target feature regions to be segmented. The segmentation requirement can be user-defined or generated based on the physiological data of the target object; this embodiment does not impose any limitations on this. The physiological data of the target object may include information such as the target object's symptoms, medications taken, and medical history; this embodiment does not impose any limitations on this. The server can obtain the structural image and segmentation requirement of the target object from the terminal, or it can obtain the structural image and segmentation requirement of the target object from the server's database; this embodiment does not impose any limitations on this.

[0070] In some embodiments, the server generates segmentation requirements based on the physiological data of the target object. Accordingly, the process of the server obtaining the segmentation requirements of the target object includes: the server acquiring the physiological data of the target object; then, the server performing feature recognition on the physiological data of the target object to obtain abnormal information about the target object. The abnormal information is used to indicate the presence of abnormal feature areas in the target object. Then, the server determines the segmentation requirements of the target object based on the abnormal information. The solution provided in this application embodiment, by performing feature recognition on the physiological data of the target object, enables more accurate identification of abnormal feature areas in the target object. The resulting segmentation requirements can then instruct subsequent segmentation of the abnormal feature areas, facilitating timely detection of abnormal feature areas and meeting user needs.

[0071] During the feature recognition process of the target object's physiological data, the server can perform text recognition on keywords within the physiological data. Keywords in the physiological data can be at least one of the following: the name of the medication taken by the target object, the name of the target object's symptoms, or the name of a disease in the medical history. The physiological data may also include images such as the target object's electrocardiogram (ECG), electrogastric electrogram (ECG), or electromyography (EMG). Furthermore, during the feature recognition process of the target object's physiological data, the server can also perform image recognition on the images within the physiological data.

[0072] For example, medical staff can use a terminal to capture structural images of a target subject and input the subject's physiological data. The terminal then sends the structural images and physiological data to a server. Based on information such as the subject's symptoms, medications, and medical history from the physiological data, the server determines the segmentation requirements for the structural images. If the subject has digestive system problems, the segmentation requirement might be for segmenting organs such as the esophagus, stomach, and intestines; if the subject has respiratory system problems, the segmentation requirement might be for segmenting organs such as the trachea and lungs. The terminal then sends the acquired structural images and segmentation requirements to the server, which performs the organ segmentation in the structural images.

[0073] 302. The server processes the segmentation requirements through the gating network of the image segmentation model to obtain the gating vector. The image segmentation model includes multiple sub-networks, each of which is used to identify its corresponding feature regions. The gating vector is used to indicate which sub-network to call.

[0074] In this embodiment, the image segmentation model can be constructed using a Mixture of Experts (MoE) system, and this embodiment does not impose any limitations on this. The image segmentation model includes multiple sub-networks and a gating network. Each sub-network is used to identify corresponding feature regions. That is, different sub-networks are used to identify different feature regions. Sub-networks can also be called expert networks. Sub-networks can be constructed using Feed Forward Networks (FFNs), and this embodiment does not impose any limitations on this. The gating network is used to interpret the predictions made by each expert network and helps determine which expert network to trust for a given structural image. That is, the gating network selects some sub-networks to make predictions for the structural image of the target object to be segmented. The gating network is a neural network and can be represented by a gating function. The gating network outputs the contribution that each sub-network should make when predicting the structural image. That is, based on the segmentation requirements of the structural image, the gating network outputs a gating vector to indicate which multiple target sub-networks to call. An image segmentation model can be viewed as a conditional computation model. During the image segmentation process, the image segmentation model only calls a few sub-networks, that is, a subset of the model.

[0075] In some embodiments, the server can determine the gate vector output by the gate network in the image segmentation model based on the confidence score of each sub-network in the image segmentation model when predicting the structured image. Accordingly, the process of the server determining the gate vector in the image segmentation model is as follows: the server's gate network of the image segmentation model processes the structured image and the segmentation requirements to determine the confidence scores of multiple sub-networks. Then, the server determines the gate vector in the image segmentation model based on the confidence scores of the multiple sub-networks. The confidence score represents the recognition performance of the corresponding sub-network for each target feature region. That is, for any given sub-network, the accuracy of its predictions for different target feature regions varies. For any target feature region, the prediction accuracy indicates the similarity between the predicted result and the actual result, and can be represented numerically. The server can weight the accuracy of multiple target feature regions to obtain a confidence score. This confidence score represents the recognition performance of the sub-network for multiple target feature regions. The server can also sum the accuracies of multiple target feature regions to obtain a confidence score, and this application embodiment does not limit this. The solution provided in this application embodiment determines the confidence scores of multiple sub-networks for the target feature regions to be segmented based on the structural image of the target object and the segmentation requirements. Since the confidence score can represent the recognition performance of the sub-network for multiple target feature regions, it can reflect the contribution that the sub-network can make to the prediction of the target feature regions. The gate vector determined based on the confidence score can be used to instruct the sub-network with a high confidence score to make predictions, which can improve the accuracy of segmentation. Furthermore, in the image segmentation process, only the code corresponding to the required multiple target sub-networks needs to be loaded into memory and run to achieve the purpose of segmenting multiple feature regions at once. This not only reduces the consumption of hardware resources and eliminates the need to maintain multiple image segmentation models, but also improves the segmentation efficiency because it can segment multiple feature regions at once.

[0076] The process by which the server determines the gating vector based on the confidence scores of multiple sub-networks is as follows: The server selects a predetermined number of target scores from the confidence scores of multiple sub-networks sorted from high to low. Then, the server sets the weights of the sub-networks corresponding to the predetermined number of target scores as a first value, and sets the weights of the remaining sub-networks in the image segmentation model as a second value. Then, the server determines the gating vector based on the weights corresponding to the multiple sub-networks in the image segmentation model. The first value indicates that the corresponding sub-network contributed to the recognition of multiple target feature regions. The second value indicates that the corresponding sub-network did not contribute to the recognition of multiple target feature regions. This embodiment does not limit the magnitude of the first and second values. The gating vector is a sparse vector. The solution provided in this application sets the weights of the sub-networks corresponding to the top-ranked target scores as a first value, and sets the weights of the remaining sub-networks as a second value. This allows the image segmentation model to make predictions based on the sub-networks with higher confidence scores, thereby achieving the segmentation of multiple target feature regions. This not only improves the accuracy of segmentation but also achieves the goal of segmenting multiple feature regions with a single image segmentation model, eliminating the need to maintain multiple image segmentation models, reducing hardware resource consumption, and enabling the segmentation of multiple feature regions at once, thus improving segmentation efficiency.

[0077] For example, the gating vector is [0,0,0,0,1,0,0,0,1,0,0,0,1]. Here, 1 represents the first value in the gating vector, indicating that the corresponding sub-network contributed to the recognition of multiple target feature regions. 0 represents the second value in the gating vector, indicating that the corresponding sub-network did not contribute to the recognition of multiple target feature regions. Most elements in the gating vector are zero, making it a sparse vector.

[0078] In some embodiments, the server can determine the gating vector using the following formula 1.

[0079] Formula 1:

[0080] g(x)=TopK(softmax(f(x)+∈))

[0081] wherein, g(x) is used to represent a gating vector; x is used to represent a structural image of a target object; f(x) is used to represent preliminary features extracted based on a segmentation requirement, the preliminary features can be extracted based on an attention mechanism or based on convolution, which is not limited in the embodiments of the present application, and f(x) is a linear transformation; ∈ is used to represent Gaussian noise applied to sub-networks; softmax(f(x)+∈) is used to represent confidence scores; TopK is used to select sub-networks that rank in the first preset K positions when confidence scores are sorted from high to low. When K<<E, most elements in g(x) are zero, and g(x) is a sparse vector.

[0082] 303, the server calls a plurality of target sub-networks among the plurality of sub-networks based on the gating vector, and the plurality of target sub-networks are configured to respectively identify corresponding target feature parts.

[0083] In the embodiment of the present application, the gating vector includes a plurality of values. The number of the plurality of values is equal to the number of sub-networks in the image segmentation model. The values in the gating vector can be divided into two categories, one is a first value and the other is a second value. Since the first value indicates that the corresponding sub-network contributes to the identification of the plurality of target feature parts, and the second value indicates that the corresponding sub-network does not contribute to the identification of the plurality of target feature parts, the server determines the sub-network corresponding to the first value as the target sub-network. During image segmentation, the server calls a plurality of target sub-networks, wherein calling can also be referred to as activation.

[0084] 304, the server processes the structural image through a plurality of target sub-networks of the image segmentation model to obtain outputs of the plurality of target sub-networks.

[0085] In the embodiment of the present application, the server inputs the structural image of the target object into a plurality of target sub-networks of the image segmentation model. For any target sub-network, the target sub-network can extract features from the structural image of the target object based on its own network parameters. The most prominent part of the extracted features is the feature of the target feature part corresponding to the target sub-network. During model training, since the input of the image segmentation model comes from different data sets, different scenarios, and data cases of different modalities, data cases from different segmentation tasks will dynamically and sparsely call part of sub-networks of the image segmentation model, so that different sub-networks have different network parameters trained according to different data sets, so that the trained sub-networks can specially segment feature parts of features, realize information separation and avoid conflict between information. Wherein, the data set used in the model training process can be LITS data set (a data set labeled with liver), KITS data set (a data set labeled with kidney), Pancreas data set (a data set labeled with pancreas), etc., which is not limited in the embodiments of the present application.

[0086] After the structural image is segmented by multiple target sub-networks, the server obtains the outputs of each target sub-network. The server can obtain the individual outputs of each target sub-network or the total output of the multiple target sub-networks; this embodiment does not impose any limitation on this.

[0087] Optionally, the server can obtain the total output of multiple target sub-networks. Accordingly, the server can calculate the total output of multiple target sub-networks using the following Formula 2.

[0088] Formula 2:

[0089]

[0090] Where MoE(x) represents the total output of multiple target subnetworks; i represents the i-th subnetwork; g(x) i Used to represent the gating vector; e(x) i The output of the i-th subnetwork is represented by E; the total number of subnetworks is represented by E; and the structure image of the target object is represented by x.

[0091] 305. Based on the structural image and segmentation requirements, the server generates multiple prediction kernels. These kernels are used to identify their respective target feature parts, and the element values ​​in the multiple prediction kernels are different.

[0092] In this embodiment, a lightweight dynamic segmentation task head is designed to assign specific prediction kernels to each segmentation task for segmenting specific feature regions or tumors. The dynamic segmentation task head comprises three stacked convolutional layers with 1×1×1 kernels. That is, the dynamic segmentation task head includes three convolutional layers, each with a 1×1×1 kernel. The kernel set of the dynamic segmentation task head is ω = {ω1, ω2, ω3}. This kernel set is dynamically generated by the filter prediction head based on feature region embeddings. In other words, the server dynamically generates multiple prediction kernels based on the structural image and segmentation requirements. For any target feature region indicated by the segmentation requirements, the server determines the element values ​​in the prediction kernel based on the target feature region in the structural image. The prediction kernel comprises the aforementioned three 1×1×1 convolutional kernels. The element values ​​in the prediction kernel refer to the parameters in the aforementioned three convolutional kernels. For any given convolutional kernel, during the subsequent prediction of target feature regions in the structural image, the kernel can convolve the features output by the target sub-network. In other words, the parameters in the convolution kernel are used as weights to weight the features output by the target subnetwork, and the resulting features can indicate the location of the target feature parts.

[0093] In some embodiments, the server can generate the prediction kernel using the following formula 3.

[0094] Formula 3:

[0095] ω=MLP(t)

[0096] Where ω represents the prediction kernel; t represents the target feature region in the structured image indicated by the segmentation requirement; and MLP represents the multilayer perceptron algorithm.

[0097] In the process of generating the prediction kernel, the server can determine the element values ​​in the prediction kernel based on the attribute information of the target feature parts; the server can also determine the element values ​​in the prediction kernel based on feature parts that are related to the target, and this application embodiment does not limit this. The attribute information of the target feature parts can be the structural attributes, positional attributes, or association attributes of the target feature parts with other feature parts, etc., and this application embodiment does not limit this.

[0098] In some embodiments, the server can determine the element values ​​in the prediction kernel based on feature regions associated with the target. Accordingly, the process of the server generating multiple prediction kernels is as follows: the server determines multiple target feature regions to be segmented based on the structural image and segmentation requirements. For any target feature region, the server determines its associated feature regions. Then, the server determines the element values ​​in the prediction kernel corresponding to the target feature region based on the target feature region and the associated feature regions. Here, there is an attribute association between the associated feature regions and the target feature regions. Attribute association can refer to location association, functional association, or structural association, etc., and this application embodiment does not impose any limitations on this. The solution provided in this application embodiment determines the prediction kernel for the target feature regions by using feature regions associated with the target. Since there are commonalities among the associated feature regions, segmentation tasks for some similar feature regions can help each other improve. Therefore, the generated prediction kernel can more accurately process the features of the target feature regions, thereby improving the prediction accuracy.

[0099] It should be noted that the process of generating multiple prediction kernels in step 305 can occur after step 304, or at any time after step 301 and before step 306; this embodiment of the application does not impose any restrictions on this. Optionally, the execution time of step 305 can be the same as the execution time of step 302. This step execution method allows the generated prediction kernels to be used for prediction when the outputs of multiple target sub-networks are obtained, thereby improving the image segmentation efficiency.

[0100] 306. The server determines the positions of multiple target feature regions in the structured image based on the outputs of multiple prediction kernels and multiple target sub-networks. The positions of the multiple target feature regions are used to represent the segmentation results.

[0101] In this embodiment, the server processes the outputs of multiple target sub-networks using multiple prediction kernels to predict the positions of multiple target feature parts in the structural image. The server can make predictions based on the total output of the multiple target sub-networks; alternatively, the server can make predictions separately for each target sub-network's output. This embodiment does not impose any limitations on this approach.

[0102] In some embodiments, the server can make predictions based on the total output of multiple target sub-networks. Accordingly, the process of the server predicting the positions of multiple target feature regions is as follows: the server fuses the features output by multiple target sub-networks to obtain target features. Then, the server convolves multiple prediction kernels with the target features respectively to obtain the positions of multiple target feature regions in the structured image. The method provided in this application, which uses the total output of multiple target sub-networks for prediction, enables the prediction of the positions of multiple target feature regions at once, improving segmentation efficiency.

[0103] In some embodiments, the server can predict the outputs of multiple target sub-networks separately. Accordingly, the process of the server predicting the positions of multiple target feature regions is as follows: for the identification of any target feature region, the server convolves the prediction kernel corresponding to the target feature region with the features output by the target sub-network corresponding to the target feature region to obtain the position of the target feature region in the structured image. The method provided in this application, which predicts the outputs of multiple target sub-networks separately, can improve the accuracy of segmentation.

[0104] In some embodiments, the server can also make predictions based on common information between different feature parts and information specific to the target feature parts. Accordingly, the process of the server predicting the positions of multiple target feature parts is as follows: the server processes the structural image through a convolutional network to obtain overall features. Then, the server fuses the features output by multiple target sub-networks in the image segmentation model to obtain target features. Then, the server fuses the overall features and target features to obtain fused features. Then, the server determines the positions of multiple target feature parts in the structural image based on multiple prediction kernels and the fused features. The method provided in this application extracts overall features of the structural image through a convolutional network to obtain common information between different feature parts in the structural image; then, by fusing the overall features and the features output by multiple target sub-networks, the fused features more accurately reflect the relationships between different feature parts in the structural image, thereby making the predicted positions of the target feature parts more accurate.

[0105] Overall, this technical solution employs a dynamic inference mechanism. The entire image segmentation model can be divided into two parts: a dynamic backbone network and a dynamic segmentation task head. For the dynamic backbone network, network pathways are sparsely invoked to obtain efficient, non-conflicting features. This corresponds to steps 301 to 304. Then, the dynamic segmentation task head generates a prediction kernel for the target task, predicting the location of the target object's feature parts in the structural image to achieve segmentation. This corresponds to steps 305 to 306. To further clarify this solution, it will be described again below with reference to the accompanying drawings. Figure 4 This is a schematic diagram of a dynamic backbone network in an image segmentation model provided according to an embodiment of this application. See also... Figure 4 The dynamic backbone network in this image segmentation model includes a multi-head attention mechanism layer and a MoE layer. A normalization layer precedes both the multi-head attention mechanism layer and the MoE layer. The server extracts initial features from the structured image through the multi-head attention mechanism layer. Then, the server performs residual processing on the extracted initial features and the features input to the normalization layer. The residual features then enter the next normalization layer for normalization. The server then processes the output of the previous normalization layer through the MoE layer to obtain the total output of multiple target sub-networks. Specifically, the MoE layer determines the gating vector of the gating network based on the features of the structured image output from the previous layer, thereby determining the multiple target sub-networks for image segmentation. Finally, the server predicts the outputs of the multiple target sub-networks using a dynamic segmentation task head to obtain the segmentation result of the structured image.

[0106] This application provides an image-based structural segmentation method. By inputting the structural image of the target object and the segmentation requirements into an image segmentation model, multiple target sub-networks in the image segmentation model can be called according to the segmentation requirements. The structural image of the target object is then processed by the called target sub-networks to identify multiple target feature regions to be segmented. This achieves the goal of segmenting multiple feature regions through a single image segmentation model, eliminating the need to maintain multiple image segmentation models, reducing hardware resource consumption, and improving segmentation efficiency by segmenting multiple feature regions at once.

[0107] Figure 5 This is a block diagram of an image-based structure segmentation apparatus according to an embodiment of this application. The image-based structure segmentation apparatus is used to perform the steps described above during image-based structure segmentation. (See also...) Figure 5 The image-based structural segmentation device includes: an acquisition module 501, a first processing module 502, a calling module 503, a second processing module 504, and a prediction module 505.

[0108] The acquisition module 501 is used to acquire the structural image and segmentation requirements. The structural image contains multiple feature regions, and the segmentation requirements are used to indicate the multiple target feature regions that need to be segmented.

[0109] The first processing module 502 is used to process the segmentation requirements through the gating network of the image segmentation model to obtain the gating vector. The image segmentation model includes multiple sub-networks, each of which is used to identify its corresponding feature parts. The gating vector is used to indicate the sub-network to be called.

[0110] Module 503 is used to call multiple target subnetworks from multiple subnetworks based on the gated vector. The multiple target subnetworks are used to identify the corresponding target feature parts.

[0111] The second processing module 504 is used to process the structured image through multiple target sub-networks of the image segmentation model to obtain the output of multiple target sub-networks;

[0112] The prediction module 505 is used to predict the positions of multiple target feature regions in the structured image based on the output of multiple target sub-networks. The positions of the multiple target feature regions are used to represent the segmentation result.

[0113] This application provides an image-based structural segmentation device. By inputting the structural image and segmentation requirements into an image segmentation model, multiple target sub-networks in the image segmentation model can be called according to the segmentation requirements. The structural image is then processed by the called target sub-networks to identify multiple target feature regions to be segmented. During the image segmentation process, only the code corresponding to the required multiple target sub-networks needs to be loaded into memory and run to achieve the goal of segmenting multiple feature regions at once. This not only reduces the consumption of hardware resources and eliminates the need to maintain multiple image segmentation models, but also improves segmentation efficiency because it can segment multiple feature regions at once.

[0114] In some embodiments, Figure 6 This is a block diagram of another image-based structural segmentation device provided according to an embodiment of this application. See also... Figure 6 The first processing module 502 includes:

[0115] The first determining unit 5021 is used to determine the confidence scores of multiple sub-networks based on the structural image and segmentation requirements. The confidence scores are used to represent the highest recognition accuracy of the corresponding sub-network for each target feature part.

[0116] The second determining unit 5022 is used to determine the gate vector in the image segmentation model based on the confidence scores of multiple sub-networks.

[0117] In some embodiments, see continue to see Figure 6The first determining unit 5021 is used to select a preset number of target scores from the confidence scores of multiple sub-networks sorted from high to low; set the weights of the sub-networks corresponding to the preset number of target scores to a first value, and set the weights of the remaining sub-networks in the image segmentation model to a second value. The first value is used to indicate that the corresponding sub-network contributes to the recognition of multiple target feature parts, and the second value is used to indicate that the corresponding sub-network does not contribute to the recognition of multiple target feature parts; and determine the gating vector based on the weights corresponding to the multiple sub-networks in the image segmentation model.

[0118] In some embodiments, see continue to see Figure 6 Prediction module 505 includes:

[0119] The generation unit 5051 is used to generate multiple prediction kernels based on the structural image and segmentation requirements. The multiple prediction kernels are used to identify their respective target feature parts, and the element values ​​in the multiple prediction kernels are different.

[0120] The third determining unit 5052 is used to determine the positions of multiple target feature parts in the structure image based on the outputs of multiple prediction kernels and multiple target sub-networks.

[0121] In some embodiments, see continue to see Figure 6 The generation unit 5051 is used to determine multiple target feature regions to be segmented based on the structural image and segmentation requirements; for any target feature region, determine the associated feature regions of the target feature region, and there is an attribute association between the associated feature regions and the target feature regions; and determine the element values ​​in the prediction kernel corresponding to the target feature region based on the target feature region and the associated feature regions.

[0122] In some embodiments, see continue to see Figure 6 The third determining unit 5052 is used to fuse the features output by multiple target sub-networks to obtain target features; and to convolve multiple prediction kernels on the target features to obtain the positions of multiple target feature parts in the structured image.

[0123] In some embodiments, see continue to see Figure 6 The third determining unit 5052 is used to identify any target feature part by convolving the prediction kernel corresponding to the target feature part with the features output by the target sub-network corresponding to the target feature part to obtain the position of the target feature part in the structure image.

[0124] In some embodiments, see continue to see Figure 6The third determining unit 5052 is used to process the structure image through a convolutional network to obtain overall features; fuse the features output by multiple target sub-networks in the image segmentation model to obtain target features; fuse the overall features and target features to obtain fused features; and determine the positions of multiple target feature parts in the structure image based on multiple prediction kernels and fused features.

[0125] It should be noted that the image-based structure segmentation device provided in the above embodiments is only illustrated by the division of the above functional modules when segmenting the structural image of the target object. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the image-based structure segmentation device and the image-based structure segmentation method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0126] In the embodiments of this application, the computer device can be configured as a terminal or a server. When the computer device is configured as a terminal, the terminal can act as the execution subject to implement the technical solutions provided in the embodiments of this application. When the computer device is configured as a server, the server can act as the execution subject to implement the technical solutions provided in the embodiments of this application. Alternatively, the technical solutions provided in this application can be implemented through the interaction between the terminal and the server. The embodiments of this application do not limit this.

[0127] Figure 7 This is a structural block diagram of a terminal 700 provided according to an embodiment of this application. The terminal 700 can be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 700 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0128] Typically, terminal 700 includes a processor 701 and a memory 702.

[0129] Processor 701 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 701 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 701 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 701 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 701 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0130] The memory 702 may include one or more computer-readable storage media, which may be non-transitory. The memory 702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 are used to store at least one computer program, which is executed by the processor 701 to implement the image-based structural segmentation method provided in the method embodiments of this application.

[0131] In some embodiments, the terminal 700 may also optionally include a peripheral device interface 703 and at least one peripheral device. The processor 701, memory 702, and peripheral device interface 703 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 703 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 704, a display screen 705, a camera assembly 706, an audio circuit 707, and a power supply 708.

[0132] Peripheral device interface 703 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 701 and memory 702. In some embodiments, processor 701, memory 702 and peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 701, memory 702 and peripheral device interface 703 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0133] The radio frequency (RF) circuit 704 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 704 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 704 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. In some embodiments, the RF circuit 704 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 704 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 704 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0134] Display screen 705 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 705 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 701 for processing. In this case, display screen 705 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 705, disposed on the front panel of terminal 700; in other embodiments, there may be at least two display screens 705, disposed on different surfaces of terminal 700 or in a folded design; in other embodiments, display screen 705 may be a flexible display screen, disposed on a curved or folded surface of terminal 700. Furthermore, display screen 705 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 705 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0135] The camera assembly 706 is used to acquire images or videos. In some embodiments, the camera assembly 706 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 706 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash is a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0136] The audio circuit 707 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 701 for processing, or input to the radio frequency circuit 704 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal 700. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert the electrical signals from the processor 701 or the radio frequency circuit 704 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 707 may also include a headphone jack.

[0137] Power supply 708 is used to power the various components in terminal 700. Power supply 708 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 708 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0138] In some embodiments, the terminal 700 further includes one or more sensors 709. The one or more sensors 709 include, but are not limited to: an accelerometer 710, a gyroscope 711, a pressure sensor 712, an optical sensor 713, and a proximity sensor 714.

[0139] Accelerometer 710 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal 700. For example, accelerometer 710 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 701 can control display screen 705 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 710. Accelerometer 710 can also be used for games or for acquiring user motion data.

[0140] The gyroscope sensor 711 can detect the orientation and rotation angle of the terminal 700. The gyroscope sensor 711, in conjunction with the accelerometer sensor 710, can collect 3D motion data from the user on the terminal 700. Based on the data collected by the gyroscope sensor 711, the processor 701 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0141] The pressure sensor 712 can be disposed on the side bezel of the terminal 700 and / or the lower layer of the display screen 705. When the pressure sensor 712 is disposed on the side bezel of the terminal 700, it can detect the user's grip signal on the terminal 700, and the processor 701 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 712. When the pressure sensor 712 is disposed on the lower layer of the display screen 705, the processor 701 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 705. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0142] An optical sensor 713 is used to collect ambient light intensity. In one embodiment, the processor 701 can control the display brightness of the display screen 705 based on the ambient light intensity collected by the optical sensor 713. Specifically, when the ambient light intensity is high, the display brightness of the display screen 705 is increased; when the ambient light intensity is low, the display brightness of the display screen 705 is decreased. In another embodiment, the processor 701 can also dynamically adjust the shooting parameters of the camera assembly 706 based on the ambient light intensity collected by the optical sensor 713.

[0143] The proximity sensor 714, also known as a distance sensor, is typically located on the front panel of the terminal 700. The proximity sensor 714 is used to detect the distance between the user and the front of the terminal 700. In one embodiment, when the proximity sensor 714 detects that the distance between the user and the front of the terminal 700 is gradually decreasing, the processor 701 controls the display screen 705 to switch from a screen-on state to a screen-off state; when the proximity sensor 714 detects that the distance between the user and the front of the terminal 700 is gradually increasing, the processor 701 controls the display screen 705 to switch from a screen-off state to a screen-on state.

[0144] Those skilled in the art will understand that Figure 7 The structure shown does not constitute a limitation on terminal 700, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0145] Figure 8This is a schematic diagram of a server structure according to an embodiment of this application. The server 800 can vary considerably due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 801 and one or more memories 802. The memory 802 stores at least one computer program, which is loaded and executed by the processor 801 to implement the image-based structural segmentation method provided in the above-described method embodiments. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated here.

[0146] This application also provides a computer-readable storage medium storing at least one computer program. This computer program is loaded and executed by a processor of a computer device to implement the operations performed by the computer device in the image-based structural segmentation method of the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0147] This application also provides a computer program product or computer program, which includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform the image-based structural segmentation method provided in the various optional implementations described above.

[0148] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0149] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An image-based structural segmentation method, characterized in that, The method includes: Acquire structural images and physiological data of a target object, wherein the structural images contain multiple feature regions and the structural images belong to the category of medical images; The physiological data is subjected to feature recognition to obtain abnormal information of the target object. The abnormal information is used to indicate the abnormal feature parts in the target object. Based on the abnormal information, the segmentation requirements of the target object are determined. The segmentation requirements are used to indicate the multiple target feature parts that need to be segmented. The image segmentation model uses a gated network to process the structured image and the segmentation requirements, determining the confidence scores of multiple sub-networks. These confidence scores represent the recognition performance of the corresponding sub-networks for each target feature region. Based on the confidence scores of the multiple sub-networks, a gate vector is determined in the image segmentation model. The image segmentation model includes multiple sub-networks, each used to recognize its corresponding feature region, and the gate vector indicates which sub-network to use. Based on the gating vector, multiple target sub-networks among the multiple sub-networks are invoked, and the multiple target sub-networks are used to identify the corresponding target feature parts; The structured image is processed through the multiple target sub-networks of the image segmentation model to obtain the outputs of the multiple target sub-networks; Based on the structured image and the segmentation requirements, multiple prediction kernels are generated. These multiple prediction kernels are used to identify their respective target feature regions, and the element values ​​in the multiple prediction kernels are different. Based on the multiple prediction kernels and the outputs of the multiple target sub-networks, the positions of the multiple target feature regions in the structured image are determined, and the positions of the multiple target feature regions are used to represent the segmentation results.

2. The method according to claim 1, characterized in that, Determining the gate vector in the image segmentation model based on the confidence scores of the multiple sub-networks includes: From the confidence scores of the multiple sub-networks sorted from high to low, select a predetermined number of target scores that rank highly. The weights of the sub-networks corresponding to the preset number of target scores are set to a first value, and the weights of the remaining sub-networks in the image segmentation model are set to a second value. The first value is used to indicate that the corresponding sub-network contributes to the recognition of the multiple target feature parts, and the second value is used to indicate that the corresponding sub-network does not contribute to the recognition of the multiple target feature parts. The gate vector is determined based on the weights corresponding to the multiple sub-networks in the image segmentation model.

3. The method according to claim 1, characterized in that, Based on the structural image and the segmentation requirements, multiple prediction kernels are generated, including: Based on the structural image and the segmentation requirements, the multiple target feature regions to be segmented are determined; For any target feature region, determine the associated feature region of the target feature region, and there is an attribute association between the associated feature region and the target feature region; Based on the target feature region and the associated feature region, determine the element value in the prediction kernel corresponding to the target feature region.

4. The method according to claim 1, characterized in that, The step of determining the positions of the multiple target feature regions in the structured image based on the outputs of the multiple prediction kernels and the multiple target sub-networks includes: By fusing the features output by the multiple target sub-networks, the target features are obtained; The multiple prediction kernels are convolved with the target features respectively to obtain the positions of the multiple target feature parts in the structure image.

5. The method according to claim 1, characterized in that, The step of determining the positions of the multiple target feature regions in the structured image based on the outputs of the multiple prediction kernels and the multiple target sub-networks includes: For the identification of any target feature region, the prediction kernel corresponding to the target feature region is convolved with the feature output by the target sub-network corresponding to the target feature region to obtain the position of the target feature region in the structure image.

6. The method according to claim 1, characterized in that, The step of determining the positions of the multiple target feature regions in the structured image based on the outputs of the multiple prediction kernels and the multiple target sub-networks includes: The overall features are obtained by processing the structure image using a convolutional network; By fusing the features output by the multiple target sub-networks in the image segmentation model, the target features are obtained. The overall features and the target features are fused to obtain the fused features; Based on the multiple prediction kernels and the fusion features, the positions of the multiple target feature regions in the structural image are determined.

7. An image-based structure segmentation device, characterized in that, The device includes: An acquisition module is used to acquire structural images and physiological data of a target object, wherein the structural image contains multiple feature regions and the structural image is a medical image; to perform feature recognition on the physiological data to obtain abnormal information of the target object, wherein the abnormal information is used to indicate the presence of abnormal feature regions in the target object; and to determine the segmentation requirements of the target object based on the abnormal information, wherein the segmentation requirements are used to indicate the multiple target feature regions that need to be segmented. The first processing module is used to process the structured image and the segmentation requirement through a gating network of an image segmentation model, determine the confidence scores of multiple sub-networks, the confidence scores representing the recognition performance of the corresponding sub-networks for each target feature region; and determine the gating vector in the image segmentation model based on the confidence scores of the multiple sub-networks, the image segmentation model including multiple sub-networks, each of which is used to recognize its corresponding feature region, and the gating vector indicating the sub-network to be invoked. The calling module is used to call multiple target sub-networks among the multiple sub-networks based on the gate vector, wherein the multiple target sub-networks are used to identify the corresponding target feature parts; The first processing module is used to process the structured image through the multiple target sub-networks of the image segmentation model to obtain the output of the multiple target sub-networks; The prediction module is used to generate multiple prediction kernels based on the structure image and the segmentation requirements. The multiple prediction kernels are used to identify their respective target feature regions, and the element values ​​in the multiple prediction kernels are different. Based on the multiple prediction kernels and the output of the multiple target sub-networks, the positions of the multiple target feature regions in the structure image are determined, and the positions of the multiple target feature regions are used to represent the segmentation result.

8. A computer device, characterized in that, The computer device includes a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded by the processor and executed as the image-based structural segmentation method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one computer program for performing the image-based structural segmentation method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the image-based structural segmentation method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Three-dimensional medical image segmentation model and training method and application thereof

    CN114820636A

  • Image segmentation paraphrasing method based on multi-expert mixing, electronic equipment and storage medium

    CN115690129A