Image segmentation processing method and device, storage medium and electronic equipment

By combining two-dimensional images and temporal features to segment ultrasound images, the problem of low segmentation accuracy in existing technologies is solved, achieving higher segmentation accuracy and nodule recognition effect.

CN114972368BActive Publication Date: 2025-12-19ALIBABA GROUP HOLDING LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110218572.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-26
Publication Date
2025-12-19
Estimated Expiration
2041-02-26

AI Technical Summary

Technical Problem

Existing ultrasound image segmentation methods mainly rely on labeled two-dimensional image data for training, resulting in low segmentation accuracy. This is especially true in breast ultrasound examinations, where image quality is greatly affected by external interference, nodule features are highly uncertain, and the segmentation task is complex.

Method used

This method combines two-dimensional images and temporal features. By acquiring target regions and sub-video data from the video to be processed, temporal features are used for segmentation, avoiding over-reliance on two-dimensional image data and providing segmentation guidance by combining contextual video segments.

Benefits of technology

It improves the segmentation accuracy of ultrasound images, especially in breast ultrasound examinations, enhancing the segmentation of nodules and reducing the need for annotation data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972368B_ABST
    Figure CN114972368B_ABST
Patent Text Reader

Abstract

The application discloses an image segmentation processing method and device, a storage medium and an electronic device. The method comprises the following steps: acquiring a video to be processed; determining a target region in the video to be processed; acquiring sub-video data displayed in the target region during playing of the video to be processed; and segmenting the video to be processed based on a time sequence feature of the sub-video data to obtain an image segmentation result. The application solves the technical problem that the segmentation accuracy of an ultrasonic image is low, because most of the current ultrasonic image segmentation methods are performed only by training labeled two-dimensional image data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image segmentation processing, in particular to an image segmentation processing method and device, a storage medium and an electronic device. BACKGROUND

[0002] In the prior art, when doctors use breast ultrasound examination, the ultrasound examination results are often displayed in the form of video images (continuous image frames), and the segmentation of nodules in breast ultrasound images is greatly affected by the image quality itself. Since ultrasound images are easily disturbed by external interference, the generated images often have noise and other factors that complicate the segmentation task; and the nodules themselves have great uncertainty caused by volume, shape, and morphology, so that doctors often need to watch repeatedly to check out nodules through video.

[0003] Since the current ultrasound image segmentation method is mostly performed by training labeled two-dimensional image data, but in the actual obtained data, the annotation data does not annotate every frame of image of the video, and the obtained annotation often has the following characteristics: 1. annotating one frame of image every ten frames; 2. doctors often select relatively large and obvious nodule shadows when selecting video images for annotation. Most methods at the present stage only use the annotated frame images for training on the two-dimensional segmentation network, and the segmentation accuracy of ultrasound images is low in actual use.

[0004] In view of the above problems, no effective solution has been proposed so far. SUMMARY

[0005] The embodiments of the present application provide an image segmentation processing method and device, a storage medium and an electronic device, to at least solve the technical problem that the current ultrasound image segmentation method is mostly performed by training labeled two-dimensional image data, and the segmentation accuracy of ultrasound images is low.

[0006] According to an aspect of an embodiment of the present application, an image segmentation processing method is provided, comprising: acquiring a to-be-processed video; determining a target region in the to-be-processed video; acquiring sub-video data displayed in the target region in the to-be-processed video playing process; and segmenting the to-be-processed video based on the time sequence characteristics of the sub-video data to obtain an image segmentation result.

[0007] According to another aspect of the embodiments of the present application, another image segmentation processing method is provided, which includes: a cloud server receiving a request message from a client, wherein the request message carries identification information representing a to-be-processed video; the cloud server obtaining the to-be-processed video based on the identification information and determining a target region in the to-be-processed video; the cloud server obtaining sub-video data displayed in the target region during playing of the to-be-processed video; the cloud server segmenting the to-be-processed video based on a time sequence feature of the sub-video data to obtain an image segmentation result; and the cloud server returning the image segmentation result to the client.

[0008] According to another aspect of the embodiments of the present application, another image segmentation processing method is provided, which includes: obtaining a to-be-processed video and displaying the to-be-processed video on an operation interface of a device; determining a target region in the to-be-processed video in response to a request instruction sensed on the operation interface; displaying the to-be-processed video in playing on the operation interface and capturing sub-video data displayed in the target region; and displaying an image segmentation result of the to-be-processed video on the operation interface, wherein the image segmentation result is obtained by segmenting the to-be-processed video based on a time sequence feature of the sub-video data.

[0009] According to another aspect of the embodiments of the present application, another image segmentation processing method is provided, which includes: obtaining a to-be-processed video of a lesion by a medical device and displaying the to-be-processed video on a case operation interface of the medical device; determining a target region in the to-be-processed video; obtaining sub-video data displayed in the target region during playing of the to-be-processed video; segmenting the to-be-processed video based on a time sequence feature of the sub-video data to obtain an image segmentation result; filling the image segmentation result into a text recording case information to obtain a structured filled case text; and displaying the structured filled case text on the case operation interface.

[0010] According to another aspect of the embodiments of the present application, an image segmentation processing apparatus is provided, which includes: a first obtaining module configured to obtain a to-be-processed video; a determining module configured to determine a target region in the to-be-processed video; a second obtaining module configured to obtain sub-video data displayed in the target region during playing of the to-be-processed video; and a segmentation processing module configured to segment the to-be-processed video based on a time sequence feature of the sub-video data to obtain an image segmentation result.

[0011] According to another aspect of the embodiments of the present application, a non-volatile storage medium is provided, which includes a stored program, wherein the program controls a device where the non-volatile storage medium is located to perform any one of the image segmentation processing methods.

[0012] According to another aspect of the embodiments of the present application, an electronic device is also provided, comprising: a processor; and a memory connected with the processor, configured to provide the processor with instructions to process the following processing steps: obtaining a video to be processed; determining a target region in the video to be processed; obtaining sub-video data displayed in the target region during playing of the video to be processed; and segmenting the video to be processed based on the time sequence characteristics of the sub-video data to obtain an image segmentation result.

[0013] In the embodiments of the present application, the video to be processed is obtained, the target region in the video to be processed is determined, the sub-video data displayed in the target region during playing of the video to be processed is obtained, and the video to be processed is segmented based on the time sequence characteristics of the sub-video data to obtain an image segmentation result.

[0014] It is easy to note that, in the embodiments of the present application, when the video to be processed is segmented, the time sequence characteristics of the three-dimensional image in the video to be processed are utilized, and the method combining two-dimensional images and time sequence characteristics is used to avoid excessive dependence on training of labeled two-dimensional image data for ultrasound image segmentation, and the attention of the current image resolution can be provided according to the context video segment, so that the segmentation result of the nodule in the chest ultrasound video image is obtained.

[0015] Therefore, the embodiments of the present application achieve the purpose of segmenting the ultrasound image based on the combination of two-dimensional images and time sequence characteristics, thereby realizing the technical effect of improving the segmentation accuracy of the ultrasound image, and further solving the technical problem that the current segmentation method of the ultrasound image is mostly performed by training of labeled two-dimensional image data, and the segmentation accuracy of the ultrasound image is low. BRIEF DESCRIPTION OF DRAWINGS

[0016] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application and illustrate the illustrative embodiments of the present application and their description serve to explain the present application, and do not constitute improper limitations on the present application. In the drawings:

[0017] Figure 1 Fig. 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing an image segmentation processing method;

[0018] Figure 2 Fig. 2 is a flowchart of an image segmentation processing method according to an embodiment of the present application;

[0019] Figure 3 Fig. 3 is a flowchart of an optional image segmentation processing method according to an embodiment of the present application;

[0020] Figure 4is a flow chart of an image segmentation processing method according to an embodiment of the present application;

[0021] Figure 5 is a flow chart of an image segmentation processing method according to an embodiment of the present application;

[0022] Figure 6 is a flow chart of an image segmentation processing method according to an embodiment of the present application;

[0023] Figure 7 is a structural schematic diagram of an image segmentation processing device according to an embodiment of the present application;

[0024] Figure 8 is a structural block diagram of another computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to make the personnel in the technical field better understand the present application scheme, the technical scheme in the present application embodiment will be clearly and completely described below in combination with the drawings in the present application embodiment. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the scope of protection of the present application.

[0026] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] First, some of the nouns or terms appearing in the description of the embodiments of the present application are applicable to the following explanations:

[0028] Encoder refers to a neural network component that encodes an image into multi-scale high-dimensional features by combining multiple layers of normalization layers, convolution layers, activation layers and pooling layers.

[0029] Decoder refers to a neural network component that restores the multi-scale features encoded by the encoder into a target segmentation prediction image through normalization layers, convolution layers, activation layers and up-sampling layers.

[0030] Feature fusion refers to a process of merging multi-scale features coded by two encoders.

[0031] Registration refers to a process of matching and superimposing two or more images acquired at different times, different sensors (imaging devices), or different conditions (weather, illumination, camera position and angle, etc.) into the same coordinate system, wherein rigid registration includes translation, scaling and rotation transformation; and deformable registration (non-rigid registration) refers to a process of deformation in addition to rigid registration.

[0032] Embodiment 1

[0033] According to the embodiments of the present application, an embodiment of an image segmentation processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0034] The method embodiment provided by Embodiment 1 of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the image segmentation processing method is shown. As shown in Figure 1 The computer terminal 10 (or mobile device 10) can include one or more processors 102 (the processor 102 can include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports in the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can also include more or fewer components than those shown in Figure 1 or have a different configuration than that shown in Figure 1 .

[0035] It should be noted that the one or more processors 102 and / or other data processing circuitry described above can be generally referred to herein as "data processing circuitry". The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuitry can be a single standalone processing module, or incorporated in whole or in part into any of the other elements of the computer terminal 10 (or mobile device). As referred to in embodiments of the present application, the data processing circuitry functions as a processor to control, for example, selection of a variable resistance terminal path connected to an interface.

[0036] The memory 104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the image segmentation processing method in embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, i.e. implements the image segmentation processing method described above. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0037] The transmission device 106 is used to receive or send data via a network. Specific examples of the network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet in a wireless manner.

[0038] The display can be, for example, a touch screen type liquid crystal display (LCD), which can enable a user to interact with the user interface of the computer terminal 10 (or mobile device).

[0039] In the above operating environment, the present application provides an image segmentation processing method as shown in Figure 2 Figure 2 is a flowchart of an image segmentation processing method according to embodiments of the present application, as shown in Figure 2

[0040] In step S202, a video to be processed is acquired.​​

[0041] Step S204, determining a target region in the to-be-processed video;

[0042] Step S206, acquiring sub-video data displayed in the target region in the to-be-processed video playing process;

[0043] It should be noted that the sub-video data can include multiple frames of ultrasound image data, and the multiple frames of ultrasound image data can constitute an ultrasound video sequence. In order to obtain more accurate results, part of the frames of the ultrasound video sequence can be labeled, that is, part of the frames of the ultrasound video sequence have lesion segmentation labels.

[0044] Step S208, segmenting the to-be-processed video based on the time sequence features of the sub-video data to obtain an image segmentation result.

[0045] In the embodiments provided in the above steps of the present application, the embodiments can be applied in the fields of audio and video, image, live broadcast, etc. First, a target region can be determined from a to-be-processed video currently played or stored. Sub-video data displayed in the target region in a predetermined time sequence at a playing time or in a playing process of the to-be-processed video can be captured. Then, the to-be-processed video can be segmented based on time sequence features of the sub-video data to obtain an image segmentation result.

[0046] It can be easily noted that the embodiments of the present application utilize time sequence features of sub-video data (which can be ultrasound image) in a video, and combine images (which can be two-dimensional images) and time sequence features to avoid over-reliance on training of labeled image data for ultrasound image segmentation. The attention of the current image resolution can be provided according to the context video segment, so as to obtain a nodule segmentation result in a chest ultrasound video image.

[0047] Therefore, the embodiments of the present application achieve the purpose of segmenting ultrasound images based on the combination of ultrasound images (which can be two-dimensional images) and time sequence features, thereby achieving the technical effect of improving the segmentation accuracy of ultrasound images. Thus, the technical problem of low segmentation accuracy of ultrasound images is solved.

[0048] Optionally, the to-be-processed video includes an ultrasound video image. As an optional embodiment, a doctor can display the ultrasound examination result in the form of an ultrasound video image (continuous image frames) when performing examination on a patient by using breast ultrasound.

[0049] As an optional embodiment, since there are other interference factors such as program boxes and text annotations in addition to the ultrasound images in the to-be-processed video, in the embodiment of the application, the target region in the to-be-processed video is determined, that is, the interference factors in the to-be-processed video are removed, and only the main image region, that is, the target region, in the to-be-processed video is extracted.

[0050] Since the target region includes video image groups and two-dimensional images, and the video image groups include three-dimensional images, when the to-be-processed video is segmented, the to-be-processed video can be segmented based on the time sequence characteristics of the sub-video data to obtain an image segmentation result.

[0051] Optionally, the embodiment of the application obtains the image segmentation result, that is, the segmentation result of the nodules in the chest ultrasound video image, by providing guidance attention when resolving the to-be-processed video through the context video segment based on the combination of two-dimensional images and time sequence characteristics.

[0052] Therefore, the embodiment of the application achieves the purpose of segmenting the ultrasound image based on the combination of two-dimensional images and time sequence characteristics, thereby achieving the technical effect of improving the segmentation accuracy of the ultrasound image, and further solving the technical problem that the current segmentation method of the ultrasound image is mostly performed by training labeled two-dimensional image data, and the segmentation accuracy of the ultrasound image is low.

[0053] In addition, the image segmentation processing method provided by the embodiment of the application belongs to a semi-supervised method, and can be trained by using video segment data without frame labeling, so that the amount of data required for labeling when segmenting the to-be-processed video can be significantly reduced.

[0054] In the embodiment of the application, since currently only the video segment of "labeled image frame a-unlabeled image frame group-labeled image frame b" is used in three-dimensional data, and then the two-dimensional image is trained using the labeled image frame a, only the context features in the time sequence can be combined in the training. In order to simultaneously collect the context features, "labeled image frame a-unlabeled image frame group-labeled image frame c-unlabeled image frame group-labeled image frame b" can be used, and then the two-dimensional image is trained using the labeled image frame c, so that the context features can be simultaneously collected.

[0055] As an optional embodiment, the embodiment of the application can utilize the time sequence characteristics of each two-dimensional image to perform registration and deformation, so that the neural network learns the change characteristics of the original nodule in each frame near a frame image on the video, and finally fuses the features in a feature fusion manner to the segmentation attention of the frame image, so that the neural network pays more attention to the nearby area of the nodule, thereby enhancing the segmentation effect. Compared with using video segments or single-frame images for segmentation, the guidance during neural network training can be increased.

[0056] In the embodiment of the application, the interval frame label and the characteristics of the video itself can be used to obtain pseudo labels by inserting frames for video labeling, and the learned features are fused with the features learned through single-frame images to enhance the attention to the position of the nodule near a frame image. It is equivalent to that the neural network has obtained corresponding guidance from the time sequence of the image, and finally enhances the segmentation effect of the ultrasound nodule.

[0057] In an optional embodiment, the step of determining the target region in the to-be-processed video can be implemented by the following implementation steps:

[0058] Step S302, determining the interference factors in the to-be-processed video, wherein the interference factors can include program boxes and / or text labels;

[0059] Step S304, determining the target region in the to-be-processed video by eliminating the interference factors.

[0060] Optionally, in the embodiment of the application, for the interference factors in the original image, the target region in the to-be-processed video can be determined by eliminating the interference factors.

[0061] In an optional embodiment, the step of determining the target region in the to-be-processed video can be implemented by the following implementation steps:

[0062] Step S402, performing first image segmentation processing on the to-be-processed video to obtain a first image segmentation processing result, wherein the first image segmentation processing includes grayscale processing and binary processing, and the first image segmentation processing is used to distinguish the gray scale of pixels in the to-be-processed video;

[0063] Step S404, performing connected domain processing on the first image segmentation processing result to obtain a second processing result, wherein the connected domain processing is used to eliminate connected domains that do not meet predetermined requirements in the to-be-processed video;

[0064] Step S406, calculating all connected domains in the second processing result and the constraint boxes of each connected domain in the all connected domains by using an open source computer vision library;

[0065] Step S408, taking the content in the constraint frame of the target connected domain in all the connected domains as the target region, wherein the target connected domain is a connected domain containing a nodule lesion.

[0066] Optionally, in the embodiment of the present application, the first image segmentation processing result is obtained by performing grayscale processing and binary processing on the to-be-processed video, wherein the pixels greater than a predetermined grayscale value (for example, the grayscale value is 5) in the first image segmentation processing result are set to 1, and the other pixels less than or equal to the predetermined grayscale value (for example, the grayscale value is 5) are set to 0.

[0067] As an optional embodiment, the first image segmentation processing result is processed by using image morphological opening operation to remove the thin part of the connected domain and retain the remaining large connected domain, and a second processing result is obtained. Then, all connected domains in the second processing result and the constraint frame of each connected domain in all connected domains are calculated by using an open source computer vision library. For example, the connected domain and the constraint frame of each connected domain are calculated by using an open source library (opencv) which integrates many general algorithms of image segmentation processing and computer vision.

[0068] As an optional embodiment, the content in the constraint frame of the target connected domain in all the connected domains is taken as the target region (i.e., the target segmentation region), wherein the target connected domain is a connected domain containing a nodule lesion. That is, the target connected domain containing the nodule lesion is retained in all the connected domains, and the content in the constraint frame of the target connected domain is extracted as the target region.

[0069] In an optional embodiment, the step of obtaining the sub-video data displayed in the target region in the playing process of the to-be-processed video can include the following implementation steps:

[0070] Step S502, obtaining a plurality of image data displayed in the target region in the playing process of the to-be-processed video;

[0071] Step S504, constructing a video sequence based on the plurality of image data to obtain the sub-video data;

[0072] Optionally, part of the plurality of image data has a lesion segmentation label, and the sub-video data includes a video image group.

[0073] As an optional embodiment, in the process of playing the to-be-processed video, display content in the target region can be acquired, the display content being a plurality of frames of ultrasound image data, the plurality of frames of ultrasound image data forming an ultrasound video sequence, and part of the frames of video in the ultrasound video sequence having lesion segmentation labels. At this time, the original video can be segmented based on the time sequence features of the ultrasound image data to obtain an image segmentation result.

[0074] In an optional embodiment, segmenting the to-be-processed video based on the time sequence features of the sub-video data to obtain an image segmentation result can include the following implementation steps:

[0075] Step S602: Extracting time sequence features of each two-dimensional image in the video image group;

[0076] Step S604: Generating pseudo labels of each two-dimensional image based on the time sequence features, wherein the pseudo labels are used for approximately labeling each two-dimensional image;

[0077] Step S606: Training a three-dimensional segmentation network and a two-dimensional segmentation network using the pseudo labels to obtain a trained three-dimensional segmentation network and a trained two-dimensional segmentation network;

[0078] Step S608: Segmenting the to-be-processed video using the trained two-dimensional segmentation network and the trained three-dimensional segmentation network to obtain an image segmentation result.

[0079] Optionally, in the embodiments of the present application, as shown in Figure 3 the time sequence features of each two-dimensional image in the video image group are extracted, pseudo labels of each two-dimensional image are generated based on the time sequence features, i.e., an image segmentation label, each two-dimensional image is approximately labeled using the pseudo labels, a three-dimensional segmentation network and a two-dimensional segmentation network are trained using the pseudo labels to obtain a trained three-dimensional segmentation network and a trained two-dimensional segmentation network, and the to-be-processed video is segmented using the trained two-dimensional segmentation network and the trained three-dimensional segmentation network to obtain an image segmentation result, for example, a segmentation label of a first image of a segmentation image group of a target video.

[0080] In the embodiments of the present application, the segmentation features on the three-dimensional segmentation network can be used to guide the segmentation of the two-dimensional segmentation network, so that the two-dimensional segmentation network pays more attention to the time sequence features on the three-dimensional image.

[0081] As an optional embodiment, in addition to using the encoder and decoder in the simple UNet, the memory gate and the forgetting gate of the long short-term memory artificial neural network (LSTM) can also be used to extract the time sequence features.

[0082] In an optional embodiment, the aforementioned temporal features may include registration parameters and deformation parameters. It should be further noted that the example of extracting the temporal features of each frame of the aforementioned video image group can be implemented through the following steps:

[0083] Step S702: Obtain the video image group within the target area mentioned above;

[0084] Step S704: Based on the registered two-dimensional images in the above video image group, the registration parameters are calculated, wherein the registered two-dimensional images are obtained by rigidly registering two consecutive frames of the above two-dimensional images in the above video image group.

[0085] Step S706: Calculate deformation parameters based on the cropped two-dimensional image, wherein the cropped two-dimensional image is obtained by cropping the registered two-dimensional image.

[0086] Step S708: Based on the above registration parameters and the above deformation parameters, generate the temporal features of each frame of the above two-dimensional image in the above video image group.

[0087] It should be further explained here that the registration in the above embodiment is performed on two consecutive two-dimensional images in a video image group (which may be a video sequence). The 32 consecutive two-dimensional images after registration and deformation will form a three-dimensional image, which will serve as the input to the subsequent neural network.

[0088] As can be seen from the technical solutions provided in steps S602 to S608 above, in an optional embodiment of this application, video image groups within the acquired target area can be extracted, for example... Figure 3 The segmented image group of the target video shown is processed by performing the following operation on multiple consecutive images (e.g., 32 images) within two labeled frames to obtain approximate labels for these 32 frames: By simplifying the video to be processed, the amount of subsequent calculations and the influence of noise in the ultrasound itself can be reduced.

[0089] In a specific optional solution, after extracting the video image group, the embodiment of the application can perform rigid registration on two continuous images by using the open-source tool of elasitx registration to obtain a registered two-dimensional image, and based on the registered two-dimensional image in the video image group, a registration parameter can be calculated; then, a cropped two-dimensional image is obtained by performing cropping processing on the registered two-dimensional image, for example, black edge cropping processing is performed on the registered two-dimensional image, the cropped two-dimensional image is used to calculate a deformation parameter by using the Demons deformable registration in the opencv open-source library, and then the registration parameter and the Demons deformation parameter are projected onto the labels of each frame of the video image group to generate the time sequence features of each frame of the two-dimensional image in the video image group.

[0090] In the embodiment of the application, a processing method of directly performing registration and deformation change on the to-be-processed video can also be used, and the generated time sequence features are input into the encoder of the three-dimensional segmentation network, and the decoder directly uses a 2D decoder to output a predicted value and a real label of a single frame image as a loss function.

[0091] In addition, the embodiment of the application can not only use a traditional registration method for registration, but also use a more accurate registration method such as an optical flow network to obtain the optical flow field of the pixel change in the video, and then project it onto the lesion segmentation label to generate the corresponding pseudo label. In addition, in the embodiment of the application, a method combining rigid registration and displacement field can also be used to calculate the change and inter-frame label of the inter-frame image.

[0092] In an optional embodiment, part of the frame image data in the above-mentioned multi-frame image data has a lesion segmentation label, and the pseudo label of each frame of the two-dimensional image is generated based on the time sequence features, including:

[0093] Step S802, obtaining the lesion segmentation label of each frame of the three-dimensional image;

[0094] Step S804, fusing the time sequence features and the lesion segmentation label to obtain the pseudo label.

[0095] Optionally, since the pseudo label is used for approximate labeling of each frame of the two-dimensional image, in the embodiment of the application, the lesion segmentation label of each frame of the three-dimensional image is obtained; the time sequence features of each frame of the three-dimensional image in the video image group are fused with the lesion segmentation label to obtain the pseudo label.

[0096] In an optional embodiment, the three-dimensional segmentation network and the two-dimensional segmentation network are trained by using the pseudo label to obtain a trained three-dimensional segmentation network and a trained two-dimensional segmentation network, including:

[0097] Step S902, training the three-dimensional segmentation network by using the pseudo label to obtain a trained three-dimensional segmentation network;

[0098] Step S904, obtaining three-dimensional feature values of each layer in the trained three-dimensional segmentation network, wherein the three-dimensional feature values are used as segmentation guidance when training the two-dimensional segmentation network;

[0099] Step S906, training the two-dimensional segmentation network by using the three-dimensional feature values to obtain a trained two-dimensional segmentation network.

[0100] As an optional embodiment, the image segmentation model is based on a segmentation network UNet, a three-dimensional segmentation network 3D UNet is used for a video image group, a two-dimensional segmentation network 2D UNet is used for a two-dimensional image, when the video image group is input into an encoder of the 3D UNet, three-dimensional feature values extracted in each layer and two-dimensional feature values feat2d output by the two-dimensional segmentation network 2D UNet are fused to form fused feature values which are then input into a next layer of the two-dimensional segmentation network 2D UNet.

[0101] In the embodiment of the application, the target region of the video stream is output with predicted three-dimensional feature values by the three-dimensional segmentation network 3D UNet, the loss coefficient Dice loss of medical image segmentation is calculated by using the pseudo label to train the three-dimensional segmentation network to obtain a trained three-dimensional segmentation network, and the two-dimensional segmentation network is trained by using the three-dimensional feature values to obtain a trained two-dimensional segmentation network.

[0102] In an optional embodiment, the two-dimensional segmentation network is trained by using the three-dimensional feature values to obtain a trained two-dimensional segmentation network, including:

[0103] Step S1002, obtaining two-dimensional feature values output by a first layer network of the two-dimensional segmentation network;

[0104] Step S1004, fusing the three-dimensional feature values and the two-dimensional feature values to obtain fused feature values;

[0105] Step S1006, inputting the fused feature values into a second layer network of the two-dimensional segmentation network, wherein the second layer network is a next layer network of the first layer network.

[0106] Optionally, in the above embodiment, the target region of the video to be processed is input into the two-dimensional segmentation network 2D UNet to obtain two-dimensional feature values output by a first layer network of the two-dimensional segmentation network, the pseudo label is used to approximately label three-dimensional feature values output by the 3D UNet for each frame of the two-dimensional image, the three-dimensional feature values and the two-dimensional feature values are fused to obtain fused feature values, and the fused feature values are input into a second layer network of the two-dimensional segmentation network.

[0107] As shown in Figure 3 Each two-dimensional feature value in the two-dimensional segmentation network needs to be fused with the three-dimensional feature value, and the feature value fusion processing is performed in a loop until the feature value fusion processing of all layers is completed.

[0108] In an optional embodiment, the three-dimensional feature value and the two-dimensional feature value are fused to obtain a fused feature value, including:

[0109] Step S1102, compressing the three-dimensional feature value to obtain a compressed feature value;

[0110] Step S1104, calculating the compressed feature value and the two-dimensional feature value to obtain an attention weight matrix;

[0111] Step S1106, fusing the three-dimensional feature value and the two-dimensional feature value by using the attention weight matrix to obtain the fused feature value.

[0112] As an optional embodiment, the feature fusion method in the embodiment mainly uses a feature fusion mechanism Attention, and the main calculation method of the feature fusion mechanism is to compress the time sequence feature dimension of the three-dimensional feature value feat3d to the channel dimension, then calculate the attention weight matrix, for example, the affinity matrix, of the compressed feature value and the two-dimensional feature value feat2d, and fuse the three-dimensional feature value and the two-dimensional feature value by using the attention weight matrix to obtain the fused feature value, for example, the attention weight matrix after softmax is weighted to the feat3d feature and finally added to the feat2d.

[0113] In the embodiment, the three-dimensional feature value and the image segmentation label are calculated to obtain a loss coefficient Dice loss and a cross entropy loss coefficient Cross Entropy loss.

[0114] As an optional embodiment, the nodules are segmented by combining the context of the image and the time sequence under the condition that only the interval frame image is labeled, mainly using the registration and deformation method to extract the sequential change features between the image frames, fusing these features into the lesion segmentation label to obtain pseudo labels, and then using the pseudo labels to train the three-dimensional segmentation network, taking the feature activation value of each layer after training as the guiding attention when training the two-dimensional network, so that the two-dimensional segmentation network pays more attention to whether there is a nodule or a possible nodule at the corresponding position.

[0115] According to the embodiment, an image segmentation processing method is also provided, as shown in Figure 4 Figure 4 ​is a flowchart of another image segmentation processing method according to an embodiment of the present application, as shown in the figure, the embodiment is for the application scenario deployed in the cloud server, the image segmentation processing method involves an embodiment that can be implemented by the following execution steps: Figure 4

[0116] In step S1202, the cloud server receives a request message from the client, wherein the request message carries identification information representing the to-be-processed video.

[0117] In step S1204, the cloud server obtains the to-be-processed video based on the identification information and determines a target region in the to-be-processed video.

[0118] In step S1206, the cloud server obtains sub-video data displayed in the target region during playback of the to-be-processed video.

[0119] In step S1208, the cloud server segments the to-be-processed video based on the time sequence features of the sub-video data to obtain an image segmentation result.

[0120] In step S1210, the cloud server returns the image segmentation result to the client.

[0121] In the embodiment of the present application, the client initiates a request message for requesting to obtain a video, and after the cloud server receives the request message from the client, the to-be-processed video that needs to be played or stored is obtained based on the identification information. Here, the request message carries identification information representing the to-be-processed video. After obtaining the to-be-processed video, the cloud server can determine a target region from the to-be-processed video and capture sub-video data displayed in the target region on each frame of image in the to-be-processed video. At this time, the cloud server segments the to-be-processed video based on the time sequence features of the sub-video data to obtain an image segmentation result, and finally returns the image segmentation result to the client that initiates the request message.

[0122] It is easy to note that in the embodiment of the present application, when the cloud server segments the to-be-processed video, the time sequence features of the three-dimensional image in the to-be-processed video are used, and the method combining two-dimensional images and time sequence features is used to avoid excessive dependence on training labeled two-dimensional image data for ultrasound image segmentation. After the image segmentation result is returned to the client, attention can be provided to the current image resolution according to the context video segment, so that the nodule segmentation result in the chest ultrasound video image is obtained.

[0123] ​Therefore, the embodiments of this application achieve the goal of segmenting ultrasound images based on a combination of two-dimensional images and temporal features, thereby improving the technical effect of segmenting ultrasound images and solving the technical problem that most current ultrasound image segmentation methods only rely on trained labeled two-dimensional image data, resulting in low segmentation accuracy of ultrasound images.

[0124] It should be noted that the execution subject of the image segmentation processing method provided in steps S1202 to S1210 above is the server, such as a SaaS cloud server. The video to be processed includes: ultrasound video images. As an optional embodiment, when doctors use breast ultrasound to examine patients, they can display the ultrasound examination results in the form of ultrasound video images (continuous image frames).

[0125] As an optional embodiment, since there are other interfering factors such as program frames and text annotations in addition to ultrasound images in the video to be processed, in this embodiment of the application, the target area in the video to be processed is determined, that is, the interfering factors in the video to be processed are removed, and only the main image area in the video to be processed other than the interfering factors is extracted, that is, the target area.

[0126] Since the target region includes video image groups and two-dimensional images, and the video image groups include three-dimensional images, when segmenting the video to be processed, the video to be processed can be segmented based on the temporal characteristics of the sub-video data to obtain image segmentation results.

[0127] Optionally, embodiments of this application use a method combining two-dimensional images and temporal features to provide guidance on distinguishing the video to be processed through contextual video segments, thereby obtaining the above-mentioned image segmentation result, namely the segmentation result of nodules in chest ultrasound video images.

[0128] Therefore, the embodiments of this application achieve the goal of segmenting ultrasound images based on a combination of two-dimensional images and temporal features, thereby improving the technical effect of segmenting ultrasound images and solving the technical problem that most current ultrasound image segmentation methods only rely on trained labeled two-dimensional image data, resulting in low segmentation accuracy of ultrasound images.

[0129] According to embodiments of this application, the following are also provided: Figure 5 Another image segmentation processing method shown is... Figure 5 This is a flowchart of another image segmentation processing method according to an embodiment of this application, such as... Figure 5 As shown, this embodiment is for an application scenario deployed on a client with interactive functionality. The embodiment involving this image segmentation processing method can be implemented through the following steps:

[0130] Step S1302, obtaining a to-be-processed video, and displaying the to-be-processed video on an operation interface of the device;

[0131] Step S1304, determining a target region in the to-be-processed video in response to a request instruction sensed on the operation interface;

[0132] Step S1306, displaying the to-be-processed video in play on the operation interface, and capturing sub-video data displayed in the target region;

[0133] Step S1308, displaying an image segmentation result of the to-be-processed video on the operation interface, wherein the image segmentation result is obtained by segmenting the to-be-processed video based on the time sequence features of the sub-video data.

[0134] In the embodiments of the present application, the display of the client can display the operation interface, which is a visual interactive interface, that is, the operation interface can receive external output operation instructions (such as touch operation, gesture operation, voice operation, etc.). The to-be-processed video obtained by the client can be displayed on the operation interface of the device. The to-be-processed video can be displayed and played on the operation interface. If the operation interface receives a request instruction, the request instruction can be responded to the sensed request instruction, the request instruction points to determining a target region in the to-be-processed video, and capturing sub-video data displayed in the target region. Since the image segmentation result can be obtained by segmenting the to-be-processed video based on the time sequence features of the sub-video data, the image segmentation result of the to-be-processed video can be finally displayed on the operation interface.

[0135] It is easy to note that in the embodiments of the present application, when segmenting the to-be-processed video, the time sequence features of the three-dimensional image in the to-be-processed video are used, and the method combining two-dimensional images and time sequence features is used to avoid excessive dependence on training labeled two-dimensional image data for ultrasound image segmentation. The attention of the current image resolution can be provided according to the context video segment, so as to obtain the nodule segmentation result in the chest ultrasound video image.

[0136] Therefore, the embodiments of the present application achieve the purpose of segmenting the ultrasound image based on the combination of two-dimensional images and time sequence features, thereby realizing the technical effect of improving the segmentation accuracy of the ultrasound image, and further solving the technical problem that the current ultrasound image segmentation method is mostly performed only by training labeled two-dimensional image data, and the segmentation accuracy of the ultrasound image is low.

[0137] It should be noted that the execution subject of the image segmentation processing method provided by the above steps S1302 to S1308 is a client, for example, a medical device client, which is applied to a scenario of implementing image segmentation based on a user interaction process, by acquiring a to-be-processed video, and displaying the to-be-processed video on an operation interface of the device; the user inputs a request instruction corresponding to the to-be-processed video on the operation interface, determines a target region in the to-be-processed video by responding to the request instruction sensed on the operation interface; the to-be-processed video in the playing is displayed on the operation interface, and sub-video data displayed in the target region is captured; and the image segmentation result of the to-be-processed video is displayed on the operation interface, wherein the image segmentation result is obtained by segmenting the to-be-processed video based on the time sequence characteristics of the sub-video data.

[0138] According to the embodiments of the present application, another image segmentation processing method is also provided, as shown in Figure 6 Figure 6 is a flowchart of another image segmentation processing method according to the embodiments of the present application, as shown in Figure 6 The embodiments of the present application are applied to the application scenario of processing the collected lesion video by a medical device (for example, a CT machine or other device that can acquire internal or external images of a human body or other object), and the embodiments involved in the image segmentation processing method can be implemented through the following execution steps:

[0139] Step S1402, acquiring a to-be-processed video of a lesion by a medical device, and displaying the to-be-processed video on a case operation interface of the medical device;

[0140] Step S1404, determining a target region in the to-be-processed video;

[0141] Step S1406, acquiring sub-video data displayed in the target region in the playing process of the to-be-processed video;

[0142] Step S1408, segmenting the to-be-processed video based on the time sequence characteristics of the sub-video data to obtain an image segmentation result;

[0143] Step S1410, filling the image segmentation result into a text recording case information to obtain a structured filled case text;

[0144] Step S1412, displaying the structured filled case text on the case operation interface.

[0145] ​In the embodiment of the present application, the display of the medical device can display the display content of the part being examined, and the display interface of the display is a visual interactive interface, that is, the display interface can receive external output operation instructions (such as touch operation, air gesture operation, voice operation, etc.). The medical device collects the to-be-processed video of the lesion, and displays the to-be-processed video on the case operation interface of the medical device. After determining the target region in the to-be-processed video, the sub-video data displayed in the target region during the playing of the to-be-processed video can be obtained, and the to-be-processed video is segmented based on the time sequence characteristics of the sub-video data to obtain an image segmentation result. At this time, the image segmentation result can be filled into the text recording the case information to obtain a structured filled case text, and the structured filled case text is displayed on the case operation interface. Such case text is intuitive to read and easy to consult.

[0146] It is easy to note that, in the embodiment of the present application, when the to-be-processed video is segmented, the time sequence characteristics of the three-dimensional image in the to-be-processed video are used, and a method combining two-dimensional images and time sequence characteristics is used to avoid excessive dependence on training of labeled two-dimensional image data for ultrasound image segmentation. The attention of the current image resolution can be provided according to the context video segment, so as to obtain the nodule segmentation result in the chest ultrasound video image.

[0147] Therefore, the embodiment of the present application achieves the purpose of segmenting the ultrasound image based on the combination of two-dimensional images and time sequence characteristics, thereby realizing the technical effect of improving the segmentation accuracy of the ultrasound image, and further solving the technical problem that the current ultrasound image segmentation method is mostly performed by training labeled two-dimensional image data, and the segmentation accuracy of the ultrasound image is low.

[0148] It should be noted that the image segmentation processing method provided by the steps S1402 to S1412 is applied to a multi-modal application scenario of text and vision. Optionally, the to-be-processed video includes an ultrasound video image. As an optional embodiment, when a doctor performs examination on a patient by using breast ultrasound, the doctor can display the ultrasound examination result in the form of an ultrasound video image (continuous image frames).

[0149] Due to the low efficiency, low completeness and low accuracy of manual case information collection and input in the prior art, in the embodiments of the present application, based on the image segmentation processing method provided in the embodiments of the present application, when a medical staff or a case text processing worker processes a patient's case text, a medical device collects a to-be-processed video of a lesion, and displays the to-be-processed video on a case operation interface of the medical device; then a target region in the to-be-processed video is determined; sub-video data displayed in the target region in the to-be-processed video playing process is obtained; the to-be-processed video is segmented based on the time sequence characteristics of the sub-video data to obtain an image segmentation result; the image segmentation result is filled into a text recording case information to obtain a structured filled case text; and the structured filled case text is displayed on the case operation interface.

[0150] Therefore, by the embodiments of the present application, automatic acquisition and structured filling of case texts can be realized, which not only improves the input efficiency of case information, but also guarantees the completeness and accuracy of case information input.

[0151] As an optional embodiment, since there are program boxes, text annotations and other interference factors in the to-be-processed video in addition to ultrasound images, in the embodiments of the present application, the target region in the to-be-processed video is determined, that is, the interference factors in the to-be-processed video are removed, and only the main image region, that is, the target region, in the to-be-processed video is extracted.

[0152] Since the target region includes video image groups and two-dimensional images, and the video image groups include three-dimensional images, when the to-be-processed video is segmented, the to-be-processed video can be segmented based on the time sequence characteristics of the sub-video data to obtain an image segmentation result.

[0153] Optionally, the embodiments of the present application provide guidance attention when distinguishing the to-be-processed video by combining two-dimensional images and time sequence characteristics through context video clips, so as to obtain the image segmentation result, that is, the segmentation result of the nodules in the chest ultrasound video image.

[0154] Therefore, the embodiments of the present application achieve the purpose of segmenting ultrasound images based on the combination of two-dimensional images and time sequence characteristics, thereby realizing the technical effect of improving the segmentation accuracy of ultrasound images, and further solving the technical problem that most of the current ultrasound image segmentation methods are only performed by training labeled two-dimensional image data, and the segmentation accuracy of ultrasound images is low.

[0155] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all described as a combination of a series of actions, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0156] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and a necessary general hardware platform, and of course it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product in essence or the part that contributes to the prior art, and the computer software product is stored in a non-volatile storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a plurality of instructions for causing an end device (which can be a mobile phone, computer, server, or network device, etc.) to execute the method of the above embodiments.

[0157] Embodiment 2

[0158] According to the embodiments of the present application, a device embodiment for implementing the above image segmentation processing method is also provided, Figure 7 is a structural schematic diagram of an image segmentation processing device according to an embodiment of the present application, as Figure 7 shown, the device includes a first acquisition module 400, a determination module 402, a second acquisition module 404, and a segmentation module 406, wherein:

[0159] The first acquisition module 400 is configured to acquire a to-be-processed video; the determination module 402 is configured to determine a target region in the to-be-processed video; the second acquisition module 404 is configured to acquire sub-video data displayed in the target region during playing of the to-be-processed video; and the segmentation module 406 is configured to segment the to-be-processed video based on a time sequence feature of the sub-video data to obtain an image segmentation result.

[0160] In the embodiments of the present application, by acquiring a to-be-processed video, determining a target region in the to-be-processed video, acquiring sub-video data displayed in the target region during playing of the to-be-processed video, and segmenting the to-be-processed video based on a time sequence feature of the sub-video data, an image segmentation result is obtained.

[0161] It is easy to note that in the embodiment of the present application, when the video to be processed is segmented, the time sequence characteristics of the three-dimensional image in the video to be processed are utilized, and the method combining two-dimensional images and time sequence characteristics is used to avoid excessive dependence on training of labeled two-dimensional image data for ultrasound image segmentation, and the attention of the current image resolution can be provided according to the context video segment, so as to obtain the nodule segmentation result in the chest ultrasound video image.

[0162] Therefore, the embodiment of the present application achieves the purpose of segmenting the ultrasound image based on the combination of two-dimensional images and time sequence characteristics, thereby realizing the technical effect of improving the segmentation accuracy of the ultrasound image, and further solving the technical problem that the current ultrasound image segmentation method is mostly performed only by training labeled two-dimensional image data, and the segmentation accuracy of the ultrasound image is low.

[0163] It should be noted that the first acquisition module 400, the determination module 402 and the segmentation module 406 correspond to steps S202 to S206 in Embodiment 1, and the three modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules can run in the computer terminal 10 provided in Embodiment 1 as part of the device.

[0164] It should be noted that the preferred embodiments of the present embodiment can refer to the related description in Method Embodiment 1, which will not be repeated here.

[0165] Embodiment 3

[0166] According to the embodiments of the present application, an embodiment of an electronic device is also provided, which can be any one of the computing devices in the computing device group. The electronic device comprises a processor and a memory, wherein:

[0167] The processor; and the memory, connected with the above-mentioned processor, for providing the above-mentioned processor with instructions for processing the following processing steps: acquiring a video to be processed; determining a target area in the video to be processed; acquiring sub-video data displayed in the target area during playing of the video to be processed; segmenting the video to be processed based on the time sequence characteristics of the sub-video data to obtain an image segmentation result.

[0168] In the embodiments of the present application, by acquiring a video to be processed; determining a target area in the video to be processed; acquiring sub-video data displayed in the target area during playing of the video to be processed; segmenting the video to be processed based on the time sequence characteristics of the sub-video data to obtain an image segmentation result.

[0169] It is easy to note that in the embodiment of the present application, when the video to be processed is segmented, the time sequence characteristics of the three-dimensional image in the video to be processed are utilized, and the method combining two-dimensional images and time sequence characteristics is used to avoid excessive dependence on training of labeled two-dimensional image data for ultrasound image segmentation. The attention of the current image resolution can be provided according to the context video segment, so as to obtain the nodule segmentation result in the chest ultrasound video image.

[0170] Therefore, the embodiment of the present application achieves the purpose of segmenting the ultrasound image based on the combination of two-dimensional images and time sequence characteristics, thereby realizing the technical effect of improving the segmentation accuracy of the ultrasound image, and further solving the technical problem that the current ultrasound image segmentation method is mostly performed only by training labeled two-dimensional image data, and the segmentation accuracy of the ultrasound image is low.

[0171] It should be noted that the preferred embodiments of the present embodiment can refer to the related description in Embodiment 1, which will not be repeated here.

[0172] Embodiment 4

[0173] According to the embodiments of the present application, an embodiment of a computer terminal is also provided, which can be any one of the computer terminal devices in the computer terminal group. Alternatively, in the present embodiment, the above-mentioned computer terminal can be replaced by a terminal device such as a mobile terminal.

[0174] Alternatively, in the present embodiment, the above-mentioned computer terminal can be located in at least one network device of the plurality of network devices of the computer network.

[0175] In the present embodiment, the above-mentioned computer terminal can execute program codes of the following steps in the image segmentation processing method: obtaining a video to be processed; determining a target region in the video to be processed; obtaining sub-video data displayed in the target region during playing of the video to be processed; and segmenting the video to be processed based on time sequence characteristics of the sub-video data to obtain an image segmentation result.

[0176] Alternatively, Figure 8 is another structural block diagram of a computer terminal according to the embodiments of the present application, as shown in Figure 8 The computer terminal can include one or more (only one is shown in the figure) processors 502, a memory 504, and a peripheral interface 506.

[0177] The memory can be configured to store software programs and modules, such as program instructions / modules corresponding to the image segmentation processing method and device in the embodiments of the present application. The processor executes various functions and data processing by running the software programs and modules stored in the memory, that is, implements the image segmentation processing method described above. The memory can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the computer terminal through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0178] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: obtaining a to-be-processed video; determining a target region in the to-be-processed video; obtaining sub-video data displayed in the target region during playing of the to-be-processed video; and segmenting the to-be-processed video based on a time sequence feature of the sub-video data to obtain an image segmentation result.

[0179] Optionally, the processor can further execute program codes of the following steps: determining an interference factor in the to-be-processed video, wherein the interference factor includes a program block and / or a text label; and determining the target region in the to-be-processed video by eliminating the interference factor.

[0180] Optionally, the processor can further execute program codes of the following steps: performing first image segmentation processing on the to-be-processed video to obtain a first image segmentation processing result, wherein the first image segmentation processing includes grayscale processing and binary processing, and the first image segmentation processing is used to distinguish the grayscale of pixels in the to-be-processed video; performing connected domain processing on the first image segmentation processing result to obtain a second processing result, wherein the connected domain processing is used to eliminate connected domains in the to-be-processed video that do not meet predetermined requirements; calculating all connected domains in the second processing result and constraint boxes of each connected domain in all connected domains by using an open source computer vision library; and taking content in the constraint box of a target connected domain in all connected domains as the target region, wherein the target connected domain is a connected domain containing a nodule lesion.

[0181] Optionally, the processor can further execute program codes of the following steps: extracting time sequence features of each two-dimensional image in the video image group; generating pseudo labels of each two-dimensional image based on the time sequence features, wherein the pseudo labels are used for approximately labeling each two-dimensional image; training the three-dimensional segmentation network and the two-dimensional segmentation network using the pseudo labels to obtain a trained three-dimensional segmentation network and a trained two-dimensional segmentation network; and segmenting the to-be-processed video using the trained two-dimensional segmentation network and the trained three-dimensional segmentation network to obtain an image segmentation result.

[0182] Optionally, the processor can further execute program codes of the following steps: obtaining a video image group in the target region; calculating registration parameters based on the registered two-dimensional images in the video image group, wherein the registered two-dimensional images are obtained by rigidly registering two consecutive two-dimensional images in the video image group; calculating deformation parameters based on the cropped two-dimensional images, wherein the cropped two-dimensional images are obtained by cropping the registered two-dimensional images; and generating time sequence features of each two-dimensional image in the video image group according to the registration parameters and the deformation parameters.

[0183] Optionally, the processor can further execute program codes of the following steps: obtaining two frames of data with labels in the three-dimensional image and lesion segmentation labels of the two frames of data; and fusing the time sequence features and the lesion segmentation labels to obtain pseudo labels of unlabelled data between the two frames of data.

[0184] Optionally, the processor can further execute program codes of the following steps: training the three-dimensional segmentation network using the pseudo labels to obtain a trained three-dimensional segmentation network; obtaining three-dimensional feature values of each layer in the trained three-dimensional segmentation network, wherein the three-dimensional feature values are used as segmentation guidance when training the two-dimensional segmentation network; and training the two-dimensional segmentation network using the three-dimensional feature values to obtain a trained two-dimensional segmentation network.

[0185] Optionally, the processor can further execute program codes of the following steps: obtaining two-dimensional feature values output by a first layer of the two-dimensional segmentation network; fusing the three-dimensional feature values and the two-dimensional feature values to obtain fused feature values; and inputting the fused feature values to a second layer of the two-dimensional segmentation network, wherein the second layer is a next layer of the first layer.

[0186] Optionally, the processor can further execute program codes of the following steps: compressing the three-dimensional feature values to obtain compressed feature values; calculating the compressed feature values and the two-dimensional feature values to obtain an attention weight matrix; and fusing the three-dimensional feature values and the two-dimensional feature values using the attention weight matrix to obtain the fused feature values.

[0187] By adopting the embodiment of the present application, a scheme of image segmentation processing is provided, which comprises: acquiring a to-be-processed video; determining a target region in the to-be-processed video; acquiring sub-video data displayed in the target region in a playing process of the to-be-processed video; and segmenting the to-be-processed video based on a time sequence feature of the sub-video data to obtain an image segmentation result.

[0188] It is easily noticed that, when segmenting the to-be-processed video, the embodiment of the present application utilizes the time sequence feature of the three-dimensional image in the to-be-processed video, and based on the method combining the two-dimensional image and the time sequence feature, the ultrasound image segmentation can be avoided from excessively relying on the training of the labeled two-dimensional image data, and the attention can be provided to the current image resolution according to the context video segment, so as to obtain the nodule segmentation result in the chest ultrasound video image.

[0189] Therefore, the embodiment of the present application achieves the purpose of segmenting the ultrasound image based on the combination of the two-dimensional image and the time sequence feature, so as to realize the technical effect of improving the segmentation accuracy of the ultrasound image, and further solves the technical problem that the current ultrasound image segmentation method is mostly performed by training the labeled two-dimensional image data, and the segmentation accuracy of the ultrasound image is low.

[0190] Those skilled in the art can understand that, Figure 8 The structure shown is only schematic, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, and the like. Figure 8 It does not limit the structure of the electronic device. For example, the computer terminal can further include more or less components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 8 It does not limit the structure of the electronic device. For example, the computer terminal can further include more or less components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 8 It does not limit the structure of the electronic device. For example, the computer terminal can further include more or less components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure.

[0191] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by programs instructing the related hardware of the terminal device, and the programs can be stored in a computer readable nonvolatile storage medium, which can include a flash disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, etc.

[0192] Embodiment 5

[0193] According to the embodiments of the present application, an embodiment of a non-volatile storage medium is also provided. Optionally, in the present embodiment, the non-volatile storage medium can be used to save the program code executed by the image segmentation processing method provided in the above embodiment 1.

[0194] Optionally, in the present embodiment, the non-volatile storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0195] Optionally, in the present embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining a to-be-processed video; determining a target region in the to-be-processed video; obtaining sub-video data displayed in the target region during playing of the to-be-processed video; and segmenting the to-be-processed video based on the time sequence characteristics of the sub-video data to obtain an image segmentation result.

[0196] Optionally, in the present embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: determining an interference factor in the to-be-processed video, wherein the interference factor includes program boxes and / or text annotations; and determining a target region in the to-be-processed video by eliminating the interference factor.

[0197] Optionally, in the present embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: performing first image segmentation processing on the to-be-processed video to obtain a first image segmentation processing result, wherein the first image segmentation processing includes grayscale processing and binary processing, and the first image segmentation processing is used to distinguish the grayscale of pixels in the to-be-processed video; performing connected domain processing on the first image segmentation processing result to obtain a second processing result, wherein the connected domain processing is used to eliminate connected domains in the to-be-processed video that do not meet predetermined requirements; calculating all connected domains in the second processing result and the constraint boxes of each connected domain in all connected domains using an open source computer vision library; and taking the content in the constraint box of a target connected domain in all connected domains as the target region, wherein the target connected domain is a connected domain containing a nodule lesion.

[0198] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: extracting a time sequence feature of the three-dimensional image in each frame of the video image group; generating a pseudo label of the two-dimensional image in each frame based on the time sequence feature, wherein the pseudo label is used for approximately labeling the two-dimensional image in each frame; training the three-dimensional segmentation network and the two-dimensional segmentation network using the pseudo label to obtain a trained three-dimensional segmentation network and a trained two-dimensional segmentation network; and segmenting the to-be-processed video using the trained two-dimensional segmentation network and the trained three-dimensional segmentation network to obtain an image segmentation result.

[0199] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining a video image group in the target region; calculating a registration parameter based on a registered two-dimensional image in the video image group, wherein the registered two-dimensional image is obtained by rigidly registering two consecutive two-dimensional images in the video image group; calculating a deformation parameter based on a cropped two-dimensional image, wherein the cropped two-dimensional image is obtained by cropping the registered two-dimensional image; and generating a time sequence feature of the two-dimensional image in each frame of the video image group according to the registration parameter and the deformation parameter.

[0200] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining a lesion segmentation label of each frame of the three-dimensional image; and fusing the time sequence feature and the lesion segmentation label to obtain the pseudo label.

[0201] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: training the three-dimensional segmentation network using the pseudo label to obtain a trained three-dimensional segmentation network; obtaining a three-dimensional feature value of each layer in the trained three-dimensional segmentation network, wherein the three-dimensional feature value is used as a segmentation guide when training the two-dimensional segmentation network; and training the two-dimensional segmentation network using the three-dimensional feature value to obtain a trained two-dimensional segmentation network.

[0202] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining a two-dimensional feature value of a first layer network output of the two-dimensional segmentation network; fusing the three-dimensional feature value and the two-dimensional feature value to obtain a fused feature value; and inputting the fused feature value to a second layer network of the two-dimensional segmentation network, wherein the second layer network is a next layer network of the first layer network.

[0203] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: compressing the three-dimensional feature values to obtain compressed feature values; calculating the compressed feature values and the two-dimensional feature values to obtain an attention weight matrix; and fusing the three-dimensional feature values and the two-dimensional feature values by using the attention weight matrix to obtain the fused feature values.

[0204] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: a cloud server receiving a request message from a client, wherein the request message carries identification information representing a to-be-processed video; the cloud server obtaining the to-be-processed video based on the identification information and determining a target region in the to-be-processed video; the cloud server obtaining sub-video data displayed in the target region during playing of the to-be-processed video; the cloud server segmenting the to-be-processed video based on a time sequence feature of the sub-video data to obtain an image segmentation result; and the cloud server returning the image segmentation result to the client.

[0205] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining a to-be-processed video and displaying the to-be-processed video on an operation interface of a device; determining a target region in the to-be-processed video in response to a request instruction sensed on the operation interface; displaying the to-be-processed video in playing on the operation interface and capturing sub-video data displayed in the target region; and displaying an image segmentation result of the to-be-processed video on the operation interface, wherein the image segmentation result is obtained by segmenting the to-be-processed video based on a time sequence feature of the sub-video data.

[0206] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining a to-be-processed video and displaying the to-be-processed video on an operation interface of a device; determining a target region in the to-be-processed video in response to a request instruction sensed on the operation interface; displaying the to-be-processed video in playing on the operation interface and capturing sub-video data displayed in the target region; and displaying an image segmentation result of the to-be-processed video on the operation interface, wherein the image segmentation result is obtained by segmenting the to-be-processed video based on a time sequence feature of the sub-video data.

[0207] The serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0208] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0209] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented in other ways. Among them, the above-mentioned device embodiments are only schematic, for example, the division of the above-mentioned units is only a logical function division, and actual implementation can have another division mode, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some interface, unit or module, which can be electrical or other forms.

[0210] The units described above as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0211] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of software functional unit.

[0212] The integrated unit, if realized in the form of software functional unit and sold or used as an independent product, can be stored in a computer readable nonvolatile storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art or the whole or part of the technical solutions can be embodied in the form of software product, and the computer software product is stored in a nonvolatile storage medium, including a plurality of instructions for making a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the method described in each embodiment of the present application. The above-mentioned nonvolatile storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and various program code storage media.

[0213] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principle of the present application, a number of improvements and refinements can be made, which should be regarded as the protection scope of the present application.

Claims

1. An image segmentation processing method characterized by, The method comprises the following steps: acquiring a to-be-processed video; determining a target region in the to-be-processed video; acquiring sub-video data displayed in the target region during playing of the to-be-processed video, wherein the sub-video data comprises a plurality of frames of image data, and part of the frames of image data comprises lesion segmentation labels; generating pseudo labels of each frame of two-dimensional image of a video image group in the sub-video data based on time sequence features of the sub-video data, and performing segmentation on the to-be-processed video by using the pseudo labels to obtain an image segmentation result; wherein, based on the time sequence features of the sub-video data, the pseudo labels of each frame of two-dimensional image of the video image group in the sub-video data are generated, comprising: acquiring two frames of data with labels in three-dimensional image in the plurality of frames of image data and lesion segmentation labels of the two frames of data; and fusing the time sequence features and the lesion segmentation labels to obtain pseudo labels of un-labeled data between the two frames of data, wherein the pseudo labels are segmentation labels for segmenting each frame of the two-dimensional image, and the pseudo labels are used for approximately labeling each frame of the two-dimensional image.

2. The method of claim 1, wherein, Determining the target region in the to-be-processed video comprises: determining interference factors in the to-be-processed video, wherein the interference factors include program boxes and / or text labels; determining the target region in the to-be-processed video by eliminating the interference factors.

3. The method of claim 1, wherein, Determining the target region in the to-be-processed video comprises: performing first image segmentation processing on the to-be-processed video to obtain a first image segmentation processing result, wherein the first image segmentation processing comprises grayscale processing and binary processing, and the first image segmentation processing is used to distinguish the gray scale of pixels in the to-be-processed video; performing connected domain processing on the first image segmentation processing result to obtain a second processing result, wherein the connected domain processing is used to eliminate connected domains in the to-be-processed video that do not meet predetermined requirements; calculating all connected domains in the second processing result and constraint boxes of each connected domain in all connected domains by using an open source computer vision library; taking content in a constraint box of a target connected domain in all connected domains as the target region, wherein the target connected domain is a connected domain containing a nodule lesion.

4. The method of claim 1, wherein, Acquiring sub-video data displayed in the target region during playing of the to-be-processed video comprises: acquiring a plurality of frames of image data displayed in the target region during playing of the to-be-processed video; constructing a video sequence based on the plurality of frames of image data to obtain the sub-video data.

5. The method of claim 4, wherein, Based on the time sequence features of the sub-video data, generating pseudo labels of each frame of two-dimensional image of a video image group in the sub-video data, and performing segmentation on the to-be-processed video by using the pseudo labels to obtain an image segmentation result, comprising: extracting time sequence features of each frame of two-dimensional image in the video image group; The pseudo label is used to train a three-dimensional segmentation network to obtain a trained three-dimensional segmentation network, and the three-dimensional feature values of each layer in the trained three-dimensional segmentation network are used to train a two-dimensional segmentation network to obtain a trained two-dimensional segmentation network, wherein the three-dimensional feature values are used as segmentation guidance when training the two-dimensional segmentation network; The trained two-dimensional segmentation network and the trained three-dimensional segmentation network are used to segment the to-be-processed video to obtain the image segmentation result.

6. The method of claim 5, wherein, The timing features include registration parameters and deformation parameters, wherein the timing features of the two-dimensional images in each frame of the video image group are extracted, including: Obtaining a video image group in the target area; Registration parameters are calculated based on the registered two-dimensional images in the video image group, wherein the registered two-dimensional images are obtained by rigid registration of two consecutive two-dimensional images in the video image group; Deformation parameters are calculated based on the cropped two-dimensional images, wherein the cropped two-dimensional images are obtained by cropping the registered two-dimensional images; According to the registration parameters and the deformation parameters, the timing features of the two-dimensional images in each frame of the video image group are generated.

7. The method of claim 5, wherein, The three-dimensional feature values are used to train a two-dimensional segmentation network to obtain a trained two-dimensional segmentation network, including: Obtaining two-dimensional feature values output by the first layer network of the two-dimensional segmentation network; The three-dimensional feature values and the two-dimensional feature values are fused to obtain fused feature values; The fused feature values are input into the second layer network of the two-dimensional segmentation network, wherein the second layer network is the next layer network of the first layer network.

8. The method of claim 7, wherein, The three-dimensional feature values and the two-dimensional feature values are fused to obtain fused feature values, including: The three-dimensional feature values are compressed to obtain compressed feature values; The compressed feature values and the two-dimensional feature values are calculated to obtain an attention weight matrix; The three-dimensional feature values and the two-dimensional feature values are fused using the attention weight matrix to obtain the fused feature values.

9. An image segmentation processing method characterized by, It includes: The cloud server receives a request message from the client, wherein the request message carries identification information representing a to-be-processed video; The cloud server obtains the to-be-processed video based on the identification information and determines a target area in the to-be-processed video; The cloud server obtains sub-video data displayed in the target area during the playback of the to-be-processed video, wherein the sub-video data includes multiple image data, and some of the multiple image data include lesion segmentation annotations; The cloud server generates pseudo labels for two-dimensional images in a video image group in the sub-video data based on the timing features of the sub-video data, and segments the to-be-processed video using the pseudo labels to obtain an image segmentation result; The cloud server returns the image segmentation result to the client; The pseudo label of each two-dimensional image frame of a video image group in the sub-video data is generated based on a time sequence feature of the sub-video data, including: obtaining two frames of data with annotations in three-dimensional images in the multiple frames of image data and lesion segmentation annotations of the two frames of data; and fusing the time sequence feature and the lesion segmentation annotations to obtain a pseudo label of unannotated data between the two frames of data, wherein the pseudo label is a segmentation label for segmenting each two-dimensional image frame, and is used for approximately annotating each two-dimensional image frame.

10. An image segmentation processing method characterized by comprising: It includes: Obtaining a to-be-processed video, and displaying the to-be-processed video on an operation interface of a device; In response to a request instruction sensed on the operation interface, determining a target region in the to-be-processed video; Displaying the to-be-processed video in play on the operation interface, and capturing sub-video data displayed in the target region, wherein the sub-video data includes multiple frames of image data, and some frames of image data in the multiple frames of image data have lesion segmentation annotations; Displaying an image segmentation result of the to-be-processed video on the operation interface, wherein the image segmentation result is obtained by generating a pseudo label of each two-dimensional image frame of a video image group in the sub-video data based on a time sequence feature of the sub-video data, and segmenting the to-be-processed video using the pseudo label; The pseudo label of each two-dimensional image frame of a video image group in the sub-video data is generated based on a time sequence feature of the sub-video data, including: obtaining two frames of data with annotations in three-dimensional images in the multiple frames of image data and lesion segmentation annotations of the two frames of data; and fusing the time sequence feature and the lesion segmentation annotations to obtain a pseudo label of unannotated data between the two frames of data, wherein the pseudo label is a segmentation label for segmenting each two-dimensional image frame, and is used for approximately annotating each two-dimensional image frame.

11. An image segmentation processing method characterized by comprising: It includes: Collecting a to-be-processed video of a lesion by a medical device, and displaying the to-be-processed video on a case operation interface of the medical device; Determining a target region in the to-be-processed video; Obtaining sub-video data displayed in the target region during playing of the to-be-processed video, wherein the sub-video data includes multiple frames of image data, and some frames of image data in the multiple frames of image data have lesion segmentation annotations; Generating a pseudo label of each two-dimensional image frame of a video image group in the sub-video data based on a time sequence feature of the sub-video data, and segmenting the to-be-processed video using the pseudo label to obtain an image segmentation result; Filling the image segmentation result into a text recording case information to obtain a structured filled case text; Displaying the structured filled case text on the case operation interface; The pseudo label of each two-dimensional image frame of a video image group in the sub-video data is generated based on a time sequence feature of the sub-video data, including: obtaining two frames of data with annotations in three-dimensional images in the multiple frames of image data and lesion segmentation annotations of the two frames of data; and performing fusion processing on the time sequence feature and the lesion segmentation annotations to obtain a pseudo label of unannotated data between the two frames of data, wherein the pseudo label is a segmentation label for segmenting each two-dimensional image frame, and is used for approximately annotating each two-dimensional image frame.

12. An image segmentation processing apparatus characterized by comprising: The method comprises: A first obtaining module is configured to obtain a to-be-processed video. A determining module is configured to determine a target region in the to-be-processed video. A second obtaining module is configured to obtain sub-video data displayed in the target region during playing of the to-be-processed video, wherein the sub-video data comprises multiple frames of image data, and some frames of image data in the multiple frames of image data have lesion segmentation annotations. A segmentation processing module is configured to generate a pseudo label of each two-dimensional image frame of a video image group in the sub-video data based on a time sequence feature of the sub-video data, and perform segmentation on the to-be-processed video by using the pseudo label to obtain an image segmentation result. The pseudo label of each two-dimensional image frame of a video image group in the sub-video data is generated based on a time sequence feature of the sub-video data, including: obtaining two frames of data with annotations in three-dimensional images in the multiple frames of image data and lesion segmentation annotations of the two frames of data; and performing fusion processing on the time sequence feature and the lesion segmentation annotations to obtain a pseudo label of unannotated data between the two frames of data, wherein the pseudo label is a segmentation label for segmenting each two-dimensional image frame, and is used for approximately annotating each two-dimensional image frame.

13. A non-volatile storage medium, comprising: The non-volatile storage medium comprises a stored program, wherein the program controls a device in which the non-volatile storage medium is located to perform the image segmentation processing method of any one of claims 1 to 10 when the program is running.

14. An electronic device, comprising: The method comprises: A processor; and A memory connected with the processor, configured to provide the processor with instructions for processing the following processing steps: Obtain a to-be-processed video. Determine a target region in the to-be-processed video. Obtain sub-video data displayed in the target region during playing of the to-be-processed video, wherein the sub-video data comprises multiple frames of image data, and some frames of image data in the multiple frames of image data have lesion segmentation annotations. Generate a pseudo label of each two-dimensional image frame of a video image group in the sub-video data based on a time sequence feature of the sub-video data, and perform segmentation on the to-be-processed video by using the pseudo label to obtain an image segmentation result. The generating pseudo labels of two-dimensional images of each frame of a video image group in the sub-video data based on a time sequence feature of the sub-video data comprises: obtaining two frames of data with annotations in three-dimensional images in the multiple frames of image data and lesion segmentation annotations of the two frames of data; and performing fusion processing on the time sequence feature and the lesion segmentation annotations to obtain pseudo labels of unannotated data between the two frames of data, wherein the pseudo labels are segmentation labels for segmenting each frame of the two-dimensional images and are used for approximately annotating each frame of the two-dimensional images.

Citation Information

Patent Citations

  • Image retrieval database establishing method

    CN104462111A

  • Childbirth monitoring method and device

    CN109657571A

  • Focus tracking method under digestive endoscope based on sequential feature learning

    CN111915573A

  • IHC digital preview recognition and organization foreground segmentation method and system

    CN112270683A