Image segmentation model training method, image segmentation method, electronic equipment and storage medium
By extracting and enhancing the timing characteristics of ultrasonic video frame groups in the image segmentation model, the problem of failure to utilize the spatial and temporal correlation of video sequences in the prior art is solved, and more accurate lesion area recognition is achieved.
Patent Information
- Application Number
- CN202510358728.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
The existing image segmentation model fails to fully utilize the temporal correlation between adjacent spatial information of the lesion region in the video sequence and continuous video frames when processing ultrasound video, resulting in a limited overall accuracy of the segmentation model.
By obtaining video frame groups with timing, labeling keyframes and extracting keyframe features and preamble frame features, performing timing enhancement, inputting the prediction network, calculating supervision loss information, and iterating the model, and using feature encoding network to enhance the timing relationship between keyframe features and preamble frame features.
The model's learning ability of image representation in lesion areas is improved and the recognition accuracy of image segmentation model is enhanced.
Smart Images

Figure CN120259341A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular, to an image segmentation model training method, an image segmentation method, an electronic device, and a storage medium. Background Art
[0002] With the rapid development of artificial intelligence and computer vision technologies, ultrasound image-assisted diagnosis technologies based on artificial intelligence have been clinically applied. Among them, convolutional neural networks have achieved remarkable results in processing static medical images, but their capabilities are limited to some extent when processing ultrasound videos with temporal information.
[0003] In related technologies, segmentation models are usually limited to independently processing single images, failing to fully utilize the adjacent spatial information of lesion regions in video sequences and also ignoring the temporal correlation between consecutive frames of videos, thereby imposing a certain degree of restriction on the overall accuracy of the segmentation model. Summary of the Invention
[0004] In view of this, the purpose of the embodiments of the present invention is to provide an image segmentation model training method, an image segmentation method, an electronic device, and a storage medium to at least partially improve the above problems.
[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, an embodiment of the present invention provides an image segmentation model training method, and the method includes: Obtain a set of video frame groups with time series; the last frame of the video frame group is a key frame, and the remaining frames of the video frame group are pre-order frames; the key frame is labeled with a true label; Input the video frame group into the feature extraction network of the image segmentation model to obtain key frame features and multiple pre-order frame features; After combining the key frame features with each of the pre-order frame features and performing temporal enhancement, input them into the prediction network of the image segmentation model to obtain a prediction result; Calculate supervised loss information based on the prediction result and the true label, and update and iterate the image segmentation model according to the supervised loss information.
[0006] Optionally, the step of inputting the video frame group into the feature extraction network of the image segmentation model to obtain key frame features and multiple pre-order frame features includes: Stack the key frame and each of the pre-order frames batch by batch to obtain a four-dimensional tensor stacked in batches; Input the four-dimensional tensor into the feature extraction network of the image segmentation model, and process multiple frames of images in the four-dimensional tensor in parallel to obtain multiple frame features; Split the multi-frame features according to the timing of the video frame group to obtain key frame features and multiple pre-order frame features.
[0007] Optionally, after combining the key frame features with each of the pre-order frame features and performing temporal enhancement, inputting them into the prediction network of the image segmentation model to obtain a prediction result, including: Input the key frame features into the first feature encoding network to obtain key frame vectors; Input each of the pre-order frame features into the second feature encoding network to obtain corresponding pre-order frame vectors; According to the key frame vectors and each of the pre-order frame vectors, perform temporal enhancement on the key frame features to obtain temporally enhanced key frame features; Input the temporally enhanced key frame features into the prediction network of the image segmentation model to obtain a prediction result.
[0008] Optionally, the performing temporal enhancement on the key frame features according to the key frame vectors and each of the pre-order frame vectors to obtain temporally enhanced key frame features includes: Calculate the similarity between each of the pre-order frame vectors and the key frame vectors to obtain the similarity of each of the pre-order frame vectors, and use each similarity as the weight of the corresponding pre-order frame feature; Multiply each of the pre-order frame features by its own weight bit by bit to obtain each weighted pre-order frame feature; Use an activation function to activate each of the weighted pre-order frame features respectively, and stack the activated results to obtain historical frame information; Stack the historical frame information and the key frame features to obtain temporally enhanced key frame features.
[0009] Optionally, the similarity is cosine similarity, and the calculation formula of the cosine similarity is:
[0010] Where is the key frame vector, The i th pre-order frame vector.
[0011] Optionally, the activation function is the sigmoid function.
[0012] Optionally, the method further includes: Obtain at least one ultrasound video data; Perform frame cutting on each of the ultrasound video data to obtain multiple video frames numbered in chronological order; Group the video frames according to the numbering order to obtain multiple groups of video frame groups with timing.
[0013] In a second aspect, an embodiment of the present invention provides an image segmentation method, the method comprising: Obtaining video frames of ultrasonic video data in real time; When the number of the video frames is greater than or equal to N, batch-stacking the current frame and the previous N - 1 video frames of the current frame to obtain stacked video frames; Inputting the stacked video frames into an image segmentation model to obtain a segmentation result of the current frame; the image segmentation model is trained by the image segmentation model training method according to any one of the first aspect.
[0014] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, where the processor implements the method according to any one of the above when executing the program.
[0015] In a fourth aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, and the computer program implements the method according to any one of the above when being executed by a processor.
[0016] An image segmentation model training method, an image segmentation method, an electronic device, and a storage medium provided by an embodiment of the present invention extract features from a video frame group to obtain key frame features and multiple pre-order frame features, and combine the key frame features with each pre-order frame feature and perform temporal enhancement, so that the temporal features in the ultrasonic video can be associated, enabling the model to further learn a more extensive image representation of the lesion area, and further making the model recognition more accurate.
[0017] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. Description of the Drawings
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 A schematic structural block diagram of the electronic device provided by an embodiment of the present invention; Figure 2 A schematic flowchart of an image segmentation model training method provided by an embodiment of the present invention; Figure 3Another flowchart of an image segmentation model training method provided by an embodiment of the present invention; Figure 4 Another flowchart of an image segmentation model training method provided by an embodiment of the present invention; Figure 5 A flowchart of step S233 provided by an embodiment of the present invention; Figure 6 A processing flowchart of an image segmentation model training method provided by an embodiment of the present invention; Figure 7 A flowchart of an image segmentation method provided by an embodiment of the present invention; Figure 8 A flowchart of image segmentation model training and image segmentation provided by an embodiment of the present invention.
[0020] Icons: 100 - electronic device; 101 - memory; 102 - communication interface; 103 - processor; 104 - communication bus. Detailed implementation manners
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. Components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0022] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0023] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, terms such as "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0024] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0025] Currently, for the detection of various lesions (such as ovarian space-occupying lesions), it mainly relies on medical images such as ultrasound, MRI, etc. In the field of ultrasound, the main way of diagnosis is that doctors observe the ultrasound images to find abnormal tissues in the ovaries and distinguish the benign and malignant nature of the tissues. However, traditional ultrasound diagnosis mainly relies on doctors' experience and subjective judgment, and is easily affected by human factors, resulting in differences in the accuracy and consistency of diagnostic results.
[0026] With the rapid development of artificial intelligence and computer vision technologies, ultrasound image-assisted diagnosis technologies based on artificial intelligence have been clinically applied. Among them, convolutional neural networks (CNNs) have achieved remarkable results in processing static medical images, but their capabilities are limited when processing ultrasound videos with temporal information. In the general diagnosis and treatment process of the ultrasound imaging department in a hospital, only a few images with clear imaging of the patient's lesion area and for the attached drawings of the ultrasound report are left, and the ultrasound video images during the patient's diagnosis are not saved. Currently, most of the ultrasound-assisted diagnostic image processing algorithms are trained and developed on the above-mentioned images, failing to fully utilize the adjacent spatial information of the lesion area in the video sequence, and at the same time ignoring the temporal correlation between consecutive frames of the video, thus imposing a certain degree of restriction on the overall accuracy of the segmentation model.
[0027] Based on the above situation, the embodiments of the present invention provide an image segmentation model training method, an image segmentation method, an electronic device and a storage medium, which can connect the temporal features in the ultrasound video, enabling the model to further learn a more extensive image representation of the lesion area, and thus making the model recognition more accurate.
[0028] To implement the process steps and functions of each example of the present invention, please refer to Figure 1 , Figure 1A schematic structural block diagram of an electronic device provided by an embodiment of the present invention. The electronic device 100 may be a model training device, including a memory 101 and a processor 103, and the memory 101 and the processor 103 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components may be electrically connected to each other through one or more communication buses 104 or signal lines. The memory 101 may be used to store software programs and modules, and the processor 103 executes the software programs and modules stored in the memory 101, thereby performing various functional applications and data processing.
[0029] The electronic device 100 may be, but is not limited to, a personal computer (PC), a server, a distributed computer, and so on. It can be understood that the electronic device 100 is not limited to a physical server, and may also be a virtual machine on a physical server, a virtual machine built on a cloud platform, etc., which can provide the same functions as the server or virtual machine. The operating system of the electronic device 100 may be, but is not limited to, the Windows system, the Linux system, and so on.
[0030] Among them, the memory 101 may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), and so on.
[0031] The communication connection between the electronic device 100 and an external device is realized through at least one communication interface 102 (which may be wired or wireless).
[0032] The processor 103 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the embodiments of the present invention can be completed by the integrated logic circuit in the hardware of the processor 103 or the instructions in the form of software. The processor 103 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0033] It can be understood that Figure 1 The structure shown is only schematic, and the electronic device 100 may also include more or fewer components than those shown Figure 1 in the figure, or have a different configuration from that shown Figure 1 in the figure. Figure 1 Each component shown in the figure can be implemented by hardware, software, or a combination thereof.
[0034] Next, an exemplary description will be given of the image segmentation model training method provided by the present invention. Figure 2 It is a schematic flowchart of an image segmentation model training method provided by an embodiment of the present invention. Refer to Figure 2 , the execution subject of this method can be the above Figure 1 shown electronic device 100, and this method includes the following steps as described Figure 2 below: S210: Obtain a set of video frame groups with time series.
[0035] Among them, the last frame of the video frame group is the key frame, and the remaining frames of the video frame group are the previous frames; the key frame is labeled with a true label.
[0036] S220: Input the video frame group into the feature extraction network of the image segmentation model to obtain the key frame feature and multiple previous frame features.
[0037] S230: After combining the key frame feature with each previous frame feature and performing temporal enhancement, input it into the prediction network of the image segmentation model to obtain a prediction result.
[0038] S240: Calculate the supervised loss information according to the prediction result and the true label, and update and iterate the image segmentation model according to the supervised loss information.
[0039] First, obtain a set of video frame groups with time series. For example, perform frame extraction and grouping on the ultrasound diagnosis video of the ovarian region, obtain one of the video frame groups, and perform sequential update and iteration. Assume that the number of frames in each group is n, then the nth frame is defined as the key frame, the nth frame is the key frame annotated by the annotating doctor, and the first n - 1 frames are defined as the previous frames. Extract features from the key frame and each previous frame in the video frame group to obtain the corresponding key frame features and multiple previous frame features. Combine each previous frame feature with the key frame feature and input them into the prediction network to obtain the prediction result. Calculate the supervised loss information between the prediction result and the true label of the key frame, and perform backpropagation of the reverse gradient on the image segmentation model according to the supervised loss information to update the parameters of the image segmentation model. A set of video frame groups is one iteration, and multiple sets of video frame groups can be used for multiple iterations to obtain a better image segmentation model.
[0040] By extracting features from the key frame and each previous frame of the video frame group, obtaining the key frame features and multiple previous frame features, and combining the key frame features with each previous frame feature, the time series features in the ultrasound video can be linked, enabling the model to further learn a more extensive image representation of the lesion area, and thus making the image segmentation model more accurate in recognition.
[0041] Before training the image segmentation model, multiple sets of video frame group data can also be prepared first. This image model training method can also include: Obtain at least one ultrasound video data.
[0042] Perform frame extraction on each ultrasound video data to obtain multiple video frames numbered in chronological order.
[0043] Group the video frames in the order of the numbers to obtain multiple sets of video frame groups with time series.
[0044] First, collect ultrasound video data of the gynecological ovarian region from multiple hospitals, and then perform patient information removal and desensitization processing on the collected videos. For each ultrasound video data, first collect all the frames in the video and number them in chronological order, and group them according to the frame number order. The number of frames within a group can be selected according to actual needs, but the number of frames between groups remains the same. Further, assume that the number of frames in each group is n, then the nth frame is defined as the key frame, and the first n - 1 frames are defined as the previous frames. Thus, multiple sets of video frame groups with time series can be obtained.
[0045] Only medical staff such as doctors need to mark the lesion area on the key frames to outline the lesion range. The marking doctor can observe the specific lesion characteristics of the patient based on all the frames of the video, but only needs to mark the lesion area on the key frames. This method not only utilizes all the frame information of the video to provide the doctor with a complete picture of the patient's lesions, avoiding the drawbacks of previous methods that marked on single lesion images, but also saves the workload of the annotator.
[0046] There are various ways to extract the features of a video frame group. For example, each frame image in the video frame group can be sequentially input into the feature extraction network of the image segmentation model to obtain the features of each frame image. Features can also be extracted in parallel for each frame image. To accelerate the feature extraction speed, in this embodiment, the key frame and each previous frame are batch stacked and then feature extraction is performed. See Figure 3 , the above S220 may include the following steps: S221: Batch stack the key frame and each previous frame to obtain a four-dimensional tensor of batch stacking.
[0047] S222: Input the four-dimensional tensor into the feature extraction network of the image segmentation model, and process multiple frame images in the four-dimensional tensor in parallel to obtain multiple frame features.
[0048] S223: Split the multiple frame features according to the time sequence of the video frame group to obtain the key frame features and multiple previous frame features.
[0049] For a group of consecutive video frame groups as input, the key frame and the previous frames are batch stacked to obtain a four-dimensional tensor including multiple frame images. The batch-stacked four-dimensional tensor is input into the feature extraction network. The feature extraction network processes multiple frame images in parallel and outputs multiple frame features, assumed to be F. Split the multiple frame features in the order of the previous frames and the key frame to obtain the features of the key frame and the previous frames, respectively assumed to be F k and (F p1 , F p2 , F p3 , …, F pi ), where i is the number of previous frames used. If the number of frames in the video frame group is n, then i = n - 1.
[0050] There are various ways to increase the temporal relationship between the key frame features and the features of each previous frame. For example, the features can be directly added or concatenated to obtain new features. To accelerate the processing speed of each feature and the training efficiency, the previous frame features can be processed first and then the key frame features can be enhanced, and then predictions can be made on the enhanced key frame features. See Figure 4 , the above step S230 may include the following steps: S231: Input the key frame features into the first feature encoding network to obtain the key frame vector.
[0051] S232: Input each previous frame feature into the second feature encoding network to obtain corresponding previous frame vectors for each one.
[0052] S233: According to the key frame vector and each previous frame vector, perform temporal enhancement on the key frame feature to obtain a temporally enhanced key frame feature.
[0053] S234: Input the temporally enhanced key frame feature into the prediction network of the image segmentation model to obtain a prediction result.
[0054] Input the feature of the key frame into the first feature encoding network. The first feature encoding network is an independent algorithm network, which extracts based on the key frame feature and outputs a multi-dimensional query vector, called the key frame vector, denoted as A. Input all previous frame features into the second feature encoding network. The second feature encoding network is an algorithm network independent of the second feature encoding network, but it is shared for all previous frame features, that is, the same second feature encoding network is used for all previous frame features. The second feature encoding network outputs a multi-dimensional vector to be queried for each feature of each previous frame, called the previous frame vector, denoted as (B1, B2, B3,..., B i ). Among them, the structural parameters of the first feature encoding network and the second feature encoding network can be the same or different, but the same second feature encoding network will be input for each previous frame feature. After obtaining the key frame vector and each previous frame vector, the relationship between each vector can be added to the key frame feature for temporal enhancement according to the key frame vector and each previous frame vector to obtain a temporally enhanced key frame feature. Input this temporally enhanced key frame feature into the prediction network of the image segmentation model to obtain a prediction result.
[0055] There can be various ways to perform temporal enhancement on the key frame feature. For example, directly fuse each previous frame feature into the key frame feature. In order to obtain the temporal correlation in the video frame group and not cause a large change to the key frame feature, in a possible implementation, see Figure 5 , step S233 may include the following steps: S2331: Calculate the similarity between each previous frame vector and the key frame vector to obtain the similarity of each previous frame vector, and use each similarity as the weight of the corresponding previous frame feature.
[0056] Among them, the similarity can be Euclidean distance, Manhattan distance, cosine similarity, etc. In this embodiment, cosine similarity is used. The calculation formula of cosine similarity is:
[0057] Among them, is the key frame vector, the i th previous frame vector.
[0058] S2332: Multiply each pre-frame feature by its own weight bit by bit to obtain each weighted pre-frame feature.
[0059] S2333: Use an activation function to activate each weighted pre-frame feature respectively, and stack the activated results to obtain historical frame information.
[0060] S2334: Stack the historical frame information and the key-frame feature to obtain the temporally enhanced key-frame feature.
[0061] Calculate the cosine similarity between each pre-frame vector and the key-frame vector as the weight of each pre-frame feature. The weights of all pre-frame features can be expressed as ( , , , …, ). Multiply the weights of all pre-frame features by themselves bit by bit, and use an activation function to limit the weight range to the interval (0, 1). This activation function can be the sigmoid function, and stack the results of this calculation to obtain historical frame information. This processing process can be represented by the following formula: F h = Concat(sigmoid(F pm · ), m = 1, 2, 3, …, n - 1) where F h is the historical frame information, F pn is the pre-frame feature, is the weight of the pre-frame feature, m is the m-th pre-frame, n is the number of frames in the video frame group, sigmoid is the activation function, and Concat is the stacking operation.
[0062] Finally, stack the historical frame information and the key-frame feature to obtain the temporally enhanced key-frame feature, which can be represented by the following formula: F z = Concat(F h , F k ) where F z is the temporally enhanced key-frame feature, F h is the historical frame information, and F k is the key-frame feature.
[0063] Exemplarily, refer to Figure 6, the processing process of the video frame group is described. First, the annotations of the key frame and each previous frame are stacked and input into the feature extraction network to obtain multi-frame features. The multi-frame features are split to obtain the key frame feature and each previous frame feature. The key frame feature is input into the first feature encoding network to obtain the key frame vector, and each previous frame is input into the second feature encoding network to obtain each previous frame vector. The similarity between each previous frame vector and the key frame vector is calculated to obtain the weight of each previous frame. The previous frame feature is multiplied by its own previous frame weight bit by bit, and the sigmoid function is used to limit the weight range to the interval (0, 1). The calculation results of all previous frames are stacked to obtain the historical frame information. Then, the historical frame information is stacked with the key frame feature to obtain the temporally enhanced key frame feature. Finally, the temporally enhanced key frame feature is input into the prediction network to obtain the final key frame prediction output.
[0064] Further, an embodiment of the present invention also provides an image segmentation method. Refer to Figure 7 , the method includes: S310: Real-time obtain video frames of ultrasonic video data.
[0065] S320: When the number of video frames is greater than or equal to N, batch-stack the current frame and the previous N - 1 video frames of the current frame to obtain a stacked video frame.
[0066] S330: Input the stacked video frame into the image segmentation model to obtain the segmentation result of the current frame; wherein, the image segmentation model is trained by the above-mentioned image segmentation model training method.
[0067] After training a mature image segmentation model, it can be used for predicting space-occupying lesions in ultrasonic videos. For example, first capture the ultrasonic input video and create a queue to save each frame image of the video. If the number of consecutive video frames set during the training of the image segmentation model is n, the length of the queue is n. Count and save the real-time frames of the video. If the current real-time frame number is less than n, only save the real-time frame into the queue and do not call the model for prediction. If the current frame is the nth real-time frame received, stack the current frame with the previous n - 1 frames in the save queue, input it into the image segmentation model and perform prediction to obtain the segmentation result of the current frame (i.e., the nth frame). If the serial number of the current frame is greater than n, delete the first saved frame in the save queue and add the current frame to the save queue. Read all the frames in the save queue, where the end of the queue is the current frame. Stack the current frame with the previous n - 1 frames in the save queue, input it into the image segmentation model and perform prediction to obtain the segmentation result of the current frame, and obtain the segmentation result of the current frame. Until the segmentation result of the last frame of the ultrasonic input video is output.
[0068] Exemplarily, refer to Figure 8, the solution will be described. Generally, it can include three modules: data preparation, model training, and model prediction. The data preparation module includes three parts: video data collection, video frame grouping, and key frame annotation. After the data is prepared, model training is carried out. The input image segmentation model is applied to each previous frame and key frame of the video frame group to obtain the features of each previous frame and key frame. Then, the features of each previous frame and key frame are used to enhance the key frame features temporally, and the enhanced key frame features are predicted. The supervision loss is calculated by comparing with the key frame annotation. Finally, the image segmentation model is optimized iteratively by backpropagating the supervision loss information. After the image segmentation model is trained, the ultrasound video can be predicted and segmented using the image segmentation model. The ultrasound video input is captured in real time, and the previous frames are saved until the number of saved frames reaches n, and the segmentation results of the key frames are continuously output.
[0069] In summary, the image segmentation model training method, image segmentation method, electronic device, and storage medium provided by the embodiments of the present invention extract features from the video frame group to obtain key frame features and multiple previous frame features, and combine the key frame features with each previous frame feature and enhance them temporally, so that the temporal features in the ultrasound video can be related, enabling the model to further learn a more extensive image representation of the lesion area, and thus making the model recognition more accurate.
[0070] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0071] In addition, in each embodiment of the present invention, the functional modules can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0072] When a function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, optical discs, and other various media that can store program codes.
[0073] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0074] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. A method for training an image segmentation model, characterized in that The method includes: Obtaining a set of video frame groups with time series; the last frame of the video frame group is a key frame, and the remaining frames of the video frame group are pre-order frames; the key frame is labeled with a true label; Inputting the video frame group into the feature extraction network of an image segmentation model to obtain key frame features and multiple pre-order frame features; After combining the key frame features with each of the pre-order frame features and performing time series enhancement, inputting them into the prediction network of the image segmentation model to obtain a prediction result; Calculating supervised loss information based on the prediction result and the true label, and updating and iterating the image segmentation model according to the supervised loss information.
2. The method according to claim 1, wherein The step of inputting the video frame group into the feature extraction network of the image segmentation model to obtain key frame features and multiple pre-order frame features includes: Stacking the key frame and each of the pre-order frames batch by batch to obtain a four-dimensional tensor of batch stacking; Inputting the four-dimensional tensor into the feature extraction network of the image segmentation model, and processing multiple frames of images in the four-dimensional tensor in parallel to obtain multiple frame features; Splitting the multiple frame features according to the time series of the video frame group to obtain key frame features and multiple pre-order frame features.
3. The method according to claim 1, characterized in that, The step of combining the key frame features with each of the pre-order frame features, performing time series enhancement, and then inputting them into the prediction network of the image segmentation model to obtain a prediction result includes: Inputting the key frame features into a first feature encoding network to obtain key frame vectors; Inputting each of the pre-order frame features into a second feature encoding network to obtain corresponding pre-order frame vectors; Performing time series enhancement on the key frame features according to the key frame vectors and each of the pre-order frame vectors to obtain time series enhanced key frame features; Inputting the time series enhanced key frame features into the prediction network of the image segmentation model to obtain a prediction result.
4. The method according to claim 3, wherein The step of performing time series enhancement on the key frame features according to the key frame vectors and each of the pre-order frame vectors to obtain time series enhanced key frame features includes: Calculating the similarity between each of the pre-order frame vectors and the key frame vector to obtain the similarity of each of the pre-order frame vectors, and using each similarity as the weight of the corresponding pre-order frame feature; Multiplying each of the pre-order frame features by its own weight bit by bit to obtain each weighted pre-order frame feature; Activating each of the weighted pre-order frame features using an activation function respectively, and stacking the activated results to obtain historical frame information; Stacking the historical frame information and the key frame features to obtain time series enhanced key frame features.
5. The method according to claim 4, characterized in that, The similarity is a cosine similarity, and the calculation formula of the cosine similarity is: Among them, is the key frame vector, the i th previous frame vector.
6. The method according to claim 4, wherein The activation function is a sigmoid function.
7. The method according to claim 1, characterized in that, The method further includes: Obtaining at least one ultrasound video data; Performing frame cutting on each of the ultrasound video data to obtain multiple video frames numbered in time order; Grouping the video frames according to the numbering order to obtain multiple groups of video frame groups with time series.
8. An image segmentation method, characterized in that, The method includes: Obtaining video frames of ultrasound video data in real time; When the number of the video frames is greater than or equal to N, stacking the current frame and the previous N - 1 video frames of the current frame batch by batch to obtain stacked video frames; Input the stacked video frames into the image segmentation model to obtain the segmentation result of the current frame; the image segmentation model is trained by the image segmentation model training method according to any one of claims 1 to 7.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 8.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 8.
Citation Information
Cited By
Class activation graph generation model training method, generation method and related equipment
CN120672750A