Target key frame identification method for ultrasonic video stream and related device
Patent Information
- Application Number
- CN202311200697.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-15
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-09-15
AI Technical Summary
[0005]为解决上述技术问题,本发明提供了超声视频流的目标关键帧识别方法和相关设备,解决了现有技术不能针对多目标的视频流识别出每个目标的关键帧图像的问题
[0049]有益效果:本发明首先对视频流的每帧图像均应用目标提取算法,得到每帧图像上的所有目标所在的所有区域,再采集各个目标区域的各组指标,最后依据每帧图像上的各组指标,从视频流中的各帧图像中识别出每一个目标所对应的关键帧图像。从上述分析可知,由于本发明针对每个目标在图像上的区域都进行了指标计算,因此本发明可以从视频流中识别出每个目标所对应的关键帧图像。
Smart Images

Figure CN117197716B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition technology, specifically to a method and related equipment for identifying target keyframes in ultrasonic video streams. Background Technology
[0002] Medical diagnostic equipment can determine the specific location of a patient's lesion or lesion (target) by analyzing ultrasound images. For example, multiple ultrasound images are sequentially acquired by a probe to form a video stream. Current technology selects one frame from the video stream as a keyframe, and analyzes this keyframe to help doctors understand the severity of the patient's condition. However, current technology only selects one keyframe image from all lesions. When multiple targets appear in the video stream, current technology cannot select a keyframe for each target from the video stream.
[0003] In summary, existing technologies cannot identify the keyframe images of each target in a multi-target video stream.
[0004] Therefore, existing technologies still need to be improved and enhanced. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a method and related equipment for identifying keyframes of targets in ultrasonic video streams, solving the problem that existing technologies cannot identify keyframe images of each target in video streams with multiple targets.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] In a first aspect, the present invention provides a method for identifying target keyframes in an ultrasonic video stream, comprising:
[0008] Target extraction is performed on each frame of the video stream to obtain the target regions where each target is located in each frame of the image. The target regions are used to characterize the position of the target in the image.
[0009] Determine a set of indicators for each target region on each frame of the image, the indicators being used to identify the target;
[0010] Based on the indicators in each frame of the image, the keyframe image corresponding to each target is identified from each frame of the video stream.
[0011] In one implementation, target extraction is performed on each frame of the video stream to obtain target regions where each target is located in each frame of the image. These target regions characterize the position of the target in the image and include:
[0012] Determine the current frame image in each frame of the image;
[0013] Extract the current depth feature map of the current frame image;
[0014] A group of preceding depth feature maps is determined, consisting of individual preceding depth feature maps, wherein each preceding depth feature map is a depth feature map of a previous frame image, and each previous frame image is located before the current frame image in the video stream;
[0015] Determine the current estimated position of each of the targets in the current frame image;
[0016] Based on the current estimated position of each target, the current depth feature map, and the preceding depth feature map group, target extraction is performed on the current frame image to obtain the current target region in each target region of the current frame image. Each current target region is used to characterize the position of each target in the current frame image after adjustment relative to the current estimated position of the target.
[0017] In one implementation, the step of extracting targets from the current frame image based on the current estimated position of each target, the current depth feature map, and the preceding depth feature map group, to obtain the current region of each target in the target region of each target in the current frame image, wherein each current region of the target is used to characterize the position of each target in the current frame image after adjustment relative to the current estimated position of the target, includes:
[0018] The current depth feature map and the previous depth feature map group are fused to obtain a fused depth feature map;
[0019] The current estimated position of each target and the current depth feature map are fused to obtain a position-depth fusion feature map, which is used to characterize the enhanced depth features of each target's current estimated position on the current frame image.
[0020] Based on the location depth fusion feature map and the fusion depth feature map, target extraction is performed on the current frame image to obtain the current region of each target in the current frame image.
[0021] In one implementation, the method for updating the preceding deep feature map set includes:
[0022] The current regions of each target and the current frame image are input into the classification network to obtain the target region accuracy output by the classification network. The target region accuracy is used to characterize the degree to which the current target region covers the target on the current frame image.
[0023] Determine the total number of feature maps formed by each of the preceding deep feature maps within the preceding deep feature map group;
[0024] Determine the current image index number of the current frame image on the video stream;
[0025] The image index number corresponding to the preceding deep feature map with the largest sequence number within the preceding deep feature map group is determined and denoted as the latest image index number;
[0026] Based on the total number of feature maps, the precision of the target region, the current image index number, and the latest image index number, determine whether to update the preceding depth feature map group.
[0027] In one implementation, determining whether to update the preceding depth feature map group based on the total number of feature maps, the precision of the target region, the current image index number, and the latest image index number includes:
[0028] Determine the difference obtained by subtracting the current image index number from the latest image index number;
[0029] Determine an exponential function with the natural constant as the base and the difference mentioned above as the exponent;
[0030] The sum obtained by adding the precision of the target region to the exponential function is determined;
[0031] When the total number of feature maps is greater than the storage threshold, and the sum of the feature maps is greater than the set update threshold, the preceding depth feature map corresponding to the smallest sequence number in the preceding depth feature map group is deleted, and the current depth feature map is placed in the preceding depth feature map group to update the preceding depth feature map group.
[0032] In one implementation, the current estimated position of each of the targets is obtained through a tracker group, and the tracker group is updated in the following ways:
[0033] Each of the targets is identified as a corresponding sub-tracker, and each of the sub-trackers is used to form the tracker group;
[0034] When there are target regions that do not appear in each of the target regions, a tracker is added to the target regions that do not appear, and it is called a new tracker. The target regions that do not appear are the regions that did not appear in the previous frame image.
[0035] Add the new tracker to the tracker group to update the tracker group.
[0036] In one implementation, identifying the keyframe image corresponding to each target from each frame of the video stream based on each set of indicators on each frame of the image includes:
[0037] Determine each sub-indicator in each group of indicators for each target region;
[0038] Based on the identification results of each sub-indicator and / or the preset importance level for each sub-indicator, a target score is obtained for each target region, and the identification results are used to characterize the index value of the sub-indicator;
[0039] The target scores of each of the previous frame images are compared to obtain the highest target score, and the previous frame image corresponding to the highest target score is recorded as the preferred frame image;
[0040] When the target score of the current frame image is greater than the highest target score, the current frame image is used as the keyframe image of the target corresponding to the target score;
[0041] When the target score of the current frame image is less than or equal to the highest target score, the preferred frame image is used as the keyframe image of the target corresponding to the target score;
[0042] The target scores of each remaining frame image of the video stream are compared with the target scores of the keyframe images. When there is a target score of a remaining frame image that is greater than the target score of the keyframe image, the final keyframe image of each target is selected from the remaining frame images.
[0043] Secondly, embodiments of the present invention also provide a target keyframe recognition device for a video stream, wherein the device comprises the following components:
[0044] The target extraction module is used to extract targets from each frame of the video stream to obtain the target regions where each target is located in each frame of the image. The target regions are used to characterize the position of the target in the image.
[0045] An indicator acquisition module is used to determine a set of indicators for each target region on each frame of the image, the indicators being used to identify the target.
[0046] The keyframe recognition module is used to identify the keyframe image corresponding to each target from each frame image in the video stream based on each set of indicators on each frame image.
[0047] Thirdly, embodiments of the present invention also provide a terminal device, wherein the terminal device includes a memory, a processor, and a target keyframe recognition program for an ultrasonic video stream stored in the memory and executable on the processor. When the processor executes the target keyframe recognition program for the ultrasonic video stream, it implements the steps of the target keyframe recognition method for the ultrasonic video stream described above.
[0048] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a target keyframe recognition program for an ultrasonic video stream. When the target keyframe recognition program for the ultrasonic video stream is executed by a processor, it implements the steps of the target keyframe recognition method for the ultrasonic video stream described above.
[0049] Beneficial Effects: This invention first applies a target extraction algorithm to each frame of the video stream to obtain all regions where all targets are located in each frame. Then, it collects various sets of indicators for each target region. Finally, based on these indicators, it identifies the keyframe image corresponding to each target from each frame of the video stream. As can be seen from the above analysis, because this invention calculates indicators for each target's region in the image, it can identify the keyframe image corresponding to each target from the video stream. Attached Figure Description
[0050] Figure 1 This is an overall flowchart of the present invention;
[0051] Figure 2 This is a flowchart illustrating the identification of keyframe images in an embodiment of the present invention;
[0052] Figure 3 This is a block diagram illustrating the internal structure of a terminal device provided in an embodiment of the present invention. Detailed Implementation
[0053] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments and accompanying drawings. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0054] Research has shown that medical diagnostic equipment can determine the specific location of lesions or lesions (targets) in a patient by analyzing ultrasound images. For example, multiple ultrasound images are sequentially acquired by a probe to form a video stream. Current technology selects one frame from the several frames of the video stream as a keyframe, and analyzes this keyframe to help doctors understand the severity of the patient's condition. However, current technology only selects one keyframe image from all lesions. When multiple targets appear in the video stream, current technology cannot select a keyframe for each target from the video stream.
[0055] To address the aforementioned technical problems, this invention provides a method and related equipment for identifying keyframes of targets in ultrasonic video streams, solving the problem that existing technologies cannot identify the keyframe image of each target in a multi-target video stream. Specifically, firstly, target extraction is performed on each frame of the video stream to obtain the target regions (target regions characterize the target's position in the image) of each target in each frame; then, a set of indicators (indicators are used to identify targets) are determined for each target region in each frame; finally, based on the indicators in each frame, the keyframe image corresponding to each target is identified from each frame of the video stream. This invention can identify the keyframe image of each target in a multi-target video stream.
[0056] For example, when acquiring ultrasound images of the human thyroid gland, a video stream consisting of several frames (e.g., four frames: A, B, C, and D) is obtained. If three lesions appear in the video stream (i.e., three lesions on the thyroid gland, i.e., three targets A, B, and C), then it is necessary to identify the keyframe image corresponding to each target from the video stream. The keyframe image is the one among the frames that is most helpful for the doctor to diagnose the severity of the lesion (target). This invention identifies keyframe images in the following way:
[0057] For example, for target A, the regions where target A is located are extracted from four frames of images, namely A, B, C, and D. That is, a region belonging to target A is extracted from each frame of the image (assuming target A is present in all four frames of the image). This results in four target regions of target A. Then, four sets of indicators corresponding to the four target regions of target A are calculated. If the set of indicators corresponding to frame A has the highest score (the score is related to the importance of the indicator and the recognition result), then frame A is the keyframe image of target A.
[0058] Perform the same operation as described above on both target B and target C to obtain keyframe images belonging to target B and keyframe images belonging to target C.
[0059] Exemplary methods
[0060] The target keyframe identification method for ultrasound video streams in this embodiment can be applied to terminal devices, which can be terminal products with image acquisition capabilities, such as ultrasound diagnostic instruments. In this embodiment, for example... Figure 1 As shown, the target keyframe identification method for the ultrasonic video stream specifically includes the following steps:
[0061] S100, target extraction is performed on each frame of the video stream to obtain each target region where each target is located in each frame of the image, and the target region is used to characterize the position of the target in the image.
[0062] S200, determine each set of indicators for each target region on each frame of the image, the indicators being used to identify the target.
[0063] S300, based on the indicators in each frame of the image, identify the key frame image corresponding to each target from each frame of the video stream.
[0064] In one embodiment, the image in step S100 is a real-time acquired ultrasound image, wherein the ultrasound image is acquired by an ultrasound diagnostic device, and the ultrasound diagnostic device operates as follows:
[0065] Preset virtual buttons can be set on the touch screen of the ultrasound diagnostic equipment, or physical buttons can be set on the operation panel of the ultrasound diagnostic equipment. When the doctor clicks the virtual button or the physical button, the ultrasound diagnostic equipment will start to acquire ultrasound images of the patient through the probe.
[0066] In one embodiment, step S100 includes the following specific steps S101 to S107:
[0067] S101, determine the current frame image in each frame of the image.
[0068] S102, Extract the current depth feature map of the current frame image.
[0069] In this embodiment, such as Figure 2 As shown, an artificial neural network model is used to extract depth features from the image to improve the robustness of the extracted depth features. In this embodiment, the convolution in the neural network model is a 2D convolution. 2D convolution is only used to extract depth features and cannot extract the temporal information (information between different frames) of the current frame image in the video stream. This temporal information is obtained through subsequent steps (i.e.,...). Figure 2 The information is obtained from the preceding feature fusion module and the tracking result fusion module. In this embodiment, 2D convolution and two fusion modules are used to extract depth information and temporal information respectively. Compared with 3D convolution, which extracts both depth information and temporal information, the former can reduce resource consumption.
[0070] S103, determine a group of preceding depth feature maps consisting of each preceding depth feature map, wherein each preceding depth feature map is a depth feature map of a previous frame image, and each previous frame image is located before the current frame image in the video stream.
[0071] Step S102 is to extract depth features from the current frame image. Step S103 uses the same method as S102 to extract depth feature maps from each previous frame image, and stores the depth feature maps of each previous frame image in the form of a group of previous depth feature maps in the previous feature storage module.
[0072] S104, determine the current estimated position of each of the targets on the current frame image.
[0073] In this embodiment, the approximate position of each target in the current frame image is calculated by the target tracking algorithm built into the tracker. The tracking algorithm usually learns the change pattern of the target by using the position information of the target in the previous frame image, thereby estimating the approximate position of the target in the current frame image.
[0074] One target corresponds to only one tracker in the video stream. If no target was extracted in a previous frame (i.e., the number of trackers is 0), no tracking is performed in the current frame. During tracking, the current frame is input into a multi-target tracker, and each tracker outputs the corresponding target's bounding box in the current frame. Available tracking algorithms include, but are not limited to, the KCF tracking algorithm. Tracking allows prediction of the target's approximate location in the current frame, which helps in the accurate extraction of subsequent targets.
[0075] For example, if the current frame is the fifth frame in the video stream, then the first to fourth frames are the previous frames. If there are two targets in the first to fourth frames, then two target trackers are set up to track the two targets in the current frame.
[0076] S105, the current depth feature map and the previous depth feature map group are fused to obtain a fused depth feature map.
[0077] The current depth feature map contains the target's features in the current frame image, while the preceding depth feature map contains the target's features in different previous frame images. In other words, all preceding depth feature maps record the target's feature change information (i.e., time sequence information). Therefore, the fused depth feature map obtained after fusing the current depth feature map and the preceding depth feature map contains all the target's features from the first previous frame image to the current frame image.
[0078] S106, the current estimated position of each target and the current depth feature map are fused to obtain a position-depth fusion feature map, which is used to characterize the enhanced depth features of each target's current estimated position on the current frame image.
[0079] Apply the estimated current location and current depth feature map of a target in the current frame image as follows: Figure 2 The tracking result fusion module shown obtains a location-depth fusion feature map. A location-depth fusion feature map is used to record the depth features of a target in an image. For example, if there is a target in the middle of an image, fusing the middle location of the target with the image's depth feature map results in a location-depth fusion feature map that records the depth features at the middle location of the image.
[0080] The fusion method in step S106 includes, but is not limited to, attention-based feature fusion, wherein the attention-based feature fusion method is as follows:
[0081] The mask of the target tracker's detection box is concatenated with the current depth feature map of the current frame image. A convolution module and a sigmoid function are then applied to the concatenation result to obtain an attention map. The attention map is then multiplied with the current depth feature map to obtain the fused feature.
[0082] The mask is obtained as follows: a zero matrix M with the same resolution as the current frame image is set, the tracking boxes of all targets are drawn on M and filled with 1s, and then scaled down to the resolution of the current depth feature map to obtain the mask. If no target is tracked in the current frame, i.e. the tracking box is empty, then the value of each element on the mask is 0.
[0083] S107, Based on the location depth fusion feature map and the fusion depth feature map, target extraction is performed on the current frame image to obtain the current region of each target in the current frame image.
[0084] The location-depth fusion feature map and the fusion-depth feature map are input into the target extraction module, which outputs the region where each target is located in the current frame image (i.e., the current region of each target). The target extraction module includes, but is not limited to, the decoder module of the target segmentation network or the detection module of the target detection network.
[0085] The preceding depth feature map group in step S103 and the tracker in step S104 are updated based on the extraction results of the target extraction module. In one embodiment, the updating of the preceding depth feature map group includes the following specific steps S1031 to S1038:
[0086] S1031, the current regions of each target and the current frame image are input into the classification network to obtain the target region accuracy score1 output by the classification network. The target region accuracy is used to characterize the degree to which the current region of the target covers the target on the current frame image.
[0087] The classification network, but not limited to the Mobilenetv3 classification network, takes the target region extracted from the current frame image and the current frame image as input to the Mobilenetv3 classification network. The Mobilenetv3 classification network outputs the target region precision score1, and the value of score1 ranges from (0,1). The closer score1 is to 1, the better the extraction effect, that is, the closer the target region extracted from the current frame image is to the actual region of the target in the current frame image.
[0088] S1032, determine the total number of feature maps formed by each of the preceding deep feature maps within the preceding deep feature map group.
[0089] The total number of feature maps is Figure 2 The total number of feature maps stored in the preceding feature storage module is limited by the storage capacity of the preceding feature storage module.
[0090] S1033, determine the current image index number i of the current frame image on the video stream.
[0091] A video stream consists of several frames, each frame of which is assigned an index number, with adjacent index numbers differing by one.
[0092] S1034, determine the image index number corresponding to the preceding deep feature map with the largest sequence number within the preceding deep feature map group, and denot it as the latest image index number.
[0093] For example, the preceding depth feature map group contains preceding depth feature maps with index numbers 1, 5, 8, and 12. Among them, index number 12 is the largest index number, and the image index number corresponding to this index number is 12.
[0094] S1035, determine the difference index-i obtained by subtracting the current image index from the latest image index number index.
[0095] S1036, Determine the exponential function e with the natural constant e as the base and the difference index-i as the exponent. index-i .
[0096] S1037, determine the target region accuracy score1 plus the exponential function e. index-i The obtained sum score:
[0097] score = score1 + e index-i
[0098] S1038, when the total number of feature maps is greater than the storage threshold T1 and the sum score is greater than the set update threshold T2 (a constant set manually), delete the preceding depth feature map corresponding to the smallest sequence number inside the preceding depth feature map group, and place the current depth feature map in the preceding depth feature map group to update the preceding depth feature map group.
[0099] For example, if the preceding deep feature map group contains preceding deep feature maps with indices 1, 5, 8, and 12, and the total number of feature maps is greater than the storage threshold T1 and the score is greater than T2, then the preceding deep feature map with index 1 is deleted, and the current deep feature map is added to the preceding deep feature map group.
[0100] In one embodiment, the specific process of updating the tracker is as follows: determine each sub-tracker corresponding to each target, and each sub-tracker is used to form the tracker group; when there is a target region that has not appeared in each target region, add a tracker for the target region that has not appeared, and denot it as a new tracker, the target region that has not appeared is the region that has not appeared in the previous frame image; add the new tracker to the tracker group to update the tracker group.
[0101] In other words, it determines whether a new target region has been extracted in the current frame (a new target region is a target region that did not appear in previous frames). If a new target region is extracted, a tracker is added and initialized for that target. If one or more trackers fail to track the target in the current frame, the corresponding trackers are deleted. This process of adding and deleting trackers completes the tracker update.
[0102] In one embodiment, when the image is a breast ultrasound image, the indicators in step S200 include the shape, orientation, edge, internal echo, posterior echo, and calcification of the lesion (target area); when the image is a thyroid image, the indicators in step S200 include the edge, composition, morphology, internal echo, and strong echo of the lesion.
[0103] In one embodiment, step S300, which involves identifying the keyframe image corresponding to each target from the video stream, includes the following specific steps S301 to S305:
[0104] S301, determine each sub-indicator in each group of indicators for each target region.
[0105] For example, if the video stream is a video stream of the human breast area, there are two target regions in the current frame image of the video stream. For either of these two target regions, various sub-indicators of the region are collected. The sub-indicators include the shape, orientation, edge, internal echo, posterior echo, and calcification of the target region.
[0106] S302, based on the identification results of each sub-indicator and / or the preset importance level for each sub-indicator, a target score for each target region is obtained, wherein the identification results are used to characterize the index value of the sub-indicator.
[0107] For example, the recognition result of the shape sub-index is whether the shape is an ellipse, circle, or irregular shape, etc., and different index values are assigned to different shapes. Importance is determined by assigning different levels of importance to the aforementioned sub-indexes such as shape, orientation, edge, internal echo, posterior echo, and calcification based on experience. A deep learning algorithm is applied to the recognition results and importance of shape, orientation, edge, internal echo, posterior echo, and calcification for each target region to obtain a target score for each target region in the current frame image. Alternatively, a deep learning algorithm can be applied to either the recognition result or the importance to obtain a target score for each target region in the current frame image.
[0108] S303, compare the target scores of each of the previous frame images to obtain the highest target score, and record the previous frame image corresponding to the highest target score as the preferred frame image.
[0109] S304, when the target score of the current frame image is greater than the highest target score, the current frame image is used as the keyframe image of the target corresponding to the target score.
[0110] Alternatively, when the target score of the current frame image is less than or equal to the highest target score, the preferred frame image is used as the keyframe image of the target corresponding to the target score.
[0111] Step S302 obtains the target score for each target region in the current frame image (denoted as frame j). For example, if there are two targets A and B in the current frame image, meaning there are two lesions, the region of target A in the current frame image is region a1, and the region of target B in the current frame image is region b1. The target score for region a1 corresponding to target A is Sa1, and the target score for region b1 corresponding to target B is Sb1. The previous frame image containing the highest target score Samax for target A in each previous frame image is frame j-5 (preferred frame image), and the previous frame image containing the highest target score Sbmax for target B in each previous frame image is frame j-3 (preferred frame image). If Sa1 is greater than Samax, then the current frame image is the keyframe image for target A; otherwise, frame j-5 is the keyframe image for target A. If Sb1 is greater than Sbmax, then the current frame image is the keyframe image for target B; otherwise, frame j-3 is the keyframe image for target B.
[0112] S305, compare the target scores of each remaining frame image of the video stream with the target scores of the keyframe images. When there is a target score of a remaining frame image that is greater than the target score of the keyframe image, select the final keyframe image of each target from the remaining frame images.
[0113] For example, if the remaining frames are frame j+1, ..., frame j+m, and the target score of target A in frame j+2 is greater than the target score of target A in the keyframe image in step S304, then frame j+2 is the final keyframe image of target A. Otherwise, the keyframe image of target A in step S304 is the final keyframe image of target A.
[0114] In summary, this invention first applies a target extraction algorithm to each frame of the video stream to obtain all regions where all targets are located in each frame. Then, it collects various sets of indicators for each target region. Finally, based on these indicators, it identifies the keyframe image corresponding to each target from each frame of the video stream. As can be seen from the above analysis, because this invention calculates indicators for each region, it can identify the keyframe image corresponding to each target from the video stream.
[0115] Exemplary device
[0116] This embodiment also provides a target keyframe recognition device for a video stream, the device comprising the following components:
[0117] The target extraction module is used to extract targets from each frame of the video stream to obtain the target regions where each target is located in each frame of the image. The target regions are used to characterize the position of the target in the image.
[0118] An indicator acquisition module is used to determine a set of indicators for each target region on each frame of the image, the indicators being used to identify the target.
[0119] The keyframe recognition module is used to identify the keyframe image corresponding to each target from each frame image in the video stream based on each set of indicators on each frame image.
[0120] Based on the above embodiments, the present invention also provides a terminal device, the principle block diagram of which can be as follows: Figure 3 As shown, the terminal device includes a processor, memory, network interface, display screen, and temperature sensor connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a target keyframe recognition method for ultrasonic video streams. The display screen can be an LCD screen or an e-ink screen. The temperature sensor is pre-installed inside the terminal device to detect the operating temperature of the internal components.
[0121] Those skilled in the art will understand that Figure 3 The block diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal device to which the present invention is applied. The specific terminal device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0122] In one embodiment, a terminal device is provided, comprising a memory, a processor, and a target keyframe recognition program for an ultrasonic video stream stored in the memory and executable on the processor. When the processor executes the target keyframe recognition program for the ultrasonic video stream, it implements the following operation instructions:
[0123] Target extraction is performed on each frame of the video stream to obtain the target regions where each target is located in each frame of the image. The target regions are used to characterize the position of the target in the image.
[0124] Determine a set of indicators for each target region on each frame of the image, the indicators being used to identify the target;
[0125] Based on the indicators in each frame of the image, the keyframe image corresponding to each target is identified from each frame of the video stream.
[0126] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying target keyframes in an ultrasonic video stream, characterized in that, include: Target extraction is performed on each frame of the video stream to obtain the target regions where each target is located in each frame of the image. The target regions are used to characterize the position of the target in the image, and the target is the lesion. Determine each set of indicators for each target region on each frame of the image, the indicators being used for The indicators for identifying the target include the shape, orientation, edges, internal echoes, posterior echoes, and calcification of the target area; Based on the indicators in each frame of the image, the key frame image corresponding to each target is identified from each frame of the video stream. The key frame image is an image that is helpful for diagnosing lesions. The process involves extracting targets from each frame of the video stream to obtain target regions where each target is located in each frame. These target regions characterize the position of the targets in the image and include: Determine the current frame image in each frame of the image; Extract the current depth feature map of the current frame image; A group of preceding depth feature maps is determined, wherein each preceding depth feature map is a depth feature map of a previous frame image, and each previous frame image is located before the current frame image in the video stream. All preceding depth feature maps record the feature change information of the target. Determine the current estimated position of each of the targets in the current frame image; The current depth feature map and the previous depth feature map group are fused to obtain a fused depth feature map; The current estimated position of each target and the current depth feature map are fused to obtain a position-depth fusion feature map, which is used to characterize the enhanced depth features of each target's current estimated position on the current frame image. Based on the location depth fusion feature map and the fusion depth feature map, target extraction is performed on the current frame image to obtain the current region of each target in the current frame image.
2. The target keyframe identification method for ultrasonic video streams as described in claim 1, characterized in that, The update method for the preceding deep feature map group includes: The current regions of each target and the current frame image are input into the classification network to obtain the target region accuracy output by the classification network. The target region accuracy is used to characterize the degree to which the current target region covers the target on the current frame image. Determine the total number of feature maps formed by each of the preceding deep feature maps within the preceding deep feature map group; Determine the current image index number of the current frame image on the video stream; The image index number corresponding to the preceding deep feature map with the largest sequence number within the preceding deep feature map group is determined and denoted as the latest image index number; Based on the total number of feature maps, the precision of the target region, the current image index number, and the latest image index number, determine whether to update the preceding depth feature map group.
3. The target keyframe identification method for ultrasonic video streams as described in claim 2, characterized in that, The step of determining whether to update the preceding depth feature map group based on the total number of feature maps, the precision of the target region, the current image index number, and the latest image index number includes: Determine the difference obtained by subtracting the current image index number from the latest image index number; Determine an exponential function with the natural constant as the base and the difference mentioned above as the exponent; The sum obtained by adding the precision of the target region to the exponential function is determined; When the total number of feature maps is greater than the storage threshold, and the sum of the feature maps is greater than the set update threshold, the preceding depth feature map corresponding to the smallest sequence number in the preceding depth feature map group is deleted, and the current depth feature map is placed in the preceding depth feature map group to update the preceding depth feature map group.
4. The target keyframe identification method for ultrasonic video streams as described in claim 1, characterized in that, The estimated current position of each of the targets is obtained through a tracker group, and the tracker group is updated in the following ways: Each of the targets is identified as a corresponding sub-tracker, and each of the sub-trackers is used to form the tracker group; When there are target regions that do not appear in each of the target regions, a tracker is added to the target regions that do not appear, and it is called a new tracker. The target regions that do not appear are the regions that did not appear in the previous frame image. Add the new tracker to the tracker group to update the tracker group.
5. The target keyframe identification method for ultrasonic video streams as described in claim 1, characterized in that, The step of identifying the keyframe image corresponding to each target from each frame of the video stream based on each set of indicators on each frame of the image includes: Determine each sub-indicator in each group of indicators for each target region; Based on the identification results of each sub-indicator and / or the preset importance level for each sub-indicator, a target score is obtained for each target region, and the identification results are used to characterize the index value of the sub-indicator; The target scores of each of the previous frame images are compared to obtain the highest target score, and the previous frame image corresponding to the highest target score is recorded as the preferred frame image; When the target score of the current frame image is greater than the highest target score, the current frame image is used as the keyframe image of the target corresponding to the target score; When the target score of the current frame image is less than or equal to the highest target score, the preferred frame image is used as the keyframe image of the target corresponding to the target score; The target scores of each remaining frame image of the video stream are compared with the target scores of the keyframe images. When there is a target score of a remaining frame image that is greater than the target score of the keyframe image, the final keyframe image of each target is selected from the remaining frame images.
6. A target keyframe recognition device for a video stream, characterized in that, The device comprises the following components: The target extraction module is used to extract targets from each frame of the video stream to obtain the target regions where each target is located in each frame of the image. The target regions are used to characterize the position of the target in the image, and the target is the lesion. The indicator acquisition module is used to determine each set of indicators for each target region on each frame of the image. The indicators are used to identify the target and include the shape, orientation, edge, internal echo, posterior echo, and calcification of the target region. The keyframe recognition module is used to identify the keyframe image corresponding to each target from each frame of the video stream based on each set of indicators on each frame of the image. The keyframe image is an image that is helpful for diagnosing lesions. The process involves extracting targets from each frame of the video stream to obtain target regions where each target is located in each frame. These target regions characterize the position of the targets in the image and include: Determine the current frame image in each frame of the image; Extract the current depth feature map of the current frame image; A group of preceding depth feature maps is determined, wherein each preceding depth feature map is a depth feature map of a previous frame image, and each previous frame image is located before the current frame image in the video stream. All preceding depth feature maps record the feature change information of the target. Determine the current estimated position of each of the targets in the current frame image; The current depth feature map and the previous depth feature map group are fused to obtain a fused depth feature map; The current estimated position of each target and the current depth feature map are fused to obtain a position-depth fusion feature map, which is used to characterize the enhanced depth features of each target's current estimated position on the current frame image. Based on the location depth fusion feature map and the fusion depth feature map, target extraction is performed on the current frame image to obtain the current region of each target in the current frame image.
7. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a target keyframe recognition program for an ultrasonic video stream stored in the memory and executable on the processor. When the processor executes the target keyframe recognition program for the ultrasonic video stream, it implements the steps of the target keyframe recognition method for an ultrasonic video stream as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a target keyframe recognition program for an ultrasonic video stream. When the target keyframe recognition program for the ultrasonic video stream is executed by a processor, it implements the steps of the target keyframe recognition method for an ultrasonic video stream as described in any one of claims 1-5.
Citation Information
Patent Citations
Structured target detection method and device, equipment and storage medium
CN114663648A