Method, device, electronic device and readable storage medium for displaying subject detection frame

By using different extraction frequencies and frame image calculation methods in the video stream, the problem of rapid offset of the main detection frame position is solved, and a more stable main detection frame display is achieved, which improves the user experience and reduces the computing burden.

CN115601775BActive Publication Date: 2025-08-29BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211303281.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-24
Publication Date
2025-08-29
Estimated Expiration
2042-10-24

AI Technical Summary

Technical Problem

During the subject detection process, there may be a large difference in the positions between adjacent subject detection boxes, resulting in the displayed subject detection box position rapidly offset and affecting the user experience.

Method used

By using different extraction frequencies in the frame image sequence of the video stream, the display frame image and the detection frame image are extracted respectively, and the position offset of the main display frame is calculated in the front and back frame images of the main detection frame detection frame to reduce the position offset of the main display frame between adjacent frame images.

Benefits of technology

It effectively reduces position jitter of the main detection box, improves user experience, and reduces the computing overhead of the main detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115601775B_ABST
    Figure CN115601775B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for displaying a subject detection frame, which relates to the field of artificial intelligence technology, specifically the field of deep learning, image processing, and computer vision technology. The specific implementation scheme is: obtaining a frame image sequence in a video stream captured for a target object; extracting a display frame image from the frame image sequence based on a preset first extraction frequency, and extracting a detection frame image from the frame image sequence based on a preset second extraction rate; in response to determining a second detection frame image from the detection frame image, for any target second display frame image in the second display frame image, based on the subject display frame in the previous display frame image of the target second display frame image, and the subject detection frame in the second detection frame image, determine the subject display frame in the target second display frame image. The present disclosure can reduce the offset of the subject display frames of the previous and next display frame images, reduce the jitter of the picture, and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically deep learning, image processing, and computer vision technology, and can be applied to scenarios such as OCR (Optical Character Recognition). Background Art

[0002] Subject detection is a widely used technique used to detect the location of one or more subjects within an image, and then crop the image region corresponding to the subject for subsequent image processing. For example, subject detection can be performed on images of paper bills to detect the image region containing the bill subject and extract the bill's structural information.

[0003] In related technologies, users can capture images with a camera and then perform subject detection on the captured images to obtain subject detection results. This subject detection process takes some time, during which the subject detection results can be displayed to the user in real time to enhance the user experience. Subject detection results are typically displayed in the form of subject detection frames.

[0004] When multiple consecutive subject detection frames are displayed in sequence, there may be a large difference in the positions of two adjacent subject detection frames, which causes the positions of the displayed subject detection frames to shift rapidly. Summary of the Invention

[0005] In order to solve at least one of the above-mentioned deficiencies, the present disclosure provides a method for displaying a subject detection frame, a display device for a subject detection frame, an electronic device, and a computer program product.

[0006] According to a first aspect of the present disclosure, a method for displaying a subject detection frame is provided, the method comprising:

[0007] Acquire a frame image sequence in a video stream captured of a target object;

[0008] Extracting a display frame image from the frame image sequence based on a preset first extraction frequency, and extracting a detection frame image from the frame image sequence based on a preset second extraction rate, wherein the first extraction frequency is greater than the second extraction frequency, a display frame image located between two adjacent detection frame images in the frame image sequence is associated with a previous detection frame image of the two adjacent detection frame images, the detection frame images are used to perform subject detection on a subject corresponding to the target object, and the display frame images are used to perform display, and in response to a subject detection frame corresponding to the target object being detected in a first detection frame image among the detection frame images, a subject display frame determined based on the subject detection frame in the first detection frame image is displayed in a first display frame image associated with the first detection frame image;

[0009] In response to determining a second detection frame image from the detection frame image, for any target second display frame image in the second display frame image, based on the subject display frame in the previous frame display frame image of the target second display frame image and the subject detection frame in the second detection frame image, the subject display frame in the target second display frame image is determined, wherein subject detection frames are detected in both the second detection frame image and the previous frame detection frame image of the second detection frame image, the second display frame image is a detection frame image associated with the second detection frame image, and a first offset between the subject display frame in the previous frame display frame image of the target second display frame image and the subject display frame in the target second display frame image is less than a second offset between the subject display frame in the previous frame display frame image of the target second display frame image and the subject detection frame in the second detection frame image.

[0010] According to a second aspect of the present disclosure, a device for displaying a subject detection frame is provided, the device comprising:

[0011] An image sequence acquisition module is used to acquire a frame image sequence in a video stream captured of a target object;

[0012] a frame image extraction module, configured to extract display frame images from the frame image sequence based on a preset first extraction frequency, and extract detection frame images from the frame image sequence based on a preset second extraction rate, wherein the first extraction frequency is greater than the second extraction frequency, a display frame image located between two adjacent detection frame images in the frame image sequence is associated with a preceding detection frame image of the two adjacent detection frame images, the detection frame images are used to perform subject detection on a subject corresponding to the target object, and the display frame images are used to display, and in response to a subject detection frame corresponding to the target object being detected in a first detection frame image among the detection frame images, a subject display frame determined based on the subject detection frame in the first detection frame image is displayed in a first display frame image associated with the first detection frame image;

[0013] A frame image calculation module is used to, in response to determining a second detection frame image from the detection frame image, determine, for any target second display frame image in the second display frame image, a subject display frame in the target second display frame image based on the subject display frame in the previous display frame image of the target second display frame image and the subject detection frame in the second detection frame image, wherein subject detection frames are detected in both the second detection frame image and the previous detection frame image of the second detection frame image, the second display frame image is a detection frame image associated with the second detection frame image, and a first offset between the subject display frame in the previous display frame image of the target second display frame image and the subject display frame in the target second display frame image is less than a second offset between the subject display frame in the previous display frame image of the target second display frame image and the subject detection frame in the second detection frame image.

[0014] According to a third aspect of the present disclosure, an electronic device is provided, including:

[0015] at least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute the method for displaying the subject detection frame.

[0018] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above-mentioned method for displaying a subject detection frame.

[0019] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the above-mentioned method for displaying a subject detection frame when executed by a processor.

[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0022] Figure 1 This is a flow chart of a method for displaying a subject detection frame provided by an embodiment of the present disclosure;

[0023] Figure 2 This is a flowchart illustrating some steps of another method for displaying a subject detection frame provided by an embodiment of the present disclosure;

[0024] Figure 3 This is a flowchart illustrating some steps of another method for displaying a subject detection frame provided by an embodiment of the present disclosure;

[0025] Figure 4 This is a flowchart illustrating some steps of another method for displaying a subject detection frame provided by an embodiment of the present disclosure;

[0026] Figure 5 This is a flowchart illustrating some steps of another method for displaying a subject detection frame provided by an embodiment of the present disclosure;

[0027] Figure 6 A flowchart of a specific embodiment of a method for displaying a subject detection frame provided by an embodiment of the present disclosure;

[0028] Figure 7 1 is a schematic structural diagram of a display device for a subject detection frame provided by an embodiment of the present disclosure;

[0029] Figure 8 4 is a block diagram of an electronic device for implementing the method for displaying a subject detection frame according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0030] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0031] In related technologies, users can capture images through a camera device, and then perform subject detection on the captured images to obtain subject detection results. The images captured by the camera device are generally output in the form of a video stream. When performing subject detection, images are generally extracted from the video stream at a certain frequency as detection frame images, and then the extracted detection frame images are input into the subject detection model in turn, thereby obtaining the subject detection results of each detection frame image.

[0032] The process of inputting the detection frame image into the subject detection model to obtain the subject detection result takes some time. During this period, the user experience can be improved by displaying the detection results of the detection model to the user in real time.

[0033] In actual use, the robustness of the subject detection model may not be strong. There may be large differences in the positions of the subject detection frames in two adjacent detection frame images detected by the subject detection model. This will cause the position of the displayed subject detection frame to shift rapidly, causing the display image seen by the user to be violently jittery, affecting the user experience.

[0034] The subject detection frame display method, device, electronic device, and computer-readable storage medium provided by the embodiments of the present disclosure are intended to solve at least one of the above technical problems in the prior art.

[0035] The subject detection frame display method provided in the embodiments of the present disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be an in-vehicle device, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, an in-vehicle device, a wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in a memory. Alternatively, the method can be executed by a server.

[0036] Figure 1 FIG. 1 shows a flow chart of a method for displaying a subject detection frame provided by an embodiment of the present disclosure, such as Figure 1 As shown in , the method may mainly include:

[0037] Step S110: Acquire a frame image sequence in a video stream captured of the target object.

[0038] For example, in step S110 , the target object may be a text carrier of important structured information such as a bill.

[0039] In some possible implementations, the video stream captured of the target object may specifically be a video stream composed of multiple frame images of the target object captured by the user through a camera device. These frame images are arranged in the order of capture time from earliest to latest to form a frame image sequence.

[0040] The camera device may be a mobile phone or other terminal device with a camera function.

[0041] Step S120: extracting a display frame image from the frame image sequence based on a preset first extraction frequency; and extracting a detection frame image from the frame image sequence based on a preset second extraction frequency;

[0042] The first extraction frequency is greater than the second extraction frequency; the display frame image located between two adjacent detection frame images in the frame image sequence is associated with the previous detection frame image of the two adjacent detection frame images;

[0043] The detection frame image is used to perform subject detection on the subject corresponding to the target object;

[0044] The display frame image is used for display. In response to the subject detection frame corresponding to the target object detected in the first detection frame image in the detection frame image, the first display frame image associated with the first detection frame image will display the subject display frame determined based on the subject detection frame in the first detection frame image.

[0045] The detection frame image is used to perform subject detection on the subject corresponding to the target object. Through subject detection, it is obtained whether the detection frame image contains the target object, and the position of the subject corresponding to the target object in the detection frame image. The position of the subject corresponding to the target object in the detection frame image is represented in the form of a subject detection frame.

[0046] In some possible implementations, the subject detection frame is the maximum circumscribed rectangle of the subject corresponding to the target object, and the position of the subject corresponding to the target object in the detection frame image is recorded by saving the pixel coordinates of the four vertices of the circumscribed matrix in the detection frame image.

[0047] In some possible implementations, subject detection is performed on the subject corresponding to the target object by inputting a detection frame image into a pre-trained subject detection model for detecting whether the frame image includes the detection subject and the specific position of the detection subject in the detection frame image.

[0048] In some possible implementations, in order to reduce the computational overhead of subject detection, a detection frame image is extracted from the frame image sequence every 100 ms.

[0049] The display frame images are used for display. In some possible implementations, in order to avoid screen freezes during the display of the display frame images, the display frame images may be extracted at a frequency of 30 frames per second.

[0050] Since the first extraction frequency of the detection frame image is greater than the second extraction frequency of the display frame image, there is generally at least one display frame image between two adjacent detection frame images, and this at least one display frame image is associated with the previous detection frame image between the two adjacent detection frame images.

[0051] In some specific implementations, a detection frame image is extracted from a frame image sequence every 100 ms. When the frequency of extracting display frame images is 30 frames per second, the frame image sequence has 1-20 images, and the 10th, 10th, and 20th frame images are extracted as detection frame images. Taking the 10th frame image as an example, the detection frame image and the 20th frame image are adjacent detection frame images, and the 3rd, 6th, 9th, 12th, 15th, and 18th frame images are extracted as display frame images. Then, the display frame images between the 10th frame image and the 20th frame image are the 12th, 15th, and 18th frame images. These display frame images are all display frame images associated with the 10th frame image; similarly, the 3rd, 6th, and 9th frame images are all display frame images associated with the 1st frame image.

[0052] Obviously, when the acquisition time interval between adjacent frame images in a frame image sequence is short, the camera position, camera angle, etc. when the camera equipment acquires the target object generally will not change greatly, the image difference between adjacent frame images is small, and the change in the position of the target object in adjacent frame images is also small.

[0053] Both the detection frame image and the display frame image are extracted from a frame image sequence in a video stream captured from the target object. The time interval between the acquisition time corresponding to the display frame image and the acquisition time corresponding to its associated detection frame image is smaller than the time interval between the acquisition times corresponding to two adjacent detection frame images.

[0054] When the detection frame images are extracted based on the second extraction frequency (such as extracting one frame every 100ms), the acquisition time difference between adjacent detection frame images is only 100ms, the image difference between adjacent detection frame images is small, and the position change of the target object in adjacent detection frame images is also small. The acquisition time difference between the display frame image associated with the detection frame image and the detection frame image is smaller than the acquisition time difference between the adjacent detection frame images. Therefore, the image difference between the display frame image associated with the detection frame image and the detection frame image is also small.

[0055] That is to say, when the detection frame image detects the target object, there is a high probability that the target object can be detected in the display frame image corresponding to the detection frame image, and the position of the target object in the detection frame image is slightly different from the position of the target object in the display frame image corresponding to the detection frame image.

[0056] Therefore, when the first detection frame image in the detection frame image is detected as the subject detection frame corresponding to the target object, there is a high probability that the target object can be detected in the display frame image associated with the first detection frame image (i.e., the first display frame image), and the position of the target object is close to the position of the subject detection frame. Therefore, the subject display frame of the first display frame image can be determined according to the position of the subject detection frame. The subject display frame is used to display the position of the target object in the first display frame image. Like the subject detection frame, it is also a rectangle. The subject display frame can be saved by saving the pixel coordinates of the four vertices of the rectangle in the display frame image.

[0057] Step S130: In response to determining the second detection frame image from the detection frame images, for any target second display frame image in the second display frame images, determining a subject display frame in the target second display frame image based on a subject display frame in a display frame image immediately preceding the target second display frame image and a subject detection frame in the second detection frame image;

[0058] In which, the subject detection frame is detected in both the second detection frame image and the previous frame detection frame image of the second detection frame image, the second display frame image is a detection frame image associated with the second detection frame image, and the first offset between the subject display frame in the previous frame display frame image of the target second display frame image and the subject display frame in the target second display frame image is less than the second offset between the subject display frame in the previous frame display frame image of the target second display frame image and the subject detection frame in the second detection frame image.

[0059] When a detection frame image is detected as a subject detection frame, and the previous detection frame image of the detection frame image is also detected as a subject detection frame, then the detection frame image is the second detection frame image, and the display frame image associated with the detection frame image is the target second display frame image. For any target second display frame image, its subject display frame is determined based on the subject detection frame of the second detection frame image and the subject display frame in the previous display frame image of the target second display frame image.

[0060] Among them, the difference between the position of the main display frame in the previous frame display frame image of the target second display frame image in the previous frame display frame image of the target second display frame image and the position of the main display frame in the target second display frame image in the target second display frame image is smaller than the difference between the position of the main display frame in the previous frame display frame image of the target second display frame image in the previous frame display frame image of the target second display frame image and the position of the main detection frame in the second detection frame image in the second detection frame image.

[0061] Similarly, taking the extraction of detection frame images from the frame image sequence every 100ms, and the extraction frequency of display frame images being 30 frames per second as an example, the frame image sequence has frame images 1-100, and the 10th, 20th, etc. frame images are extracted as detection frames. Taking the 10th frame image as an example, if the 1st frame image and the 10th frame image are both detected in the subject detection frame, then the 10th frame image can be used as the second detection frame image.

[0062] For the display frame images associated with the 10th frame image, that is, the 12th, 15th, and 18th frame images are the second display frame images, for the 12th frame image, the 9th frame image is its previous display frame, and the main body display frame of the 12th frame image is determined according to the main body display frame of the 9th frame image and the main body detection frame of the 10th frame image; for the 15th frame image, the 12th frame image is its previous display frame, and the main body display frame of the 15th frame image is determined according to the main body display frame of the 12th frame image and the main body detection frame of the 10th frame image; and so on, for the 18th frame image, the 15th frame image is its previous display frame, and the main body display frame of the 18th frame image is determined according to the main body display frame of the 15th frame image and the main body detection frame of the 10th frame image.

[0063] When a target second display frame image is the target second display frame image whose acquisition time is closest to the second detection frame image among all target second display frame images, then the display frame image preceding the target second display frame image is the display frame image associated with the detection frame image preceding the second detection frame image, and its main body display frame is determined based on the main body detection frame of the detection frame image preceding the second detection frame image, and is close in position to the main body detection frame of the detection frame image preceding the second detection frame image. The main body display frame of the target second display frame image should be close in position to the main body detection frame of the second detection frame image.

[0064] When the camera device shakes between the previous frame detection frame image and the second detection frame image of the second detection frame image, the position of the subject detection frame of the previous frame detection frame image of the second detection frame image and the position of the subject detection frame of the second detection frame image are greatly different. For example, the position of the subject detection frame of the previous frame detection frame image of the second detection frame image is biased to the left, and the position of the subject detection frame of the second detection frame image is biased to the right. If the subject display frame of the target second display frame image is determined only based on the subject detection frame of the second detection frame image, the subject display frame of the previous display frame image of the target second display frame image will be biased to the left, and the subject display frame of the target second display frame image will be biased to the right. The positions of the subject display frames of the previous and next display frame images will shift rapidly, causing the display image seen by the user to shake violently, affecting the user experience.

[0065] When determining the main body display frame of the target second display frame image, not only the main body detection frame in the second detection frame image is considered, but also the main body display frame in the previous frame display frame image of the target second display frame image is considered, so that the first offset between the main body display frame in the previous frame display frame image of the target second display frame image and the main body display frame in the target second display frame image is less than the second offset between the main body display frame in the previous frame display frame image of the target second display frame image and the main body detection frame in the second detection frame image, which is equivalent to making the change in the position of the main body display frame of the previous frame display image of the target second display frame image and the position of the main body display frame of the target second display frame image less than the change in the position of the main body detection frame of the second detection frame image and the position of the main body display frame of the previous frame display image of the target second display frame image, thereby reducing the degree of offset of the positions of the main body display frames of the previous and next display frame images, reducing the jitter of the picture, and thereby improving the user experience.

[0066] Based on a similar idea, when a target second display frame image is not the target second display frame image whose acquisition time is closest to the second detection frame image among all target second display frame images, then the previous display frame image of the target second display frame image is the display frame image associated with the second detection frame image.

[0067] When determining the main body display frame of the target second display frame image, the first offset between the main body display frame in the previous display frame image of the target second display frame image and the main body display frame in the target second display frame image is made smaller than the second offset between the main body display frame in the previous display frame image of the target second display frame image and the main body detection frame in the second detection frame image. This is equivalent to making the position of the main body display frame of the target second display frame image closer to the position of the main body detection frame of the second detection frame image than the position of the main body display frame of the previous display frame image of the target second display frame image, that is, closer to the possible position of the target object of the target second display frame image.

[0068] The target second display frame image located after the target second display frame image is determined based on the main display frame of the target second display frame image and the main detection frame of the second detection frame image. Then, the position of the main display frame of the target second display frame image located after the target second display frame image is closer to the position of the main detection frame of the second detection frame image than the position of the main display frame of the target second display frame image.

[0069] By analogy, the position of the main display frame of each target second display frame image will be closer to the position of the main detection frame of the second detection frame image than the position of the main display frame of the previous target second display frame image, which is equivalent to the position of the main display frame of each target second display frame image will be closer to the possible position of the target object of the target second display frame image than the position of the main display frame of the previous target second display frame image.

[0070] At the same time, since the position of the main display frame moves slowly from the second display frame image of the first frame target to the second display frame image of the last frame target, it avoids the jitter of the display image caused by the large movement of the position of the main display frame between the previous and next display frame images, affecting the user experience.

[0071] That is to say, in the display method of the subject detection frame provided by the embodiment of the present disclosure, the detection frame image is extracted at a sampling frequency lower than the sampling frequency of the display frame image, and the subject detection frame is detected in both the second detection frame image and the detection frame image before the second detection frame image. Based on the subject display frame in the display frame image before the target second display frame image associated with the second detection frame image and the subject detection frame in the second detection frame image, the subject display frame in the target second display frame image is determined. When there is a large difference in the position of the subject detection frames in two adjacent detection frame images, the position of the subject display frame between the front and rear display frame images does not follow the subject detection frame in the detection frame image and shift rapidly, but the shift is "dispersed" to multiple display frame images, thereby reducing the shift of the subject display frames of the front and rear display frame images, reducing the jitter of the picture, and improving the user experience.

[0072] At the same time, the extraction frequency of the display frame image in the embodiment of the present disclosure is greater than the extraction frequency of the detection frame image, which reduces the number of detection frame images. Compared with performing subject detection on each display frame, the computational overhead of subject detection is reduced.

[0073] Figure 2 FIG. 1 is a flow chart showing some steps of another method for displaying a subject detection frame provided by an embodiment of the present disclosure, such as Figure 2 As shown in , in an optional manner of the present disclosure, the step of determining the subject display frame in the target second display frame image based on the subject display frame in the previous display frame image of the target second display frame image and the subject detection frame in the second detection frame image includes:

[0074] Step S210: Determine the vertex coordinates of the main display frame in the target second display frame image based on the vertex coordinates of the main display frame in the previous frame display frame image of the target second display frame image and the first weight, and based on the vertex coordinates of the main display frame in the second detection frame image and the second weight, where the sum of the first weight and the second weight is 1, and the first weight is a value greater than 0 and less than 1.

[0075] Taking a vertex as an example, the vertex coordinates corresponding to the main display frame of the target second display frame image can be calculated using the following formula:

[0076] current.x=front_point.x*weight+detect_point.x*(1-weight);

[0077] current.y=front_point.y*weight+detect_point.y*(1-weight);

[0078] Among them, front_point.x is the x-coordinate of the corresponding vertex of the main display frame of the previous frame display frame image of the target second display frame image, and front_point.y is the y-coordinate of the corresponding vertex of the main display frame of the previous frame display frame image of the target second display frame image; detect_point.x is the x-coordinate of the corresponding vertex of the main detection frame in the second detection frame image, and front_point.y is the y-coordinate of the corresponding vertex of the main detection frame in the second detection frame image; weight is the first weight, which is a value greater than 0 and less than 1, and 1-weight is the second weight; current_point.x is the x-coordinate of the corresponding vertex of the main display frame of the target second display frame image, and current_point.y is the y-coordinate of the corresponding vertex of the main display frame of the target second display frame image.

[0079] The coordinates of each vertex of the main body display frame of the target second display frame image are calculated using the above formula to obtain the main body display frame of the target second display frame image.

[0080] Among them, weight is a parameter that can be adjusted according to the actual display and detection situation of the target object. By adjusting the weight, it is possible to control whether the main display frame of the target second display frame image is closer to the main detection frame in the second detection frame image or closer to the main display frame of the previous frame display frame image of the target second display frame image (that is, control the offset degree of the main display frame of the adjacent display frame images), so that the display effect of the main display frame of the target second display frame image becomes smooth, there will be no severe jitter, and the visual perception is more comfortable.

[0081] Figure 3 A flowchart of some steps of another method for displaying a subject detection frame provided by an embodiment of the present disclosure is shown as follows: Figure 3 As shown in , in an optional manner of the present disclosure, the method for displaying the subject detection frame may further include:

[0082] Step S310, in response to determining a third detection frame image from the detection frame image, based on the subject detection frame in the third detection frame image, determine the subject display frame in the third display frame image associated with the third detection frame image, wherein the subject detection frame is detected in the third detection frame image, and the subject detection frame is not detected in the previous detection frame image of the third detection frame image.

[0083] Similarly, taking the extraction of detection frames from a frame image sequence every 100ms, and the extraction frequency of display frame images being 30 frames per second as an example, the frame image sequence has frames 1-100, and the 1st, 10th, 20th, etc. frames are extracted as detection frames. Taking the 10th frame image as an example, if the subject detection frame is not detected in the 1st frame image, but is detected in the 10th frame image, then the 10th frame image can be used as the third detection frame image. Similarly, since there are no other detection frames before the 1st frame image, when the subject detection frame is detected in the 1st frame image, it can also be used as the third detection frame image.

[0084] For the display frame images associated with the 10th frame image, that is, the 12th, 15th, and 18th frame images are the subject display frames of the third display frame images, which are determined based on the subject detection frame of the 10th frame image.

[0085] When a third display frame image is the third display frame image whose acquisition time is closest to the third detection frame image among all third display frame images, then the display frame image immediately preceding the third display frame image is the display frame image associated with the detection frame image immediately preceding the third detection frame image, and its subject display frame is determined based on the subject detection frame of the detection frame image immediately preceding the third detection frame image. Since the detection frame image immediately preceding the third detection frame image has no subject detection frame detected, the display frame image immediately preceding the third display frame image also has no corresponding subject display frame. Regardless of the position of the subject display frame of the third display frame image, it will change compared to the display frame image immediately preceding it. Therefore, the subject display frame of the third display frame image associated with the third detection frame image can be determined by the subject detection frame of the third detection frame image, that is, it can be close to the possible position of the target object.

[0086] Based on a similar idea, when a third display frame image is not the third display frame image whose acquisition time is closest to the third detection frame image among all third display frame images, the previous display frame image of the third display frame image is the display frame image associated with the third detection frame image, and the corresponding main display frame is determined by the main detection frame of the third detection frame image. The main display frame of the third display frame image is also determined by the main detection frame of the third detection frame image, which can ensure that the position of the main display frame from the previous display frame image of the third display frame image to the main display frame of the third display frame image does not change, eliminating the jitter of the main display frame and improving the user's visual experience.

[0087] Figure 4 A flowchart of some steps of another method for displaying a subject detection frame provided by an embodiment of the present disclosure is shown as follows: Figure 4 As shown in , in an optional manner of the present disclosure, the method for displaying the subject detection frame may further include:

[0088] Step S410: In response to the continuous detection of subject detection frames in the detection frame image, which includes not less than a first preset number of first target detection frame images, a display style of the subject display frame in the display frame image associated with the first target detection frame image is determined based on the confidence level of the subject detection frame in each first target detection frame image.

[0089] Performing subject detection on a detection frame image may include detecting a subject detection frame of the detection frame image and a confidence level of the subject detection frame (i.e., the probability that the subject detection frame contains the target object; a greater confidence level indicates a greater probability that the subject detection frame contains the target object). The confidence level is a value greater than 0 and less than or equal to 1.

[0090] In some specific implementations, only when the confidence level of the subject detection frame is greater than a preset threshold is it considered that the subject detection frame exists in the detection frame image.

[0091] When a subject detection frame is detected in continuous detection frame images, a display style of the subject display frame in the display frame images associated with the detection frame images is determined based on the confidence of the subject detection frame detected by the detection frame images.

[0092] Similarly, taking the case where a detection frame image is extracted from a frame image sequence every 100 ms, the frequency of display frame image extraction is 30 frames per second, and the first preset number is 1, the frame image sequence has 1-100 frames, and the 1st, 10th, 20th, etc. frames are extracted as detection frames. When the 1st frame image is not detected as a subject detection frame, and the 10th and 20th frames image are both detected as subject detection frames, then the 10th and 20th frames image can be used as the first target detection frame image. The display style of the subject display frame associated with the 10th frame (i.e., the 12th, 15th, and 18th frames image) is determined by the confidence level of the subject detection frame detected in the 10th frame image, and the display style of the subject display frame in the display frame images associated with the 20th frame image (i.e., the 21st, 24th, and 27th frames image) is determined by the confidence level of the subject detection frame detected in the 10th and 20th frames image.

[0093] Since the display style of the subject display frame is determined by the confidence of the subject detection frame, the user can judge the confidence of the subject detection frame based on the display style of the subject display frame. When the confidence of the subject detection frame is high, it means that the camera device has a high probability of having captured a high-quality subject containing the target object, and the user can continue to maintain the shooting angle; when the confidence of the subject detection frame is low, it means that the quality of the subject containing the target object captured by the camera device is low, and the user can adjust the shooting angle, shooting position, etc. to obtain a higher-quality subject.

[0094] In some possible implementations, determining, based on the confidence level of the subject detection frame in each first target detection frame image, a display style of the subject display frame in the display frame image associated with the first target detection frame image may include:

[0095] Determining a confidence mean based on the confidence of the subject detection frame in each first target detection frame image;

[0096] Based on the preset correspondence between the confidence mean and the display style and based on the confidence mean corresponding to the first target detection frame image, the display style of the subject display frame in the display frame image associated with the first target detection frame image is determined.

[0097] The confidence mean can be obtained by dividing the sum of the confidences of the subject detection frames in the first target detection frame images by the number of the first target detection frame images.

[0098] The preset correspondence between the mean confidence value and the display style can be such that the higher the mean confidence value, the closer the color of the main display frame is to a safe color, and the lower the mean confidence value, the closer the color of the main display frame is to a dangerous color. In some specific embodiments, when the mean confidence value is greater than or equal to 0.85, the color of the main display frame is green, and when the mean confidence value is less than 0.85, the color of the main display frame is orange.

[0099] Similarly, taking the example of extracting detection frames from a frame image sequence every 100ms, with a display frame extraction frequency of 30 frames per second and a first preset number of 2, the frame image sequence has frames 1-100, with the 1st, 10th, and 20th frames being extracted as detection frames. If the subject detection frame is not detected in the 1st frame, but the 10th and 20th frames are both detected as subject detection frames, and the confidence level of the subject detection frame for the 10th frame is 0.8 and the confidence level of the subject detection frame for the 20th frame is 0.85, then the 20th frame can be used as the first target detection frame. The display style of the subject display frame in the display frames associated with the 20th frame (i.e., the 21st, 24th, and 27th frames) is determined by the average confidence level of the subject detection frames detected in the 10th and 20th frames, i.e., 0.825, and is orange.

[0100] When the subject detection frame is also detected in the 30th frame image and the confidence of the subject detection frame is 0.9, the display style of the subject display frame in the display frame image associated with the 30th frame image (i.e., the 30th, 33rd, 36th, and 39th frame images) is determined by the average confidence value of the subject detection frame detected in the 10th, 20th, and 30th frame images, i.e., 0.85, and is green.

[0101] The display style of the main display frame is determined by the average confidence value corresponding to the first target detection frame image, which can avoid the image jitter caused by excessive changes in the style of the main display frames of adjacent display frame images due to excessive differences in the confidence values ​​of the main detection frames of adjacent detection frames.

[0102] In some related technologies, after an electronic image of a target object (such as a paper bill) is generated by a camera device, the generated electronic image is sent to a pre-trained subject detection model for detecting the bill subject in the electronic image. If the subject detection model detects the bill subject in the electronic image and the confidence level is greater than a certain threshold, the collection of the paper bill can be considered complete, and the bill subject detected in the electronic image is used as the electronic data corresponding to the paper bill for further processing.

[0103] However, from the perspective of actual application, the camera device is too sensitive in capturing images. When the camera device is in motion, it may generate an electronic image. As a result, although the subject detection model can detect the bill subject in the electronic image, the quality of the bill subject is not high (for example, the content of the bill subject is blurred, etc.), and it cannot be used for further operations on the bill subject.

[0104] Figure 5 A flowchart of some steps of another method for displaying a subject detection frame provided by an embodiment of the present disclosure is shown as follows: Figure 5 As shown in , in an optional manner of the present disclosure, the method for displaying the subject detection frame may further include:

[0105] Step S510: In response to the presence of not less than a second preset number of subject detection frames continuously detected in the detection frame image, determine the stability of the subject detection frame in each second target detection frame image based on the position change of the subject detection frame in each second target detection frame image.

[0106] When a subject detection frame is detected in continuous detection frame images and the number of these connected detection frame images is not less than a second preset number (such as 20), these detection frame images are second target detection frames, and the stability of the subject detection frame in each second target detection frame image is determined based on the position change of the subject detection frame in the second target detection frame image.

[0107] Similarly, for example, a detection frame image is extracted from a frame image sequence every 100ms, with a display frame extraction frequency of 30 frames per second and a first preset number of 20. The frame image sequence has frames 1-200, and the 1st, 10th, 20th, 30th, and so on, are extracted as detection frames. When these frames are all detected as subject detection frames, these frames can be used as second target detection frames. The stability of the subject detection frames in these detection frames is determined based on the changes in the positions of the subject detection frames.

[0108] In some specific implementations, the eight coordinates of the four vertices of the subject detection frame of the second target detection frame can be extracted to form eight lists. Each list has an x-coordinate or y-coordinate of a vertex, and eight values ​​are obtained by calculating the square difference of each list. The average value Q of the eight values ​​is calculated. If Q is less than a preset stability threshold, the subject detection frame is considered to be stable; otherwise, it is considered that the subject detection frame is not stable.

[0109] Of course, it is also possible to calculate the variance of each list or other values ​​that can represent the stability of the list. This disclosure does not limit the specific calculation method.

[0110] If the stability of the subject detection frame in each second target detection frame image is good, it means that the position change of the subject detection frame in each second target detection frame image is small, which further indicates that the change of the frame image captured by the camera device is small. Correspondingly, the camera device is more likely to be in a stable state. At this time, the quality of the subject corresponding to the target object obtained by the camera device is more likely to be high.

[0111] That is, determining the stability of the subject detection frame in each second target detection frame image can help obtain a subject corresponding to the target object with higher quality.

[0112] In some possible implementations, after determining the stability of the subject detection frame in each second object detection frame image based on the position change of the subject detection frame in each second object detection frame image, the method further includes:

[0113] In response to the stability of the subject detection frames in the second target detection frame images satisfying a preset stability threshold, a target subject detection frame is determined based on the subject detection frames in the second target detection frame images.

[0114] In some possible implementations, the image pixels corresponding to the target subject detection frame can be used as image data corresponding to the target object, and the server can further process it.

[0115] Similarly, taking the example of extracting detection frames from a frame image sequence every 100ms, with a display frame extraction frequency of 30 frames per second and a first preset number of 20, the frame image sequence has frames 1-200, and the 1st, 10th, 20th, 30th, and so on, are extracted as detection frames. When these frames are all detected in the subject detection frame, they can be used as the second target detection frame. The eight coordinates of the four vertices of the subject detection frame of these frame images can be extracted to form eight lists. Each list has the x-coordinate or y-coordinate of a vertex. By calculating the square difference of each list, eight values ​​are obtained, and the average value Q of these eight values ​​is calculated.

[0116] When Q is less than the preset stability threshold, it is considered that the position change of the subject detection frame of these frame images meets the preset stability threshold. In other words, the position change of the subject detection frame in each frame image is small, and the frame images collected by the camera device during this period of time change is small. The camera device is more likely to be in a stable state. At this time, the quality of the subject corresponding to the target object obtained by the camera device is higher.

[0117] In some possible implementations, determining the target subject detection frame based on the subject detection frame in the second target detection frame image includes:

[0118] The subject detection frame in the last second target detection frame image is determined as the target subject detection frame.

[0119] When the stability of the subject detection frame in each second target detection frame image meets a preset stability threshold, the subject detection frame in the last frame of the second target detection frame image is determined as the target subject detection frame.

[0120] When the stability of the subject detection frame in the second target detection frame image meets the preset stability threshold, the position change of the subject detection frame in each second target detection frame image is small, the change of the frame image captured by the camera device is small, and the camera device is more likely to be in a stable state.

[0121] The longer the stable state lasts, the higher the quality of the collected target object's main body is. Determining the main body detection frame in the last frame of the second target detection frame image as the target main body detection frame can ensure the quality of the target main body detection frame.

[0122] In some possible implementations, determining the target subject detection frame based on the subject detection frame in the second target detection frame image includes:

[0123] Determine a candidate subject detection frame based on the subject detection frame in the second target detection frame image;

[0124] In response to the confidence of the subject detection frame in the second target detection frame image satisfying a preset confidence condition, the candidate subject detection frame is determined as the target subject detection frame.

[0125] Taking the determination of the subject detection frame in the last frame of the second target detection frame image as a candidate subject detection frame as an example, that is, only when the confidence of the subject detection frame in the last frame of the second target detection frame image is greater than the preset threshold, the subject detection frame in the last frame of the second target detection frame image is determined as the target subject detection frame.

[0126] Since the stability is determined based on the position change of the subject detection frame of multiple consecutive frames of second target detection frame images, it is possible that the camera device shook when acquiring the last frame of the second target detection frame image, resulting in the quality of the subject detection frame of the last frame of the second target detection frame image being not high. However, since the other consecutive frames of second target detection frame images except the last frame of the second target detection frame image are acquired when the camera device is in a stable state, the stability calculated based on these consecutive frames of second target detection frame images meets the stability threshold.

[0127] At this time, using the last frame of the second target detection frame image as the target subject detection frame will result in low quality of the target subject detection frame.

[0128] By comparing the confidence of the candidate subject detection frame with the preset confidence condition, the confidence of the candidate subject detection frame can be guaranteed to be high, thereby ensuring the quality of the target subject detection frame and avoiding the occurrence of the above situation.

[0129] Figure 6 This is a flowchart of a specific embodiment of a method for displaying a subject detection frame provided by an embodiment of the present disclosure, such as Figure 6 As shown, a method for displaying a subject detection frame provided by an embodiment of the present disclosure can be used in a scenario where the target object is a bill and a mobile phone is used as a camera device to capture the target object.

[0130] When the mobile phone camera captures the target object and generates a video stream, a frame of image is extracted every 100ms as a detection frame image for subject detection. In order to facilitate the judgment of the stability of the detection frame image, a result queue is constructed.

[0131] Before starting to extract frame images from the video stream as detection frame images, first set the initial state, set the second preset number totalTime to 20, set the first weight weight to 0.5, set the result queue to empty and the length to 20, set the confidence of the subject detection frame to be greater than 0.7, and only then is it considered that the subject detection frame exists in the detection frame image, and set the flag representing whether the image acquisition is completed to False.

[0132] Start extracting frame images from the video stream as detection frame images, and perform subject detection on the detection frame images. When a subject detection frame is detected in the detection frame image and the confidence of the subject detection frame is greater than 0.7, the position of the subject detection frame of the detection frame image (that is, the pixel coordinates of the four vertices of the subject detection frame) and the confidence information of the subject detection frame are stored in the result queue, and the value of totalTime is reduced by 1.

[0133] At the same time, the subject detection results of the detection frame need to be displayed on the display interface of the mobile phone to enhance the user's interactive experience.

[0134] Specifically, the color of the subject display frame displayed on the display interface is determined according to the mean confidence value of all subject detection frames in the current result queue, that is, if the mean confidence value is greater than or equal to 0.85, the color of the subject display frame is green; if the mean confidence value is less than 0.85, the color of the subject display frame is orange.

[0135] The position of the subject display frame displayed on the next frame is calculated based on the position of the last subject detection frame in the current result queue and the position of the subject display frame displayed on the current display interface. Of course, if the result queue is empty, the subject display frame will not be displayed on the next frame display interface.

[0136] Specifically, the position of the main display frame displayed on the next frame display interface can be calculated according to the formula.

[0137] current.x=front_point.x*weight+detect_point.x*(1-weight);

[0138] current.y=front_point.y*weight+detect_point.y*(1-weight);

[0139] Among them, front_point.x is the x-coordinate of the corresponding vertex of the main display frame displayed on the current display interface, and front_point.y is the y-coordinate of the corresponding vertex of the main display frame displayed on the current display interface; detect_point.x is the x-coordinate of the corresponding vertex of the last main detection frame in the current result queue, and front_point.y is the y-coordinate of the corresponding vertex of the last main detection frame in the current result queue; weight is the first weight, which is a value greater than 0 and less than 1, and 1-weight is the second weight; current_point.x is the x-coordinate of the corresponding vertex of the main display frame displayed on the next frame display interface, and current_point.y is the y-coordinate of the corresponding vertex of the main display frame displayed on the next frame display interface.

[0140] Continue to extract the detection frame image. If the subject detection frame is not detected in the detection frame image, or the confidence of the subject detection frame is less than 0.7, clear the result queue and reset totalTime to 20. If the subject detection frame is detected in the detection frame image and the confidence of the subject detection frame is greater than 0.7, continue to execute the step of storing the position of the subject detection frame of the detection frame image and the confidence information of the subject detection frame in the result queue, and subtract 1 from the value of totalTime.

[0141] When the value of totalTime is zero, or the length of the result queue is full, the stability state is calculated based on the position of the subject detection frame in the result queue. Specifically, the 8 coordinates of the four vertices of these subject detection frames are extracted to form 8 lists. Each list has the x-coordinate or y-coordinate of a vertex. By calculating the square difference of each list, 8 values ​​are obtained. The average value Q of the 8 values ​​is calculated. When Q is less than the preset stability threshold, the camera device is considered to have reached a stable state.

[0142] At the same time, the confidence score of the last subject detection frame in the result queue is obtained and compared with the preset confidence condition (i.e., whether it is greater than 0.9). If it meets the confidence condition, the subject detection frame is determined as the target subject detection frame, and the phone's camera resources are released, ending the camera acquisition process. At the same time, the flag value is set to True, allowing external programs to determine that the camera acquisition process has ended.

[0143] If the confidence of the last subject detection frame in the result queue does not meet the confidence condition, the first subject detection frame position and the confidence of the subject detection frame in the result queue are popped out of the result queue, and the detection frame image is continued to be extracted. If the subject detection frame is not detected in the detection frame image, or the confidence of the subject detection frame is less than 0.7, the result queue is cleared and totalTime is reset to 20; if the subject detection frame is detected in the detection frame image, and the confidence of the subject detection frame is greater than 0.7, the position of the subject detection frame of the detection frame image and the confidence information of the subject detection frame are stored in the result queue, and the value of totalTime is reduced by 1.

[0144] If Q is greater than the preset stability threshold, the camera device is considered to be in an unstable state, and the first subject detection frame position and the confidence of the subject detection frame in the result queue are popped out of the result queue, and the detection frame image is continued to be extracted. If the subject detection frame is not detected in the detection frame image, or the confidence of the subject detection frame is less than 0.7, the result queue is cleared and totalTime is reset to 20; if the subject detection frame is detected in the detection frame image, and the confidence of the subject detection frame is greater than 0.7, the position of the subject detection frame of the detection frame image and the confidence information of the subject detection frame are stored in the result queue, and the value of totalTime is reduced by 1.

[0145] Based on Figure 1 The same principle as shown in the method, Figure 7 A schematic structural diagram of a display device for a subject detection frame provided by an embodiment of the present disclosure is shown. Figure 7 As shown, the display device 70 of the subject detection frame may include:

[0146] An image sequence acquisition module 710 is used to acquire a frame image sequence in a video stream captured of a target object;

[0147] A frame image extraction module 720 is configured to extract display frame images from a frame image sequence based on a preset first extraction frequency, and to extract detection frame images from the frame image sequence based on a preset second extraction rate, wherein the first extraction frequency is greater than the second extraction frequency, and a display frame image located between two adjacent detection frame images in the frame image sequence is associated with a previous detection frame image of the two adjacent detection frame images, wherein the detection frame images are used to perform subject detection on a subject corresponding to a target object, and the display frame images are used for display, and in response to a subject detection frame corresponding to a target object being detected in a first detection frame image among the detection frame images, a subject display frame determined based on the subject detection frame in the first detection frame image is displayed in a first display frame image associated with the first detection frame image;

[0148] The frame image calculation module 730 is used to determine, in response to determining a second detection frame image from the detection frame image, a subject display frame in the target second display frame image for any target second display frame image in the second display frame image based on the subject display frame in the previous display frame image of the target second display frame image and the subject detection frame in the second detection frame image, wherein subject detection frames are detected in both the second detection frame image and the previous detection frame image of the second detection frame image, the second display frame image is a detection frame image associated with the second detection frame image, and a first offset between the subject display frame in the previous display frame image of the target second display frame image and the subject display frame in the target second display frame image is less than a second offset between the subject display frame in the previous display frame image of the target second display frame image and the subject detection frame in the second detection frame image.

[0149] Compared with the prior art, the display device of the subject detection frame extracts the detection frame image at a sampling frequency lower than the sampling frequency of the display frame image, and the subject detection frame is detected in both the second detection frame image and the detection frame image before the second detection frame image. The subject display frame in the target second display frame image associated with the second detection frame image is determined based on the subject display frame in the display frame image before the second detection frame image and the subject detection frame in the second detection frame image. When there is a large difference in the position of the subject detection frames in two adjacent detection frame images, the position of the subject display frame between the front and rear display frame images does not follow the subject detection frame in the detection frame image and shift rapidly, but the shift is "dispersed" to multiple display frame images, thereby reducing the shift of the subject display frames of the front and rear display frame images, reducing the jitter of the picture, and improving the user experience.

[0150] It can be understood that the above modules of the display device of the subject detection frame in the embodiment of the present disclosure have the function of realizing Figure 1 The functions of the corresponding steps of the method for displaying the subject detection frame in the embodiment shown in . This function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The above modules can be software and / or hardware, and the above modules can be implemented separately or integrated with multiple modules. For the functional description of each module of the display device of the subject detection frame, please refer to Figure 1 The corresponding description of the method for displaying the subject detection frame in the embodiment shown in is not repeated here.

[0151] In some possible implementations, the frame image calculation module 730 may also be used to:

[0152] Based on the vertex coordinates of the main display frame in the previous frame display frame image of the target second display frame image and the first weight, and based on the vertex coordinates of the main detection frame in the second detection frame image and the second weight, the vertex coordinates of the main display frame in the target second display frame image are determined, the sum of the first weight and the second weight is 1, and the first weight is a value greater than 0 and less than 1.

[0153] In some possible implementations, the subject detection frame display device 70 further includes:

[0154] A frame image processing module is used to determine a subject display frame in a third display frame image associated with the third detection frame image in response to determining a third detection frame image from the detection frame image, based on the subject detection frame in the third detection frame image, wherein the subject detection frame is detected in the third detection frame image, and the subject detection frame is not detected in the previous detection frame image of the third detection frame image.

[0155] In some possible implementations, the subject detection frame display device 70 further includes:

[0156] The frame image style module is used to determine the display style of the subject display frame in the display frame image associated with the first target detection frame image based on the confidence of the subject detection frame in each first target detection frame image in response to the continuous detection of subject detection frames in the detection frame image.

[0157] In some possible implementations, determining a display style of a subject display frame in a display frame image associated with the first target detection frame image based on the confidence level of the subject detection frame in each first target detection frame image includes:

[0158] Determining a confidence mean based on the confidence of the subject detection frame in each first target detection frame image;

[0159] Based on the preset correspondence between the confidence mean and the display style and based on the confidence mean corresponding to the first target detection frame image, the display style of the subject display frame in the display frame image associated with the first target detection frame image is determined.

[0160] In some possible implementations, the subject detection frame display device 70 further includes:

[0161] A stability calculation module is used to determine the stability of the subject detection frame in each second target detection frame image based on the position change of the subject detection frame in each second target detection frame image in response to the continuous detection of the subject detection frame in the detection frame image.

[0162] In some possible implementations, the subject detection frame display device 70 further includes:

[0163] The stability judgment module is used to determine the target subject detection frame based on the subject detection frame in the second target detection frame image in response to the stability of the subject detection frame in each second target detection frame image meeting a preset stability threshold.

[0164] In some possible implementations, determining the target subject detection frame based on the subject detection frame in the second target detection frame image includes:

[0165] The subject detection frame in the last second target detection frame image is determined as the target subject detection frame.

[0166] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0167] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0168] The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for displaying a subject detection frame as provided in an embodiment of the present disclosure.

[0169] Compared with the existing technology, this electronic device extracts the detection frame image at a sampling frequency lower than the sampling frequency of the display frame image, and the subject detection frame is detected in both the second detection frame image and the detection frame image preceding the second detection frame image. Based on the subject display frame in the display frame image preceding the target second display frame image associated with the second detection frame image and the subject detection frame in the second detection frame image, the subject display frame in the target second display frame image is determined. When there is a large difference in the position of the subject detection frames in two adjacent detection frame images, the position of the subject display frame between the front and rear display frame images does not follow the subject detection frame in the detection frame image and shift rapidly, but the shift is "dispersed" to multiple display frame images, thereby reducing the shift of the subject display frames of the front and rear display frame images, reducing the jitter of the picture, and improving the user experience.

[0170] The readable storage medium is a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method for displaying a subject detection frame as provided in an embodiment of the present disclosure.

[0171] Compared with the prior art, the readable storage medium extracts the detection frame image at a sampling frequency lower than the sampling frequency of the display frame image, and the subject detection frame is detected in both the second detection frame image and the detection frame image preceding the second detection frame image. The subject display frame in the target second display frame image associated with the second detection frame image is determined based on the subject display frame in the display frame image preceding the target second display frame image and the subject detection frame in the second detection frame image. When there is a large difference in the position of the subject detection frames in two adjacent detection frame images, the position of the subject display frame between the front and rear display frame images does not follow the subject detection frame in the detection frame image and shift rapidly, but "disperses" the shift to multiple display frame images, thereby reducing the shift of the subject display frames of the front and rear display frame images, reducing the jitter of the picture, and improving the user experience.

[0172] The computer program product includes a computer program. When the computer program is executed by a processor, the computer program implements the method for displaying a subject detection frame provided in the embodiment of the present disclosure.

[0173] Compared with the existing technology, this computer program product extracts detection frame images at a sampling frequency lower than the sampling frequency of display frame images, and a subject detection frame is detected in both the second detection frame image and the detection frame image preceding the second detection frame image. Based on the subject display frame in the display frame image preceding the target second display frame image associated with the second detection frame image and the subject detection frame in the second detection frame image, the subject display frame in the target second display frame image is determined. When there is a large difference in the position of the subject detection frames in two adjacent detection frame images, the position of the subject display frame between the front and rear display frame images does not follow the subject detection frame in the detection frame image and shift rapidly, but instead "disperses" the shift to multiple display frame images, thereby reducing the shift of the subject display frames of the front and rear display frame images, reducing the jitter of the picture, and improving the user experience.

[0174] Figure 8 A schematic block diagram of an example electronic device 80 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0175] like Figure 8As shown, the electronic device 80 includes a computing unit 810, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 820 or a computer program loaded from a storage unit 880 into a random access memory (RAM) 830. Various programs and data required for the operation of the device 80 can also be stored in the RAM 830. The computing unit 810, the ROM 820, and the RAM 830 are connected to each other via a bus 840. An input / output (I / O) interface 850 is also connected to the bus 840.

[0176] Multiple components in device 80 are connected to I / O interface 850, including an input unit 860, such as a keyboard, mouse, etc.; an output unit 870, such as various types of displays, speakers, etc.; a storage unit 880, such as a magnetic disk, optical disk, etc.; and a communication unit 890, such as a network card, modem, wireless communication transceiver, etc. The communication unit 890 allows device 80 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0177] The computing unit 810 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 810 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 810 executes the method for displaying a subject detection frame provided in the embodiments of the present disclosure. For example, in some embodiments, the method for displaying a subject detection frame provided in the embodiments of the present disclosure can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 880. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 80 via the ROM 820 and / or the communication unit 890. When the computer program is loaded into the RAM 830 and executed by the computing unit 810, one or more steps of the method for displaying a subject detection frame provided in the embodiments of the present disclosure can be performed. Alternatively, in other embodiments, the computing unit 810 may be configured in any other appropriate manner (for example, by means of firmware) to execute the method for displaying the subject detection frame provided in the embodiments of the present disclosure.

[0178] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0179] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0180] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0181] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0182] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0183] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0184] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0185] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for displaying a subject detection frame, comprising: Acquire a frame image sequence in a video stream captured of a target object; Extracting a display frame image from the frame image sequence based on a preset first extraction frequency, and extracting a detection frame image from the frame image sequence based on a preset second extraction frequency, wherein the first extraction frequency is greater than the second extraction frequency, a display frame image located between two adjacent detection frame images in the frame image sequence is associated with a previous detection frame image of the two adjacent detection frame images, the detection frame images are used to perform subject detection on a subject corresponding to the target object, and the display frame images are used to perform display, and in response to a subject detection frame corresponding to the target object being detected in a first detection frame image among the detection frame images, a subject display frame determined based on the subject detection frame in the first detection frame image is displayed in a first display frame image associated with the first detection frame image; In response to determining a second detection frame image from the detection frame images, for any target second display frame image in the second display frame images, determining a subject display frame in the target second display frame image based on a subject display frame in a display frame image preceding the target second display frame image and a subject detection frame in the second detection frame image, wherein subject detection frames are detected in both the second detection frame image and a detection frame image preceding the second detection frame image, the second display frame image is a detection frame image associated with the second detection frame image, and a first offset between the subject display frame in the display frame image preceding the target second display frame image and the subject display frame in the target second display frame image is less than a second offset between the subject display frame in the display frame image preceding the target second display frame image and the subject detection frame in the second detection frame image; The determining of the subject display frame in the target second display frame image based on the subject display frame in the display frame image preceding the target second display frame image and the subject detection frame in the second detection frame image includes: Based on the vertex coordinates of the main body display frame in the previous frame display frame image of the target second display frame image and the first weight, and based on the vertex coordinates of the main body detection frame in the second detection frame image and the second weight, the vertex coordinates of the main body display frame in the target second display frame image are determined, the sum of the first weight and the second weight is 1, and the first weight is a value greater than 0 and less than 1.

2. The method according to claim 1, further comprising: In response to determining a third detection frame image from the detection frame image, based on the subject detection frame in the third detection frame image, a subject display frame in a third display frame image associated with the third detection frame image is determined, wherein the subject detection frame is detected in the third detection frame image, and the subject detection frame is not detected in a previous detection frame image of the third detection frame image.

3. The method according to claim 1 or 2, further comprising: In response to the presence of no less than a first preset number of subject detection frames continuously detected in the detection frame image, a display style of the subject display frame in the display frame image associated with the first target detection frame image is determined based on the confidence of the subject detection frame in each of the first target detection frame images.

4. The method according to claim 3, wherein: The determining, based on the confidence level of the subject detection frame in each of the first target detection frame images, a display style of the subject display frame in the display frame image associated with the first target detection frame image includes: Determining a confidence mean based on the confidence of the subject detection frame in each of the first target detection frame images; Based on the preset correspondence between the confidence mean and the display style and based on the confidence mean corresponding to the first target detection frame image, the display style of the main body display frame in the display frame image associated with the first target detection frame image is determined.

5. The method according to claim 1 or 2, further comprising: In response to the presence of no less than a second preset number of second target detection frame images in the detection frame image, the stability of the subject detection frame in each second target detection frame image is determined based on the position change of the subject detection frame in each second target detection frame image.

6. The method according to claim 5, after determining the stability of the subject detection frame in each second object detection frame image based on the position change of the subject detection frame in each second object detection frame image, the method further comprises: In response to the stability of the subject detection frames in the second target detection frame images satisfying a preset stability threshold, a target subject detection frame is determined based on the subject detection frames in the second target detection frame images.

7. The method according to claim 6, wherein: The determining of the target subject detection frame based on the subject detection frame in the second target detection frame image includes: The subject detection frame in the last frame of the second target detection frame image is determined as the target subject detection frame.

8. The method according to claim 6 or 7, wherein: The determining of the target subject detection frame based on the subject detection frame in the second target detection frame image includes: Determine a candidate subject detection frame based on the subject detection frame in the second target detection frame image; In response to the confidence of the subject detection frame in the second target detection frame image satisfying a preset confidence condition, the candidate subject detection frame is determined as the target subject detection frame.

9. A display device for a subject detection frame, comprising: An image sequence acquisition module is used to acquire a frame image sequence in a video stream captured of a target object; a frame image extraction module, configured to extract display frame images from the frame image sequence based on a preset first extraction frequency, and extract detection frame images from the frame image sequence based on a preset second extraction frequency, wherein the first extraction frequency is greater than the second extraction frequency, a display frame image located between two adjacent detection frame images in the frame image sequence is associated with a preceding detection frame image of the two adjacent detection frame images, the detection frame images being used to perform subject detection on a subject corresponding to the target object, and the display frame images being used for display, wherein, in response to a subject detection frame corresponding to the target object being detected in a first detection frame image among the detection frame images, a subject display frame determined based on the subject detection frame in the first detection frame image is displayed in a first display frame image associated with the first detection frame image; a frame image calculation module for, in response to determining a second detection frame image from the detection frame image, determining, for any target second display frame image in the second display frame images, a subject display frame in the target second display frame image based on a subject display frame in a display frame image preceding the target second display frame image and a subject detection frame in the second detection frame image, wherein subject detection frames are detected in both the second detection frame image and the detection frame image preceding the second detection frame image, the second display frame image is a detection frame image associated with the second detection frame image, and a first offset between the subject display frame in the display frame image preceding the target second display frame image and the subject display frame in the target second display frame image is less than a second offset between the subject display frame in the display frame image preceding the target second display frame image and the subject detection frame in the second detection frame image; The determining of the subject display frame in the target second display frame image based on the subject display frame in the display frame image preceding the target second display frame image and the subject detection frame in the second detection frame image includes: Based on the vertex coordinates of the main body display frame in the previous frame display frame image of the target second display frame image and the first weight, and based on the vertex coordinates of the main body detection frame in the second detection frame image and the second weight, the vertex coordinates of the main body display frame in the target second display frame image are determined, the sum of the first weight and the second weight is 1, and the first weight is a value greater than 0 and less than 1.

10. The device according to claim 9, wherein the display device of the subject detection frame further comprises: A frame image processing module is used to determine, in response to determining a third detection frame image from the detection frame image, a subject display frame in a third display frame image associated with the third detection frame image based on the subject detection frame in the third detection frame image, wherein the subject detection frame is detected in the third detection frame image and the subject detection frame is not detected in a previous detection frame image of the third detection frame image.

11. The device according to claim 9 or 10, wherein the display device of the subject detection frame further comprises: A frame image style module is used to determine, in response to the continuous detection of subject detection frames in the detection frame image, a first target detection frame image having no less than a first preset number of first target detection frame images, and based on the confidence level of the subject detection frame in each of the first target detection frame images, a display style of the subject display frame in the display frame image associated with the first target detection frame image.

12. The device according to claim 11, wherein The determining, based on the confidence level of the subject detection frame in each of the first target detection frame images, a display style of the subject display frame in the display frame image associated with the first target detection frame image includes: Determining a confidence mean based on the confidence of the subject detection frame in each of the first target detection frame images; Based on the preset correspondence between the confidence mean and the display style and based on the confidence mean corresponding to the first target detection frame image, the display style of the main body display frame in the display frame image associated with the first target detection frame image is determined.

13. The device according to claim 9 or 10, wherein the display device of the subject detection frame further comprises: A stability calculation module is used to determine the stability of the subject detection frame in each second target detection frame image based on the position change of the subject detection frame in each second target detection frame image in response to the continuous detection of the subject detection frame in the second target detection frame image of not less than a second preset number.

14. The device according to claim 13, wherein the display device of the subject detection frame further comprises: The stability judgment module is used to determine the target subject detection frame based on the subject detection frame in each second target detection frame image in response to the stability of the subject detection frame in each second target detection frame image meeting a preset stability threshold.

15. The device according to claim 14, wherein The determining of the target subject detection frame based on the subject detection frame in the second target detection frame image includes: The subject detection frame in the last frame of the second target detection frame image is determined as the target subject detection frame.

16. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

17. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

18. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Target detection box output method and device, terminal and storage medium

    CN110677585A

  • Video data processing method and device, electronic equipment and computer readable medium

    CN111179310A