Target object marking method, electronic equipment, storage medium and program product
By using an anchor point frame as a reference when the electronic device and the target object remain relatively stationary, combined with a preset change threshold and a static mode judgment, the problem of coordinate frame jitter is solved, achieving stable target object marking, improving user experience and resource utilization.
Patent Information
- Application Number
- CN202410745746.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-12-16
AI Technical Summary
When the electronic device and the target object remain relatively stationary or nearly stationary, the position prediction model causes the coordinate frame to jitter in the captured image, affecting the user experience.
By using the anchor point box as a reference when the electronic device and the target object remain relatively stationary, it determines whether to display the anchor point box or the first output box, avoiding unnecessary box jitter. Combined with preset change thresholds and static mode judgment, it improves recognition accuracy.
It effectively avoids coordinate frame jitter, improves resource utilization, reduces unnecessary costs, and enhances the accuracy of positioning markers and user experience.
Smart Images

Figure CN121151673A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of terminal, and in particular, to a target object marking method, an electronic device, a storage medium, and a program product. BACKGROUND
[0002] With the rapid development of science and technology, the shooting function of electronic devices is becoming more and more powerful, that is, electronic devices can not only be used for simple shooting, but also can frame a target object in a shooting picture through a coordinate frame during shooting. For example, during shooting by an electronic device, a picture being shot can be displayed on a screen, and a marking frame can be displayed in the shooting picture, so that the electronic device can focus on a target object in the marking frame.
[0003] In the conventional technology, a position prediction model is used to positionally predict a target object in a shooting picture to output a coordinate frame, and the electronic device can display the predicted coordinate frame in the shooting picture to mark the target object in the shooting picture. In the case that the target object and the electronic device are relatively static or close to being relatively static, the electronic device can perform positional detection through the position prediction model, so that the coordinate frame follows the movement, thereby always framing the target object in the shooting picture.
[0004] However, due to the instability of the position prediction model, in the case that the electronic device and the target object are relatively static or close to being relatively static, the coordinate frame displayed on the shooting picture by the position prediction model will still change in position or size, and thus the coordinate frame will easily shake on the picture. SUMMARY
[0005] Embodiments of the present application provide a target object marking method, an electronic device, a storage medium, and a program product, which can avoid shaking of a coordinate frame in the case that an electronic device and a target object are relatively static.
[0006] To achieve the above object, embodiments of the present application adopt the following technical solutions:
[0007] In a first aspect, a target object marking method is provided. The method is applied to an electronic device having a camera. In a process in which the camera continuously captures frame images, the electronic device performs positioning prediction on a target object in each captured frame image to obtain an output frame. For example, the electronic device can perform positioning prediction on the target object in each real-time captured frame image by using a position prediction model, and predict an output frame for the target object in each frame image. However, unlike the conventional method, the scheme in the embodiments of the present application does not directly display the predicted output frame. Instead, when each current frame image is processed, the electronic device determines whether the electronic device and the target object remain relatively static within a target time period, i.e., between the capture time of the historical frame image corresponding to the output frame (i.e., anchor frame) used for reference and the capture time of the current frame image. If the electronic device and the target object remain relatively static, the anchor frame can be directly displayed in the current frame image to mark the target object in the current frame image. In the case where the electronic device and the target object move relatively within the target time period, the electronic device can display a first output frame in the current frame image to mark the target object, and use the first output frame as a new anchor frame.
[0008] In the above scheme, in the case where the electronic device and the target object remain relatively static within the target time period, the first output frame predicted for the target object in the current frame image can not be displayed in the current frame image, and the anchor frame can be displayed in the current frame image to mark the target object. From the perspective of the user, in the case where the electronic device and the target object remain relatively static, a fixed anchor frame is always displayed, thereby avoiding frame jitter. If the electronic device and the target object move relatively within the target time period, the user is not overly sensitive to the change of the marking frame, and thus the first output frame predicted for the current frame image can be directly displayed, and the first output frame can be used as a new anchor frame, so that if the electronic device and the target object remain relatively static when subsequent frame images are captured, unnecessary frame jitter can be avoided.
[0009] In a possible implementation of the first aspect, the electronic device can identify whether the electronic device and the target object remain relatively static within the target time period by detecting the relative change between the first output frame and the anchor frame. Specifically, in the case where the relative change between the first output frame and the anchor frame does not exceed a preset change threshold, it is determined that the electronic device and the target object remain relatively static within the target time period, and thus the target object can be marked in the current frame image based on the anchor frame.
[0010] In the above scheme, no additional auxiliary judgment is required, and the predicted output frame is used to identify whether the electronic device and the target object remain relatively static within the target time period, thereby improving the utilization of resources and reducing unnecessary costs.
[0011] In a possible implementation manner of the first aspect, the preset change threshold comprises a preset position change threshold and / or a preset size change threshold. The electronic device can identify whether the electronic device and the target object remain relatively stationary within the target time period by detecting the position change and / or the size change between the first output frame and the anchor point frame. Specifically, in a case where the position change between the first output frame and the anchor point frame does not exceed the preset position change threshold and / or the size change between the first output frame and the anchor point frame does not exceed the preset size change threshold, the target object is marked based on the anchor point frame in the current frame image.
[0012] In the above scheme, other auxiliary judgments are not required, and whether the electronic device and the target object remain relatively stationary within the target time period is identified by the change of the position and size of the predicted output frame, which improves the utilization rate of resources, is more convenient, and also improves the accuracy of identification to a certain extent.
[0013] In a possible implementation manner of the first aspect, the preset position change threshold comprises a first threshold. In a case where the center point displacement between the first output frame and the anchor point frame does not exceed the first threshold, the electronic device determines that the position change does not exceed the preset position change threshold.
[0014] And / or,
[0015] The preset size change threshold comprises a second threshold and / or a third threshold,
[0016] In a case where the side length change between the first output frame and the anchor point frame does not exceed the second threshold and / or the diagonal length change between the first output frame and the anchor point frame does not exceed the third threshold, the electronic device determines that the size change does not exceed the preset size change threshold.
[0017] In the above scheme, other auxiliary judgments are not required, and whether the electronic device and the target object remain relatively stationary within the target time period is accurately identified by the side length change or the diagonal change between the frames, and it is more convenient.
[0018] In a possible implementation manner of the first aspect, in a case where the center point displacement between the first output frame and the anchor point frame does not exceed the first threshold and / or the side length change between the first output frame and the anchor point frame does not exceed the second threshold and / or the diagonal length change between the first output frame and the anchor point frame does not exceed the third threshold, the electronic device marks the target object based on the anchor point frame in the current frame image.
[0019] In the scheme, whether the electronic device and the target object remain relatively static within the target time period is identified based on multi-dimensional quantification of center point displacement, diagonal length change, side length change, and the like between the boxes, so that the final marking is more accurate.
[0020] In a possible implementation of the first aspect, in a case where the relative change does not exceed the preset change threshold and the electronic device has been in the static mode, the electronic device marks the target object in the current frame image based on the anchor box.
[0021] In the scheme, whether the electronic device and the target object remain relatively static within the target time period is not identified only from the single dimension of the relative change between the boxes, but is identified in combination with the static mode unique to the embodiments of the present application. The static mode refers to a mode in which the anchor box is output and displayed. By combining the two dimensions of whether the electronic device is in the static mode and the relative change between the boxes, the accuracy of identification is improved.
[0022] In a possible implementation of the first aspect, in a case where the relative change does not exceed the preset change threshold, but the electronic device is not in the static mode, and the number of stable frames does not reach the preset frame number threshold, the target object is marked in the current frame image based on the first output box, the anchor box is kept unchanged, and the number of stable frames is incremented by 1. The number of stable frames refers to the number of stable frame images existing between the historical frame image and the current frame image. The relative change between the output box predicted for the target object in the stable frame image and the anchor box does not exceed the preset change threshold. For example, the output box predicted for the target object in the stable frame image can be obtained by a position prediction model.
[0023] In the scheme, when the electronic device is not in the static mode, even if the relative change between the boxes is small, the anchor box is not directly output, and the electronic device enters the static mode. Instead, the number of stable frames is combined to avoid the situation that the electronic device is too easy to enter the static mode, and thus the situation that the electronic device frequently enters or exits the static mode and causes lag is avoided.
[0024] In a possible implementation of the first aspect, in a case where the relative change does not exceed the preset change threshold, but the electronic device is not in the static mode, and the number of stable frames reaches the preset frame number threshold, the electronic device takes the first output box as a new anchor box to mark the target object in the current frame image based on the new anchor box, and enters the static mode.
[0025] In the scheme, when the electronic device is not in the static mode, the two dimensions are checked, that is, if the relative change between the boxes is small, and the number of stable frames reaches the preset frame number threshold, the first output box is taken as a new anchor box to mark the target object in the current frame image, and the electronic device enters the static mode. In this way, the difficulty of entering the static mode is improved, and the lag is reduced.
[0026] In addition, in a case where the relative change does not exceed the preset change threshold, but the electronic device is not in the stationary mode, and the number of stable frames reaches the preset frame number threshold, although the relative change does not exceed the preset change threshold, the relative change is close to the preset change threshold to a great extent, that is, there is a certain deviation between the first output frame and the anchor point frame. If the preset change threshold is set improperly, for example, is set to be too large, it will actually cause the deviation between the first output frame and the anchor point frame to be relatively obvious. In this case, if the anchor point frame is directly output, the positioning of the target object will be inaccurate. Therefore, based on the above scheme of the present application, the first output frame is displayed as a new anchor point frame to mark the target object in the current frame image, and the stationary mode is entered, which can improve the accuracy of frame output.
[0027] In a possible implementation of the first aspect, in a case where the number of state changes of the electronic device within a preset time window exceeds a preset number threshold, the electronic device can increase the preset frame number threshold; wherein the number of state changes is used to represent the number of times of state changes of the electronic device in the process of processing the frame images collected within the preset time window; wherein the state change includes entering the stationary mode and exiting the stationary mode.
[0028] In the above scheme, the number of state changes of the electronic device is detected based on the preset time window, which can accurately identify whether the state is stable in the process of processing the frame images, that is, can accurately identify the case of frequently entering or exiting the stationary mode, so as to increase the preset frame number threshold used for comparison with the number of stable frames, which can improve the difficulty of entering the stationary mode, thereby avoiding the frequent entry and exit of the stationary mode and reducing the stuttering caused by the frequent entry and exit of the stationary mode.
[0029] In a possible implementation of the first aspect, in a case where the relative change between the first output frame and the anchor point frame exceeds the preset change threshold, the electronic device can mark the target object in the current frame image based on the first output frame, and take the first output frame as a new anchor point frame.
[0030] In the above scheme, other auxiliary judgments are not required, but the predicted output frame is used to identify whether the electronic device and the target object remain relatively stationary in the target time period, which improves the utilization rate of resources and reduces unnecessary costs.
[0031] In a possible implementation of the first aspect, in a case where the electronic device has been in the stationary mode, but the relative change exceeds the preset change threshold, the electronic device exits the stationary mode, marks the target object in the current frame image based on the first output frame, and takes the first output frame as a new anchor point frame.
[0032] In the above scheme, if the electronic device is in the static mode, and the relative change between the first output frame and the anchor frame is too large, the electronic device can exit the static mode, and mark the target object based on the first output frame in the current frame image, so as to ensure the accuracy of the positioning mark. In addition, the first output frame is taken as a new anchor frame, so that if the electronic device and the target object remain relatively static when subsequent image frames are collected, unnecessary frame jitter can be avoided.
[0033] In a possible implementation of the first aspect, if there is no anchor frame, the target object is marked based on the first output frame in the current frame image, and the first output frame is taken as an anchor frame.
[0034] In the above scheme, if there is no anchor frame, the first output frame can be taken as an anchor frame, so that if the electronic device and the target object remain relatively static when subsequent image frames are collected, unnecessary frame jitter can be avoided.
[0035] In the second aspect, the present application provides another target object marking method, applied to an electronic device with a camera. In the process of continuously collecting frame images by the camera, if the electronic device and the target object remain relatively static, a first coordinate frame for marking the target object, for example, an anchor frame, is fixedly displayed, and if the electronic device and the target object move relatively, a second coordinate frame predicted for the target object in the collected current frame image is displayed in real time.
[0036] Therefore, in the case that the electronic device and the target object remain relatively static, a fixed frame can be displayed, and in the case that the electronic device and the target object move relatively, each predicted coordinate frame can be displayed in real time.
[0037] In the third aspect, the present application provides an electronic device, which includes a memory, a camera, and a processor. The memory and the camera are coupled with the processor. The memory stores computer program code, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device executes the method in the first aspect or the second aspect or any possible implementation of the first aspect or the second aspect.
[0038] In the fourth aspect, the present application provides a computer storage medium, which includes computer instructions. When the computer instructions are run on an electronic device, the electronic device executes the method in the first aspect or the second aspect or any possible implementation of the first aspect or the second aspect.
[0039] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a computer, causes the computer to perform the method according to the first aspect or the second aspect or any possible implementation manner thereof. The computer can be the electronic device according to the third aspect or any possible implementation manner thereof.
[0040] It can be understood that the electronic device according to the third aspect, the computer storage medium according to the fourth aspect, and the computer program product according to the fifth aspect can have beneficial effects similar to those of the first aspect or the second aspect or any possible implementation manner thereof, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0041] Figure 1 A schematic diagram of a photograph preview interface of an electronic device according to an embodiment of the present application;
[0042] Figure 2 A schematic diagram of another photograph preview interface of an electronic device according to an embodiment of the present application;
[0043] Figure 3 A schematic diagram of a video shooting interface of an electronic device according to an embodiment of the present application;
[0044] Figure 4 A schematic diagram of a principle according to an embodiment of the present application;
[0045] Figure 5 Another schematic diagram of a principle according to an embodiment of the present application;
[0046] Figure 6 A schematic diagram of a structure of an electronic device according to an embodiment of the present application;
[0047] Figure 7 A schematic diagram of a system structure of an electronic device according to an embodiment of the present application;
[0048] Figure 8 A schematic diagram of a flow of a target object marking method according to an embodiment of the present application;
[0049] Figure 9 A schematic diagram of another flow of a target object marking method according to an embodiment of the present application;
[0050] Figure 10 A schematic diagram of a principle of an electronic device entering a still mode according to an embodiment of the present application;
[0051] Figure 11 A schematic diagram of a principle of an electronic device exiting a still mode according to an embodiment of the present application;
[0052] Figure 12Another target object marking method provided by an embodiment of the present application is shown in the flowchart.
[0053] Figure 13 A time window detection diagram provided by an embodiment of the present application. DETAILED DESCRIPTION
[0054] Hereinafter, the terms "first" and "second" are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features. In the description of the embodiments, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0055] Before formally introducing the method in the embodiments of the present application, the causes of frame jitter are described in detail.
[0056] In the scenario of continuous image acquisition by the electronic device, the position prediction model will respectively perform positioning prediction on each frame of the acquired images, that is, predict the position of the target object in each frame of the images, and output the corresponding coordinate frame. In order to express briefly, the coordinate frame predicted and output by the model can be simply referred to as "output frame", therefore, each frame of the images has a corresponding output frame (i.e., the coordinate frame predicted and output by the model for the target object in each frame of the images). In the traditional method, the electronic device can display the output frame predicted by the model to mark or frame the target object in each frame of the images. In order to facilitate the distinction, the frame that finally marks and displays the target object on the frame of the image is called "marking frame" below, therefore, the marking frame in the traditional method is the output frame of the model.
[0057] In the case that the electronic device and the target object are relatively static or close to being relatively static, that is, the target object and the electronic device do not produce relative displacement, the position, size, etc. of the target object in different frames of the images acquired continuously should not change. Then, in the ideal case, the position and size of the output frame predicted by the position prediction model for the target object should also not change, and further, the position and size of the marking frame finally displayed on the frame of the image should also not change. That is, the effect shown in the following should be presented. Figure 1
[0058] Figure 1 A schematic diagram of a photograph preview interface of an electronic device provided by an embodiment of the present application, Figure 1 is taken as an example to illustrate the case of continuous image acquisition. In the case of taking a photo, after the camera is started, the electronic device displays a preview interface (i.e., an interface for previewing before taking a photo) in the screen, and the real-time shooting picture in the camera field of view range is displayed in the picture preview area in the preview interface. It should be understood that the continuously and real-time acquired shooting picture is a continuously acquired frame image. As shown in Figure 1 , the shooting picture includes a person, a window, a wall clock, and the like. Before starting to take a photo, the electronic device can determine, in response to a touch operation of a user, that the user wants to focus on the person (i.e., a target object) in the real-time acquired shooting picture. Then, the position prediction model predicts the output box 101 of the person, so as to display the output box 101 in the shooting picture to mark the person. If the person and the electronic device remain relatively static or do not produce relative movement, then, in an ideal case, the marking box for the person should always be the output box 101, that is, only one fixed output box 101 is displayed to mark the person. Figure 1
[0059] However, due to the instability of the position prediction model, even in the case that the electronic device and the target object remain relatively static or are close to being relatively static, for example, in the case that the electronic device is fixed on a tripod to take a photo and the target object remains static, the position and size of the output box of the position prediction model for different frame images still change frequently, so that the marking box in the shooting picture is prone to jitter.
[0060] Specifically, Figure 2 Another schematic diagram of a photo preview interface of an electronic device provided by an embodiment of the present application is shown in FIG. 2B. In Figure 2 , in the case that the electronic device and the person 201 remain relatively static, due to the instability of the position prediction model, the size or position of the marking box on each frame image changes all the time in multiple continuous frame images. For example, the marking box 202 is a marking box in the first frame image, the marking box 203 is a marking box in the second frame image, and the marking box 204 is a marking box in the third frame image. It can be seen that the position and size of the marking box 202, the marking box 203, and the marking box 204 change all the time. However, in fact, the picture content or image content of the first frame image, the second frame image, and the third frame image does not change. Therefore, in the view of the user, the shooting picture does not change, but the marking box for the person always changes, so that the user perceives that the marking box jitters in the preview interface.
[0061] This jitter is particularly pronounced in challenging scenarios. These challenging scenarios refer to situations where the location prediction model struggles to locate and identify the target object. For example, if multiple similar objects exist in the frame, and the target object is one of them, these similar objects significantly interfere with the location prediction model's output. This makes it harder for the model to locate and identify the target object, or it reduces the model's certainty about the target object's location and detection. Consequently, the predicted bounding box becomes highly unstable. The position and size of the bounding box predicted by the model for the current frame may deviate significantly from the position and size predicted for previous frames (earlier or historical frames). This results in significant jitter in the bounding box even when the electronic device and the target object are stationary or nearly stationary.
[0062] To facilitate understanding, we will now combine Figure 3 The frame jitter in difficult scenarios is illustrated.
[0063] Figure 3 This is a schematic diagram of a video shooting interface of an electronic device provided in an embodiment of this application. Figure 3 This example illustrates the scenario of continuously acquiring images, using video recording as an example. Figure 3 The image displayed shows the video recording interface after recording begins, showing the railing. Assuming one of the crossroads 301 on the railing is the target object, the electronic device can use a position prediction model to locate and detect this crossroads 301, obtaining an output frame (i.e., a coordinate frame), which is then displayed as a marker frame. Since the railing is fixed to the ground and will not move, and assuming the electronic device is also fixed to a tripod for recording, the railing and electronic device are in a relatively static state. Therefore, ideally, the marker frame for the crossroads 301 should also remain fixed, specifically as follows: Figure 3 As shown in (a) in the figure.
[0064] However, due to Figure 3 The railing contains multiple similar intersection points, which can interfere with the output of the location prediction model. This causes the output bounding box to be unstable; that is, the position and size of the output bounding box predicted by the location prediction model for the current frame may deviate significantly from the position and size of the output bounding box for previous frames, resulting in discrepancies between the bounding boxes in different frames. Specifically, during video recording, even when the electronic device remains stationary or nearly stationary, it will still exhibit similar behavior. Figure 3 As shown in (b) in the figure, Figure 3The dashed box in (b) in FIG. 10 illustrates that the bounding box displayed at different time instants when the electronic device is stationary or nearly stationary is jittery.
[0065] It should be understood that, Figure 3 The jitter of the bounding box is illustrated by taking the interface after starting video shooting as an example. In fact, the jitter of the bounding box can also occur in the video shooting preview interface before starting video shooting. For example, before starting video shooting, the video shooting preview interface can receive a touch operation of the user to determine a target object to be focused, and then output a bounding box predicted by the position prediction model for the target object as a marking box to frame the target object for focus. Due to the instability of the model itself, the jitter of the bounding box can occur in the video shooting preview interface when the target object and the electronic device are relatively stationary.
[0066] It should be noted that the above is only an example of the shooting scene and the video shooting scene to illustrate the scenario of continuously collecting images. In fact, in addition to the above examples, the scenario of continuously collecting images can also include a target tracking or target detection scenario. Specifically, in the target detection scenario, the camera of the electronic device can be called to continuously collect shooting pictures in the detection interface, and the shooting pictures can be displayed in the detection interface in real time and continuously. The electronic device can detect a target object in the shooting pictures by using the position prediction model, and output a coordinate box (i.e., an output box) to locate and mark the target object in the shooting pictures. Similarly, in the target detection scenario, if the detected target object is stationary relative to the electronic device, but the position prediction model itself has instability defects, the marking box of the target object in the detection interface will also be jittery. Of course, in addition to the above examples of the shooting scene, the shooting scene, and the target detection scenario, other scenarios also belong to the scenario of continuously collecting images, which will not be illustrated one by one in the embodiments of the present application.
[0067] In some solutions, in order to solve the problem of frame jitter in the above-mentioned scenario of continuous image acquisition, a filter is used to smooth the image to reduce frame jitter. For example, after obtaining the output frame of the current frame image (i.e., the positioning frame output by the position prediction model for the target object in the current frame image), the electronic device does not directly display the output frame predicted by the model on the current frame image, but can combine the position information of the output frame with the position information of the mark frame displayed on the historical frame image (such as the last frame image or the previous frame image). For example, the position information of the mark frame displayed on the historical frame image and the position information of the output frame of the current frame image are fused based on adaptive weights to obtain the position information of the mark frame finally displayed on the current frame image. However, even if the electronic device remains relatively stationary with the target object, the output frame of the current frame image predicted by the position prediction model will still change randomly, and thus the mark frame finally obtained by combining the position information of the mark frame of the historical frame image and the position information of the output frame of the current frame image will also change to a large extent. Therefore, this solution can only alleviate the jitter of the mark frame, such as reducing the amplitude of the jitter, but cannot eliminate the jitter of the mark frame, which will also result in poor user experience.
[0068] Based on this, in an embodiment of the present application, a target object marking method is proposed, which can enable the stable display of the mark frame of the target object in the shooting picture of the electronic device when the electronic device remains relatively stationary or does not produce relative movement with the target object, i.e., the mark frame remains stationary in the shooting picture without jitter.
[0069] Specifically, during the process of continuous acquisition of frame images by the camera, the electronic device does not directly display the output frame obtained by real-time position prediction of the target object in each acquired frame image, but analyzes and processes the target object marking method provided by the present application to obtain the mark frame finally displayed in each frame image.
[0070] In some embodiments, during the process of continuous acquisition of frame images by the camera, in the case where the electronic device remains relatively stationary with the target object, a first coordinate frame for marking the target object, such as an anchor point frame, is fixedly displayed; in the case where the electronic device moves relatively with the target object, a second coordinate frame predicted for the target object in the currently acquired frame image is displayed in real time, i.e., the output frame predicted in real time.
[0071] It should be understood that the electronic device can use the position prediction model to perform real-time position prediction of the target object in each acquired frame image to obtain the output frame, or can use other position prediction methods to predict the position of the target object in each frame image in real time to obtain the output frame, which is not limited.
[0072] Figure 4This is a simplified schematic diagram illustrating the principle of an embodiment of this application. Figure 4 As shown, in the embodiments of this application, the target object marking method (also known as the jitter suppression method) can be set after the position prediction model to facilitate the adjustment of the output box of the position prediction model, so that the marking box finally displayed in the captured image remains stationary and no longer jitters. Specifically, the target object marking method of this application can post-process the output box of the position prediction model to obtain the marking box finally displayed on the captured image of the current frame.
[0073] It should be understood that the location prediction model in this application refers to any model with location prediction capabilities, such as a single-target tracking (SOT) model or other models capable of predicting the location of a target object in an image. Therefore, this target object labeling method can be applied to the output bounding box of any location prediction model, and it is plug-and-play. Furthermore, this application is implemented entirely through logical judgment, with almost no additional computational overhead or performance / power consumption risks. Moreover, this application can accept only the output bounding box of the location prediction model as input, making it applicable to any scenario with a coordinate bounding box exhibiting jitter. Traditional jitter removal algorithms typically trade off stability and sensitivity, but this application can maintain sufficient sensitivity while removing jitter, making it suitable for a wide range of applications.
[0074] This application provides a target object identification method applied to an electronic device with a camera. During the continuous acquisition of frame images by the camera, the electronic device can obtain a first output box predicted for a target object in the currently acquired frame image. When the electronic device and the target object remain relatively stationary within a target time period, an anchor box is displayed in the current frame image to identify the target object. The anchor box is a second output box predicted for the target object in historical frame images, used as a reference. The target time period refers to the time interval between the first acquisition time of the historical frame image corresponding to the anchor box and the second acquisition time of the current frame image. That is, in this application embodiment, the output box predicted for historical frame images acquired before the current frame image (i.e., the second output box) is used as a reference, and the anchor box is directly output when the electronic device and the target object remain relatively stationary, thereby avoiding frame jitter.
[0075] Specifically, the electronic device acquires anchor boxes, which are determined from the output boxes corresponding to the target object in historical frame images. Historical frame images are those captured before the current frame image. During the target time period between the acquisition time of the historical frame image and the acquisition time of the current frame image, if the electronic device and the target object remain relatively stationary, the electronic device can directly mark the target object in the current frame image based on the anchor boxes, without needing to output the first output box predicted by the position prediction model for the current frame image (i.e., the coordinate box predicted by the position prediction model for the target object in the current frame image). Furthermore, if the electronic device and the target object move relatively during this target time period, the electronic device can mark the target object in the current frame image based on the first output box predicted by the position prediction model and use this output box as the new anchor boxes.
[0076] Figure 5 Another simplified schematic diagram of the principle provided for an embodiment of this application, as shown below. Figure 5 As shown, Figure 5 The image displayed on the screen when an electronic device is taking a picture of a person. Figure 5 The person in the image is the target object, and the electronic device displays a marker box on the screen to indicate the person in the captured image. Assuming the marker box displayed in frame 1 is the output box predicted by the model for frame 1, the electronic device can use the marker box displayed in frame 1 as the anchor box 501. It is understood that, for illustrative purposes, the anchor box in the figure is marked with a solid line, and the output box is marked with a dashed line.
[0077] Therefore, when the model performs location detection or position prediction for the person in frame image n (n≥2), if the electronic device and the person remain relatively stationary during the time from frame image 1 to frame image n, the output box predicted by the model for frame image n is as follows: Figure 5 As shown in the dashed box 502 in (a1), the electronic device can use the anchor box 501 as the final label box of the frame image n, and display the anchor box 501 on the frame image n, that is, as shown in the figure. Figure 5 As shown in (a2).
[0078] Please continue reading. Figure 5 If, during the time interval from frame 1 to frame n, the electronic device and the person move relative to each other, the model predicts the output box for frame n as follows: Figure 5 As shown in the dashed box 503 in (b1), the electronic device can use the output box 503 as the final label box of the frame image n, and display the output box 503 on the frame image n, that is, as shown in the figure. Figure 5 As shown in (b2).
[0079] It can be understood that the reason for the jitter phenomenon is that there is a change in position or size between the mark frame displayed on the current frame image and the mark frame displayed on the previous frame image, so that the mark frame in the picture presents a jitter situation. In the embodiment of the present application, in the case that the electronic device and the target object do not produce relative movement in the target time period, the target object is marked based on the anchor point frame in the current frame image. In this way, the same frame is displayed on the current frame image and the previous frame image, that is, the anchor point frame is displayed, so that the mark frame displayed in the picture no longer appears the jitter phenomenon.
[0080] Figure 6 A structural schematic diagram of an electronic device is provided in the embodiment of the present application.
[0081] Please refer to Figure 6 In the present application, the electronic device 600 is taken as an example of a mobile phone, and the electronic device 600 provided by the present application is introduced. As shown in Figure 6 The electronic device 600 can include a processor 610, an external memory interface 620, an internal memory 621, a universal serial bus (USB) interface 630, a charge management module 640, a power management module 641, a battery 642, an antenna 1, an antenna 2, a mobile communication module 650, a wireless communication module 660, an audio module 670, a speaker 670A, a receiver 670B, a microphone 670C, a headset interface 670D, a sensor module 680, a key 690, an indicator 692, a camera 693, a display screen 694, and the like. The sensor module 680 can include a gyroscope sensor 680A, a gravity sensor 680B, a touch sensor 680C, and the like.
[0082] It can be understood that the structure shown in the embodiment of the present application does not constitute a specific limitation on the electronic device 600. In other embodiments of the present application, the electronic device 600 can include more or fewer components than shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0083] The processor 610 can include one or more processing units, for example: the processor 610 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors.
[0084] In the embodiments of the present application, when the camera 693 is turned on and shooting is performed, the processor 610 displays the content shot within the field of view of the camera 693 on the display screen 694. The processor 610 can focus on the target object in the shot picture and generate a mark box. When the relative displacement between the target object and the electronic device 600 occurs, the processor 610 detects and tracks the position and motion trail of the target object in the shot picture through a related position prediction model, so that the mark box can always frame the target object and focus on the target object, and the mark box is continuously displayed on the shot picture.
[0085] The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.
[0086] The memory in the processor 610 can also be configured to store instructions and data. In some embodiments, the memory in the processor 610 is a cache memory. The memory can save instructions or data that have just been used or are repeatedly used by the processor 610. If the processor 610 needs to use the instructions or data again, it can directly call them from the memory. This avoids repeated access and reduces the waiting time of the processor 610, thereby improving the efficiency of the system.
[0087] In some embodiments, the processor 610 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0088] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a limitation on the structure of the electronic device 600. In some other embodiments of the present application, the electronic device 600 can also use different interface connection modes or a combination of multiple interface connection modes.
[0089] The charging management module 640 is configured to receive charging input from a charger. The charging management module 640 can also supply power to the electronic device through the power management module 641 while charging the battery 642. The power management module 641 is configured to connect the battery 642, the charging management module 640, and the processor 610. The power management module 641 receives input from the battery 642 and / or the charging management module 640 to supply power to the processor 610, the internal memory 621, the display screen 694, the camera 693, and the wireless communication module 660, etc.
[0090] The wireless communication function of the electronic device 600 can be realized through the antenna 1, the antenna 2, the mobile communication module 650, the wireless communication module 660, the modem processor, and the baseband processor, etc.
[0091] The antenna 1 and the antenna 2 are configured to transmit and receive electromagnetic wave signals.
[0092] The mobile communication module 650 can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. The mobile communication module 650 can receive electromagnetic waves by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transfer the processed signals to the modem processor for demodulation. The mobile communication module 650 can also amplify the signals modulated by the modem processor, and radiate the signals as electromagnetic waves through the antenna 1.
[0093] The wireless communication module 660 can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. The wireless communication module 660 receives electromagnetic waves via the antenna 2, and performs frequency modulation and filtering on the electromagnetic wave signals, and transmits the processed signals to the processor 610. The wireless communication module 660 can also receive signals to be transmitted from the processor 610, perform frequency modulation and amplification on the signals, and radiate the signals as electromagnetic waves through the antenna 2.
[0094] In some embodiments, the antenna 1 and the mobile communication module 650 of the electronic device 600 are coupled, and the antenna 2 and the wireless communication module 660 are coupled, so that the electronic device 600 can communicate with a network and other devices through wireless communication technology.
[0095] The electronic device 600 implements a display function through a GPU, a display screen 694, and an application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display screen 694 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 610 can include one or more GPUs, which execute program instructions to generate or change display information.
[0096] The display screen 694 is configured to display images, videos, and the like. The display screen 694 includes a display panel. The display panel can be a liquid crystal display (LCD). The display panel can also be manufactured by using an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a miniled, a microled, a micro-oled, a quantum dot light emitting diodes (QLED), and the like. In some embodiments, the electronic device can include one or N display screens 694, where N is a positive integer greater than 1.
[0097] In embodiments of the present application, the display screen 694 can display the image captured by the camera 693, and can also display a mark box corresponding to the target object in the captured image.
[0098] The electronic device 600 can implement the photographing function by using the ISP, the camera 693, the video codec, the GPU, the display screen 694, and the application processor.
[0099] The ISP is configured to process the data fed back by the camera 693. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electric signal, and the camera photosensitive element transmits the electric signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can also optimize the algorithm for the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be arranged in the camera 693.
[0100] The camera 693 is configured to capture still images or videos. An object projects an optical image through a lens to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, which is then transmitted to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into a standard image signal in a format such as RGB, YUV, or the like. In some embodiments, the electronic device 600 can include one or N cameras 693, where N is a positive integer greater than 1.
[0101] The digital signal processor is configured to process digital signals, including digital image signals. For example, when the electronic device 600 is selecting a frequency point, the digital signal processor is configured to perform Fourier transform on the frequency point energy, and the like.
[0102] The video codec is configured to compress or decompress digital videos. The electronic device 600 can support one or more video codecs. In this way, the electronic device 600 can play or record videos in multiple encoding formats, such as moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, and the like.
[0103] The NPU is a neural-network (NN) computing processor that is configured to process input information quickly by imitating the structure of a biological neural network, such as the transmission mode between human neurons, and is also configured to learn continuously. Through the NPU, the electronic device 600 can implement intelligent cognition applications, such as image recognition, face recognition, voice recognition, text understanding, and the like.
[0104] In the embodiments of the present application, the electronic device 600 implements the target object marking method provided in the embodiments of the present application, first relying on the ISP, the image collected by the camera 693, and second relying on the video codec, the image computing and processing capability provided by the GPU. The electronic device 600 can implement neural network algorithms such as face recognition, human body recognition, and re-identification (ReID) through the computing and processing capability provided by the NPU.
[0105] The internal memory 621 can include one or more random access memories (RAMs) and one or more non-volatile memories (NVMs).
[0106] The random access memory can include static random-access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM, such as the fifth generation of DDR SDRAM commonly referred to as DDR5 SDRAM), and the like.
[0107] The non-volatile memory can include a magnetic disk storage device, flash memory. The flash memory can include NOR FLASH, NAND FLASH, 3D NAND FLASH, and the like according to the operating principle, and can include single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), and the like according to the storage unit potential order, and can include universal flash storage (UFS), embedded multi media Card (eMMC), and the like according to the storage specification.
[0108] The random access memory can be directly read and written by the processor 610, and can be used to store executable programs (such as machine instructions) of an operating system or other programs running, and can also be used to store data of users and application programs, and the like. The non-volatile memory can also store executable programs and store data of users and application programs, and the like, and can be loaded into the random access memory in advance for direct reading and writing by the processor 610.
[0109] In the embodiments of the present application, the code for implementing the target object marking method of the embodiments of the present application can be stored on the non-volatile memory. When the camera application is running, the electronic device 600 can load the executable code stored in the non-volatile memory into the random access memory.
[0110] The external memory interface 620 can be used to connect an external nonvolatile memory, to extend the storage capacity of the electronic device 600. The external nonvolatile memory communicates with the processor 610 via the external memory interface 620, to implement a data storage function.
[0111] The electronic device 600 can implement an audio function through an audio module 670, a speaker 670A, a receiver 670B, a microphone 670C, an earphone interface 670D, an application processor, and the like. For example, music playback, voice recording, and the like.
[0112] The audio module 670 is used to convert digital audio information into an analog audio signal output, and is also used to convert an analog audio input into a digital audio signal. The speaker 670A, also known as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 600 can listen to music or listen to a hands-free call through the speaker 670A. The receiver 670B, also known as a "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 600 answers a call or a voice message, the receiver 670B can be held close to the ear to listen to the voice. The microphone 670C, also known as a "microphone", "sounder", is used to convert a sound signal into an electrical signal. The earphone interface 670D is used to connect a wired earphone.
[0113] The gyroscope sensor 680A can be used to determine the motion posture of the electronic device 600. In some embodiments, the angular velocity of the electronic device 600 around three axes (i.e., x, y, and z axes) can be determined through the gyroscope sensor 680A, to determine the motion posture of the electronic device 600.
[0114] The gravity sensor 680B is made of a cantilever type displacement device made of an elastic sensitive element, and a storage spring made of an elastic sensitive element is used to drive an electrical contact to complete the conversion from gravity change to electrical signal. In the embodiments of the present application, whether a relative displacement is generated between the electronic device 600 and a target object can be determined through the gravity sensor 680B.
[0115] The touch sensor 680C, also known as a "touch device". The touch sensor 680C can be disposed on the display screen 694, and the touch sensor 680C and the display screen 694 form a touch screen, also known as a "touch screen". The touch sensor 680C is used to detect a touch operation acting on or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the touch event type. The display screen 694 can provide visual output related to the touch operation. In other embodiments, the touch sensor 680C can also be disposed on the surface of the electronic device 600, which is different from the position of the display screen 694.
[0116] In the embodiments of the present application, whether a relative displacement is generated between the electronic device 600 and the target object can be determined in combination with the gyro sensor 680A and the gravity sensor 680B.
[0117] The keys 690 include a power-on key, a volume key, and the like. The keys 690 can be mechanical keys. Alternatively, the keys 690 can be touch keys. The electronic device 600 can receive a key input and generate a key signal input related to user settings and function control of the electronic device 600.
[0118] In summary, in the embodiments of the present application, when it is determined by the gyro sensor 680A and the gravity sensor 680B that no relative displacement is generated between the electronic device 600 and the target object within a target time period, the electronic device enters a still mode, the target object in the current frame image is marked based on the anchor box, and the shooting picture displayed on the display screen 694.
[0119] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device can include more or fewer components than those illustrated, or combine certain components, or split certain components, or different arrangement of components. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.
[0120] Figure 7 The system structure schematic diagram of the electronic device provided in the embodiments of the present application is shown.
[0121] The software system of the electronic device 600 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. The embodiments of the present application take the Android system with a layered architecture as an example to illustrate the software structure of the electronic device 600.
[0122] The layered architecture divides the system into several layers, each layer has a clear role and division of labor. The layers communicate with each other through a software interface. In some embodiments, the system is divided into five layers, from top to bottom, the application layer, the application framework layer, the hardware abstraction layer, the driver layer, and the hardware layer.
[0123] The application layer can include a series of application packages. In the embodiments of the present application, the application packages can include a camera, a gallery, and the like.
[0124] The application framework layer provides the application programs of the application layer with application programming interfaces (APIs) and programming frameworks. The application framework layer includes some pre-defined functions. In the embodiments of the present application, the application framework layer can include a camera access interface, wherein the camera access interface can include camera management and camera devices. The camera access interface is used to provide the camera application with application programming interfaces and programming frameworks.
[0125] The hardware abstraction layer is an interface layer between the application framework layer and the driver layer, and provides a virtual hardware platform for the operating system. In the embodiment of the present application, the hardware abstraction layer can include a camera hardware abstraction layer and a camera algorithm library.
[0126] The camera hardware abstraction layer can provide virtual hardware of the camera device 1, the camera device 2 or more camera devices. The camera algorithm library can include running code and data for implementing the target object marking method provided in the embodiment of the present application.
[0127] The driver layer is a layer between hardware and software. The driver layer includes drivers of various hardware. The driver layer can include a camera device driver, a digital signal processor driver and an image processor driver, etc.
[0128] The camera device driver is used to drive the sensor of the camera to collect images and drive the image signal processor to pre-process the images. The digital signal processor driver is used to drive the digital signal processor to process the images. The image processor driver is used to drive the image processor to process the images.
[0129] The target object marking method in the embodiment of the present application will be described in detail below in combination with the above system structure:
[0130] (1) In response to the operation of the user opening the camera application, for example, the operation of clicking the camera application icon, the camera application is started by calling the camera access interface of the application framework layer, and the camera interface is entered to display the current shooting picture on the screen, and then the instruction of starting the camera is sent by calling the virtual hardware of the camera device (the camera device 1 and / or other camera devices) in the camera hardware abstraction layer. The camera hardware abstraction layer sends the instruction to the camera device driver in the kernel layer. The camera device driver can start the corresponding camera sensor and collect image light signals through the sensor. One camera device in the camera hardware abstraction layer corresponds to one camera sensor in the hardware layer.
[0131] (2) The camera sensor can transmit the collected image light signals to the image signal processor for pre-processing to obtain image electrical signals (original images), and transmit the original images to the camera hardware abstraction layer through the camera device driver.
[0132] (3) The camera hardware abstraction layer can send the original images to the camera algorithm library. The camera algorithm library stores program codes for implementing the target object marking method provided in the embodiment of the present application.
[0133] (4) Based on a digital signal processor, an image processor, and a camera algorithm library, the target object marking method determines the target object in the picture through the position prediction model first, and displays a marking box in the shooting picture to frame the target object. Specifically, in the case that the electronic device and the target object do not produce relative movement within the target time period, the target object is marked on the current shooting picture by an anchor box. In the case that the electronic device and the target object produce relative movement within the target time period, the target object is marked on the current shooting picture by an output box corresponding to the current shooting picture, and the output box is used as a new anchor box.
[0134] (5) The camera algorithm library can send the recognized original image collected by the camera to the camera hardware abstraction layer. Then, the camera hardware abstraction layer can display it.
[0135] The target object marking method of the present application will be specifically described below in combination with specific embodiments and drawings. Figure 8 The flowchart of a target object marking method provided by an embodiment of the present application is shown in Figure 8 It can be seen that the target object marking method of the present application comprises the following steps:
[0136] S801: In the process of continuously collecting frame images by the camera, a first output box predicted by a position prediction model for a target object in a collected current frame image is obtained.
[0137] In some embodiments, in the process of shooting the target object by the electronic device, the position prediction model identifies the target object in the image and predicts an output box, and the electronic device obtains the output box.
[0138] S802: It is judged whether the electronic device and the target object produce relative movement within a target time period.
[0139] In some embodiments, if no, S803 is executed, and if yes, S804 is executed.
[0140] In some embodiments, the electronic device can judge whether the relative change between the output box corresponding to the target object in the current frame image and the anchor box exceeds a preset change threshold, to determine whether the electronic device and the target object produce relative movement within the target time period. If the relative change exceeds the preset change threshold, it is considered that relative movement is produced, and if the relative change does not exceed the preset change threshold, it is considered that relative movement is not produced.
[0141] In some embodiments, in the case that the electronic device and the target object remain relatively stationary within the target time period, the anchor box is displayed in the current frame image to mark the target object.
[0142] Specifically, in a case where the relative change between the first output frame and the anchor frame does not exceed the preset change threshold, the target object is marked in the current frame image based on the anchor frame, that is, S803 is performed.
[0143] In some embodiments, in a case where the electronic device and the target object have a relative movement within the target time period, the first output frame is displayed in the current frame image to mark the target object, and the first output frame is taken as a new anchor frame.
[0144] Specifically, in a case where the relative change between the first output frame and the anchor frame exceeds the preset change threshold, the target object is marked in the current frame image based on the first output frame, and the first output frame is taken as a new anchor frame, that is, S804 is performed.
[0145] In some embodiments, judging whether the relative change between the output frame corresponding to the target object in the current frame image and the anchor frame exceeds the preset change threshold comprises judging whether the position change between the output frame corresponding to the target object in the current frame image and the anchor frame exceeds a preset position change threshold, and / or judging whether the size change between the output frame corresponding to the target object in the current frame image and the anchor frame exceeds a preset size change threshold.
[0146] In some embodiments, in a case where the position change between the first output frame and the anchor frame does not exceed the preset position change threshold and / or the size change between the first output frame and the anchor frame does not exceed the preset size change threshold, the target object is marked in the current frame image based on the anchor frame.
[0147] In some embodiments, the preset position change threshold comprises a first position change threshold and a second position change threshold, and the preset size change threshold comprises a first size change threshold and a second size change threshold. If the electronic device is not in the stationary mode, the first position change threshold is used to judge the position change between the first output frame and the anchor frame, and / or the first size change threshold is used to judge the size change between the first output frame and the anchor frame. If the electronic device is in the stationary mode, the second position change threshold is used to judge the position change between the first output frame and the anchor frame, and / or the second size change threshold is used to judge the size change between the first output frame and the anchor frame.
[0148] In this case, the purpose of it being more difficult to enter the stationary mode than to exit the stationary mode is achieved, so as to avoid easily entering the stationary mode in some special cases that are not truly relatively stationary (such as slow movement), thereby making the display of the mark frame displayed on the screen more accurate and the user experience better.
[0149] How to determine whether the position change between the output box corresponding to the target object in the current frame image and the anchor box exceeds the preset position change threshold, and / or whether the size change between the output box corresponding to the target object in the current frame image and the anchor box exceeds the preset size change threshold, can be seen from S1205 and S1211.
[0150] In some embodiments of the present application, the electronic device determines whether the electronic device moves by using sensor data collected by a gyroscope sensor and a gravity sensor.
[0151] In some embodiments of the present application, the electronic device can determine whether the picture taken by the electronic device changes within a preset time period. If yes, it can be determined that the electronic device and the target object move relatively within the target time period. If no, it can be determined that the electronic device and the target object do not move relatively within the target time period, i.e., remain relatively stationary.
[0152] S803: Display the anchor box in the current frame image to indicate the target object. The anchor box is the second output box predicted by the position prediction model for the target object in the historical frame image.
[0153] S804: Display the first output box in the current frame image to indicate the target object, and take the first output box as a new anchor box.
[0154] In some embodiments of the present application, Figure 9 Another flowchart of a target object indicating method provided by an embodiment of the present application is shown in the figure. In the figure, the specific steps are as follows:
[0155] S901: In the process of continuously collecting frame images by the camera, obtain the first output box predicted by the position prediction model for the target object in the collected current frame image.
[0156] S902: Determine whether the electronic device is in a stationary mode.
[0157] In some embodiments, the stationary mode means that when the electronic device is in the stationary mode, the second output box predicted for the target object in the historical frame image and displayed most recently is used to indicate the target object in the current frame image, i.e., the stationary mode is a mode in which the anchor box is output and displayed.
[0158] In some embodiments, if not, S903 is performed, and if yes, S908 is performed.
[0159] S903: Determine whether the position change between the first output box and the anchor box exceeds the first position change threshold, and / or whether the size change between the first output box and the anchor box exceeds the first size change threshold.
[0160] In some embodiments, if the positional change between the first output box and the anchor box does not exceed a first positional change threshold, and the size change between the first output box and the anchor box does not exceed a first size change threshold, then S904 is executed. If the positional change between the first output box and the anchor box exceeds the first positional change threshold, and / or the size change between the first output box and the anchor box exceeds the first size change threshold, then S907 is executed.
[0161] S904: Determine whether the number of stable frames has reached the preset frame number threshold.
[0162] In some embodiments, if the condition is not met, S905 is executed; if the condition is met, S906 is executed.
[0163] In some embodiments, before the electronic device enters a static mode, a determination of the number of stable frames is also required. Specifically, the relative change between the output bounding box predicted by the position prediction model for the target object and the anchor point bounding box in the stable frame image does not exceed a preset change threshold. The number of stable frames refers to the number of stable frame images existing between historical frame images and the current frame image.
[0164] S905: Mark the target object in the current frame image based on the first output box, keep the anchor point box unchanged, and increment the number of stable frames by 1.
[0165] In some embodiments, when the electronic device is not in a static mode, the relative change does not exceed a preset change threshold, and the number of stable frames does not reach a preset frame number threshold, the target object is marked in the current frame image based on the first output box, the anchor point box remains unchanged, and the number of stable frames is incremented by 1.
[0166] For example, when an electronic device takes a picture of a person, the electronic device generates multiple consecutive frame images (e.g., frame image 1 to frame image 5) in real time based on the scene during the shooting process and displays them on the screen. Figure 10 This is a schematic diagram illustrating the principle of an electronic device entering a static mode, as provided in an embodiment of this application. Figure 10 As shown, this figure contains five consecutive frame images (frame images 1 to 5). The dashed boxes in 1001a, 1002a, 1003a, 1004a, and 1005a represent the output boxes predicted by the location prediction model for the target objects in frames 1 to 5. Assuming the output box of frame image 1 is taken as the anchor box, the solid boxes in 1002a, 1003a, 1004a, and 1005a are the anchor boxes. The boxes in 1001b, 1002b, 1003b, 1004b, and 1005b represent the boxes that ultimately mark the person on frames 1 to 5.
[0167] from Figure 10It can be seen that the electronic device is not in the stationary mode before the frame image 1, and the output frame of the frame image 1 will be used as the marking frame of the person and displayed on the frame image 1, and the output frame of the frame image 1 will be used as the anchor frame, as shown in 1001b.
[0168] It can be seen from 1002a that the position change between the output frame (i.e., the dashed frame) in 1002a and the anchor frame does not exceed the first position change threshold, and the size change does not exceed the first size change threshold. In this case, the electronic device and the person do not have relative movement in the time period between the capture time of the frame image 1 and the capture time of the frame image 2, but the electronic device is not in the stationary state, and the output frame of the frame image 2 will be used as the marking frame of the person and displayed on the frame image 2, and the original anchor frame will be kept unchanged, and the number of stable frames will be increased by one.
[0169] Similarly, the output frame of the frame image will be used as the marking frame of the person and displayed on the frame image, and the original anchor frame will be kept unchanged, and the number of stable frames will be increased by one. In the frame image 1 to the frame image 4, since the number of stable frames of the electronic device does not reach the preset frame number threshold, the electronic device is not in the stationary mode.
[0170] S906: The first output frame is used as a new anchor frame to mark the target object in the current frame image based on the new anchor frame, and the stationary mode is entered.
[0171] In some embodiments, in the case that the electronic device is not in the stationary mode, the relative change does not exceed the preset change threshold, and the number of stable frames reaches the preset frame number threshold, the first output frame is used as a new anchor frame to mark the target object in the current frame image based on the new anchor frame, and the stationary mode is entered.
[0172] An example is shown in the foregoing Figure 10 As shown in the foregoing, it is assumed that the stable frame number threshold is 3, the number of stable frames is 0 at the frame image 1, the number of stable frames is increased by one at the frame image 2, the frame image 3, and the frame image 4, and the number of stable frames reaches the preset frame number threshold at the frame image 5. In this case, the output frame of the frame image 5 will be used as the marking frame of the person and displayed on the frame image 5, and the output frame of the frame image 5 will be used as the new anchor frame, and the electronic device enters the stationary mode.
[0173] In some embodiments of the present application, after the electronic device enters the stationary mode, the target object can also be marked in the current frame image based on the output frame (anchor frame) corresponding to the target object in the historical frame image.
[0174] In some other embodiments of this application, after the electronic device enters the static mode, the position and size of the output box of the current frame image can be adjusted based on the anchor point box to obtain a new output box of the current frame image, and the target object can be marked in the current frame image based on the new output box, and the new output box can be used as a new anchor point box.
[0175] S907: Mark the target object in the current frame image based on the first output box, and use the first output box as the new anchor point box.
[0176] S908: Determine whether the positional change between the first output box and the anchor box exceeds the second positional change threshold, and / or whether the size change between the first output box and the anchor box exceeds the second size change threshold.
[0177] In some embodiments, if the positional change between the first output box and the anchor box does not exceed a second positional change threshold, and the size change between the first output box and the anchor box does not exceed a second size change threshold, then S910 is executed. If the positional change between the first output box and the anchor box exceeds the second positional change threshold, and / or the size change between the first output box and the anchor box exceeds the second size change threshold, then S909 is executed.
[0178] S909: Exit still mode, mark the target object in the current frame image based on the first output box, and use the first output box as the new anchor point box.
[0179] In some embodiments, when the electronic device is in a static mode but the relative change exceeds a preset change threshold, the electronic device exits the static mode, marks the target object in the current frame image based on the first output box, and uses the first output box as a new anchor point box.
[0180] One example, Figure 11 This is a schematic diagram illustrating the principle of an electronic device exiting a static mode, as provided in an embodiment of this application. Figure 11 As shown, taking the electronic device photographing flowers as an example, the electronic device generates multiple consecutive frame images (e.g., frame image 1 to frame image 4) in real time based on the scene during the shooting process, and displays them on the screen. Figure 11 As shown, the dashed boxes in 1101a, 1102a, 1103a, and 1104a are used to indicate the output boxes predicted by the position prediction model for the target objects in frames 1 to 4, respectively. The solid boxes in 1101a, 1102a, 1103a, and 1104a are anchor boxes, while the boxes in 1101b, 1102b, 1103b, and 1104b are used to indicate the boxes that will ultimately mark the flowers on frames 1 to 4.
[0181] from Figure 11It can be seen that the electronic device is in the stationary mode before the frame image 1, and the position changes between the output frame and the anchor frame for the flowers (i.e., the target object) in the frame images 1 to 3 do not exceed the first position change threshold, and the size changes do not exceed the first size change threshold. Therefore, as shown in the frame images 1101b, 1102b and 1103b, the anchor frame will be used as the marking frame of the flowers in the frame images 1 to 3.
[0182] Due to the relative displacement between the electronic device and the target object, the position change between the output frame and the anchor frame in the frame image 4 exceeds the first position change threshold, or the size change between the output frame and the anchor frame corresponding to the target object in the frame image 4 exceeds the first size change threshold. In this case, the electronic device will exit the stationary mode, save the output frame in 1104a as a new anchor frame, and use the new anchor frame (i.e., the output frame in 1104a) as the marking frame of the flowers in the frame image 4 as shown in 1104b, and reset the stable frame number to zero. It can be understood that the output frame in 1104a is used as the new anchor frame, so it is represented by a solid line frame in 1104b.
[0183] S910: Keep the stationary mode, and mark the target object based on the anchor frame in the current frame image.
[0184] In some embodiments, the target object is marked based on the anchor frame in the current frame image when the electronic device has been in the stationary mode and the relative change does not exceed the preset change threshold.
[0185] Figure 12 Another flowchart of a target object marking method provided by an embodiment of the present application is shown in FIG. 12. The specific steps are as follows:
[0186] S1201: Obtain a first output frame of a current frame image.
[0187] The electronic device determines the position of the target object in the current frame image by using a position prediction model, and obtains an output frame predicted by the position prediction model for the target object in the current frame image.
[0188] S1202: Calculate the center point, diagonal length and side length of the first output frame of the current frame image.
[0189] S1203: Whether the current is in the stationary mode.
[0190] In some embodiments, if the electronic device is not in the stationary mode, S1204 is performed, and if the electronic device is in the stationary mode, S1211 is performed.
[0191] S1204: Whether there is an anchor frame.
[0192] In some embodiments of the present application, if the electronic device is not in the stationary mode and the anchor box exists, S1205 is performed. If the electronic device is not in the stationary mode and the anchor box does not exist, S1210 is performed. For example, when the target object is marked with a mark box on the first frame image, the first mark box is obtained. At this time, the electronic device is not in the stationary mode and the anchor box does not exist, so the mark box of the first frame image can be directly saved as the anchor box.
[0193] S1205: Whether the position change between the first output box and the anchor box exceeds the first position change threshold, and / or whether the size change between the first output box and the anchor box exceeds the first size change threshold.
[0194] In some embodiments, the first position change threshold includes a first threshold, and the target object marking method further includes: determining that the position change does not exceed the preset position change threshold if the center point displacement between the first output box and the anchor box does not exceed the first threshold. And / or, the first size change threshold includes a second threshold and / or a third threshold, and it is determined that the size change does not exceed the first size change threshold if the side length change between the first output box and the anchor box does not exceed the second threshold, and / or the diagonal length change between the first output box and the anchor box does not exceed the third threshold.
[0195] In some embodiments, if the position change between the output box corresponding to the current frame image and the anchor box does not exceed the first position change threshold, and the size change between the output box corresponding to the current frame image and the anchor box does not exceed the first size change threshold, S1206 is performed.
[0196] In some embodiments, if the position change between the output box corresponding to the current frame image and the anchor box exceeds the first position change threshold, or the size change between the output box corresponding to the current frame image and the anchor box exceeds the first size change threshold, S1210 is performed.
[0197] S1206: Whether the number of stable frames reaches a preset frame number threshold.
[0198] In some embodiments, if the number of stable frames does not reach the preset frame number threshold, S1207 is performed. If the number of stable frames reaches the preset frame number threshold, S1208 is performed.
[0199] S1207: Marking the target object in the current frame image based on the first output box, keeping the anchor box unchanged, and increasing the number of stable frames by 1.
[0200] S1208: Taking the first output box as a new anchor box to mark the target object in the current frame image based on the new anchor box.
[0201] S1209: Enter the static mode.
[0202] In some embodiments, the execution order between S1208 and S1209 is not limited. It can be understood that after S1209 is executed, the first output frame can be taken as a new anchor point frame, and subsequent frame images are processed accordingly.
[0203] S1210: Label the target object in the current frame image based on the first output frame, take the first output frame as a new anchor point frame, and reset the count of stable frames to zero.
[0204] In some embodiments, in the case that the electronic device is not in the static mode, and the position change between the output frame corresponding to the current frame image and the anchor point frame exceeds the first position change threshold, and / or the size change between the output frame corresponding to the current frame image and the anchor point frame exceeds the first size change threshold, the output frame of the current frame image is saved as a new anchor point frame, and the output frame is displayed on the current frame image as a label frame of the target object, and the count of stable frames is reset to zero.
[0205] In other embodiments, in the case that the electronic device exits the static mode, the output frame of the current frame image is saved as a new anchor point frame, and the output frame is displayed on the current frame image as a label frame of the target object, and the count of stable frames is reset to zero.
[0206] S1211: Whether the position change between the first output frame and the anchor point frame exceeds the second position change threshold, and / or whether the size change between the first output frame and the anchor point frame exceeds the second size change threshold.
[0207] In some embodiments, the second position change threshold includes a fourth threshold, and the target object labeling method further includes: in the case that the center point displacement between the first output frame and the anchor point frame does not exceed the fourth threshold, it is determined that the position change does not exceed the second position change threshold. And / or, the second size change threshold includes a fifth threshold and / or a sixth threshold, in the case that the side length change between the first output frame and the anchor point frame does not exceed the fifth threshold, and / or the diagonal length change between the first output frame and the anchor point frame does not exceed the sixth threshold, it is determined that the size change does not exceed the second size change threshold.
[0208] In some embodiments, if the position change between the output frame corresponding to the current frame image and the anchor point frame does not exceed the second position change threshold, and the size change between the output frame corresponding to the current frame image and the anchor point frame does not exceed the second size change threshold, S1213 is executed.
[0209] In some embodiments, if the position change between the output frame corresponding to the current frame image and the anchor frame exceeds the second position change threshold, or the size change between the output frame corresponding to the current frame image and the anchor frame exceeds the second size change threshold, S1212 is performed.
[0210] S1212: exit the static mode.
[0211] In some embodiments, after the electronic device exits the static mode, S1210 is performed.
[0212] S1213: maintain the static mode.
[0213] In some embodiments, after the electronic device maintains the static mode, S1214 is performed.
[0214] S1214: mark the target object based on the anchor frame in the current frame image.
[0215] The inventors of the present application found that, in the case that the electronic device is slowly moved by the user, if the method in the above embodiments is used, the electronic device may frequently enter or exit the static mode. Because, due to the slow movement, some frame images will be misrecognized as stable frame images in a short time, and then, after the number of stable frames accumulates to the preset frame number threshold, the electronic device will enter the static mode in the actual slow movement. After the electronic device enters the static mode, if the electronic device is slowly moved by the user, after a certain time length is accumulated, the position change between the output frame corresponding to the frame image and the anchor frame will exceed the second position change threshold, or the size change between the output frame corresponding to the frame image and the anchor frame will exceed the second size change threshold, and the electronic device will exit the static mode. Therefore, this will cause the electronic device to frequently enter or exit the static mode in a period of time, so that the electronic device will use the anchor frame as the marking frame of the frame image at times, and will use the output frame as the marking frame of the frame image at times, and because the change is frequent, the phenomenon of marking frame freezing may also occur in the picture.
[0216] Therefore, the inventors of the present application propose the following scheme to avoid the electronic device frequently entering and exiting the static mode in a preset time period in the case of slow movement and the like, so as to alleviate the freezing of the marking frame.
[0217] In some embodiments, in the case that the number of state changes of the electronic device in a preset time window exceeds a preset number threshold, the preset frame number threshold is increased. The number of state changes is used to represent the number of state changes of the electronic device in the process of processing the frame images collected in the preset time window. The state change includes entering the static mode and exiting the static mode.
[0218] Specifically, in the embodiments of the present application, in order to avoid the electronic device frequently entering and exiting the static mode within the preset time period, the stability of the electronic device can be determined through the time window, that is, whether the electronic device frequently enters and exits the static mode within the time window is determined, and if the stability of the electronic device is not high within the time window, that is, the electronic device frequently enters and exits the static mode within the time window, the preset frame number threshold is increased, so as to increase the difficulty of the electronic device entering the static mode, avoid the electronic device frequently entering and exiting the static mode within the preset time period, prevent the marker frame in the picture from freezing, and improve the user experience.
[0219] Specifically, in the present application, the electronic device can preset the length of a time window, and the length of the time window is used to determine the number of frame images detected each time. The electronic device determines whether the anchor frame is used as the marker frame or the output frame in each frame image in the time window, to determine the situation of the electronic device entering and exiting the static mode within the preset time period.
[0220] Figure 13 A time window detection schematic diagram is provided for the embodiments of the present application. In this diagram, the 8th frame image to the 22nd frame image are all located within the time window.
[0221] Suppose the preset frame number threshold is set to 4, and the output frame of the 5th frame image is used as the anchor frame. Since the electronic device is slowly moved by the user, the difference between the position and size of the output frame of the 6th frame image, the output frame of the 7th frame image, the output frame of the 8th frame image, the output frame of the 9th frame image, and the anchor frame, and the difference between the position and size of the output frame of the 10th frame image and the anchor frame are all less than or equal to the first preset threshold. At this time, in the 10th frame image, the number of stable frames reaches the preset frame number threshold, the electronic device enters the static mode, and the anchor frame is used as the marker frame of the 10th frame image. And between the 11th frame image and the 15th frame image, the electronic device is always in the static mode, and the anchor frame is used as the marker frame of the 11th frame image to the 15th frame image.
[0222] However, subsequently, since the electronic device is slowly moved, in the 16th frame image, the difference between the position and size of the output frame of the 16th frame image and the anchor frame is greater than the second preset threshold, at this time, the static mode is exited, and the output frame of the 16th frame image is used as the anchor frame. But since the electronic device is slowly moved by the user, the difference between the position and size of the output frame of the 17th frame image, the output frame of the 18th frame image, the output frame of the 19th frame image, the output frame of the 20th frame image, and the anchor frame, and the difference between the position and size of the output frame of the 21st frame image and the anchor frame may all be less than or equal to the first preset threshold. At this time, in the 21st frame image, the number of stable frames reaches the preset frame number threshold, the electronic device enters the static mode, and the anchor frame is used as the marker frame of the 21st frame image.
[0223] It can be understood that, from Figure 13 It can be understood that, from
[0224] Therefore, in the present application, the difficulty of entering the standby mode can be increased by increasing the preset frame number threshold, so as to prevent the electronic device from frequently entering or exiting the standby mode within the preset time period, and avoid the phenomenon of frame freezing. That is, if the stable frame number of the electronic device still reaches the preset frame number threshold after increasing the preset frame number threshold, it means that the electronic device and the target object in the frame remain in a static state or a state close to a static state, rather than the electronic device being in a slow moving state. And when the stable frame number reaches the increased preset frame number threshold, the preset frame number threshold can be reset to the value before adjustment.
[0225] Some other embodiments of the present application provide an electronic device, which can include a camera, a memory and one or more processors. The camera, the memory and the processor are coupled. The memory is configured to store computer program code, the computer program code including computer instructions. When the computer instructions are executed by the processor, the electronic device performs each function or step in the above method embodiments. The structure of the electronic device can refer to the structure of the electronic device 600 shown in Figure 6
[0226] The embodiments of the present application also provide a computer storage medium, which includes computer instructions, when the computer instructions are run on the above electronic device, the electronic device performs each function or step performed by the electronic device in the above method embodiments.
[0227] The embodiments of the present application also provide a computer program product, when the computer program product is run on a computer, the computer performs each function or step performed by the electronic device in the above method embodiments.
[0228] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions.
[0229] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the division of the apparatus embodiments is merely illustrative, and for example, the division of the modules or units can be different, and each module or unit can be integrated in a form of another module or unit, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection can be indirect coupling or communication connection through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0230] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, i.e., may be located in one place or distributed in multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0231] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0232] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The software product is stored in a storage medium, and includes a number of instructions for causing an apparatus (which can be a single-chip microcomputer, a chip, etc.) or a processor to perform all or part of the steps of the various embodiments of the method of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0233] In the above embodiments, all or part of the methods can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the methods can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state disk), etc.
[0234] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be instructed by a computer program to complete the relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. The storage medium includes ROM or random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0235] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for identifying a target object, characterized in that, Applied to an electronic device having a camera, the method includes: During the continuous acquisition of frame images by the camera, a first output box predicted for the target object in the current frame image is obtained; When the electronic device and the target object remain relatively stationary within a target time period, an anchor box is displayed in the current frame image to mark the target object; wherein, the anchor box is a second output box predicted for the target object in historical frame images as a reference; the target time period refers to the time period between the first acquisition time of the historical frame image corresponding to the anchor box and the second acquisition time of the current frame image; When the electronic device and the target object move relative to each other within the target time period, the first output box is displayed in the current frame image to mark the target object, and the first output box is used as a new anchor point box.
2. The method according to claim 1, characterized in that, The step of displaying an anchor point box in the current frame image to mark the target object when the electronic device and the target object remain relatively stationary within a target time period includes: If the relative change between the first output frame and the anchor frame does not exceed a preset change threshold, the target object is marked in the current frame image based on the anchor frame.
3. The method according to claim 2, characterized in that, The preset change thresholds include: preset position change thresholds and / or preset size change thresholds; The step of marking the target object in the current frame image based on the anchor box when the relative change between the first output box and the anchor box does not exceed a preset change threshold includes: If the positional change between the first output frame and the anchor frame does not exceed the preset positional change threshold and / or the size change between the first output frame and the anchor frame does not exceed the preset size change threshold, the target object is marked in the current frame image based on the anchor frame.
4. The method according to claim 3, characterized in that, The position change threshold includes a first threshold, and the size change threshold includes a second threshold and / or a third threshold; If the positional change between the first output frame and the anchor frame does not exceed the preset positional change threshold and / or the size change between the first output frame and the anchor frame does not exceed the preset size change threshold, the target object is marked in the current frame image based on the anchor frame, including: If the displacement of the center point between the first output frame and the anchor frame does not exceed the first threshold, and / or the change in the side length between the first output frame and the anchor frame does not exceed the second threshold, and / or the change in the diagonal length between the first output frame and the anchor frame does not exceed the third threshold, the target object is marked in the current frame image based on the anchor frame.
5. The method according to claim 2, characterized in that, The step of marking the target object in the current frame image based on the anchor box when the relative change between the first output box and the anchor box does not exceed a preset change threshold includes: When the electronic device is in a static mode and the relative change does not exceed a preset change threshold, the target object is marked in the current frame image based on the anchor point box; The static mode refers to the mode in which the anchor point box is output and displayed.
6. The method according to claim 5, characterized in that, The method further includes: When the electronic device is not in a static mode, the relative change does not exceed a preset change threshold, and the number of stable frames does not reach a preset frame number threshold, the target object is marked in the current frame image based on the first output box, the anchor point box remains unchanged, and the number of stable frames is incremented by 1. The number of stable frames refers to the number of stable frame images that exist between the historical frame image and the current frame image; the relative change between the output box predicted for the target object and the anchor box in the stable frame image does not exceed the preset change threshold.
7. The method according to claim 5, characterized in that, The method further includes: When the electronic device is not in a static mode, the relative change does not exceed a preset change threshold, and the number of stable frames reaches a preset frame number threshold, the first output box is used as a new anchor box to mark the target object in the current frame image based on the new anchor box, and the static mode is entered.
8. The method according to claim 6 or 7, characterized in that, The method further includes: If the number of times the state of the electronic device changes exceeds a preset threshold within a preset time window, the preset frame number threshold is increased. The number of state changes is used to characterize the number of times the electronic device changes state during the process of processing frame images acquired within the preset time window; the state changes include entering the still mode and exiting the still mode.
9. The method according to claim 1, characterized in that, When the electronic device and the target object move relative to each other within the target time period, displaying the first output box in the current frame image to mark the target object and using the first output box as a new anchor point box includes: If the relative change between the first output box and the anchor box exceeds a preset change threshold, the target object is marked in the current frame image based on the first output box, and the first output box is used as the new anchor box.
10. The method according to claim 9, characterized in that, When the relative change between the first output box and the anchor box exceeds a preset change threshold, marking the target object in the current frame image based on the first output box and using the first output box as a new anchor box includes: When the electronic device is in a static mode but the relative change exceeds a preset change threshold, the electronic device exits the static mode, marks the target object in the current frame image based on the first output box, and uses the first output box as a new anchor point box.
11. A method for identifying a target object, characterized in that, Applied to an electronic device having a camera, the method includes: During the continuous acquisition of frame images by the camera, while the electronic device and the target object remain relatively stationary, a first coordinate frame for marking the target object is fixedly displayed. When the electronic device moves relative to the target object, a second coordinate bounding box predicted for the target object in the current frame image is displayed in real time.
12. An electronic device, characterized in that, The electronic device includes a memory, a camera, and a processor; the memory and the camera are coupled to the processor; wherein the memory stores computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the electronic device to perform the method as described in any one of claims 1 to 10 or 11.
13. A computer storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 10 or 11.
14. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 10 or 11.
Citation Information
Patent Citations
Target detection box output method and device, terminal and storage medium
CN110677585A
Rapid movement detection method and device
CN112669351A
Target tracking method and related equipment
CN113256683A
Adjusting method, device and equipment of viewing frame and storage medium
CN116524168A
Electronic apparatus, image output method, and program
JP2011010007A