Target tracking method and electronic equipment

By displaying multi-frame images in electronic devices and extracting multiple features, using key point segmentation and feature library updates, the low accuracy problems caused by occlusion and pose changes in target tracking are solved, and higher target tracking accuracy is achieved.

CN114359335BActive Publication Date: 2025-08-26HUAWEI TECH CO LTD

Patent Information

Application Number
CN202011066347.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-30
Publication Date
2025-08-26
Estimated Expiration
2040-09-30

AI Technical Summary

Technical Problem

In the prior art, in the process of target tracking, the accuracy of target tracking is not high when facing occlusion, disappearance and reappearance, and deformation.

Method used

By displaying multi-frame images in electronic devices, extracting and matching multiple features of tracking targets and candidate targets, segmenting images with key points, and updating and replacing features in the feature library, ensuring accurate tracking of targets under different poses and partial occlusions.

Benefits of technology

It improves the accuracy of target tracking, especially when the object is partially blocked or posture changes, which enhances the target recognition ability of electronic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359335B_ABST
    Figure CN114359335B_ABST
Patent Text Reader

Abstract

The present application discloses a target tracking method and an electronic device. The method includes: the electronic device displays the Nth frame image and determines the tracking target in the Nth frame image. Then the electronic device divides the tracking target into multiple tracking target images. The electronic device extracts features from each of the multiple tracking target images to obtain multiple tracking target features of the tracking target. Next, the electronic device displays the N+1th frame image. The electronic device detects the candidate target in the N+1th frame image. The electronic device divides the candidate target into multiple candidate target images. The electronic device extracts features from each candidate target image to obtain candidate target features of the candidate target. If the first candidate target feature among the multiple candidate target features matches the first tracking target feature among the multiple tracking target features, the electronic device determines that the candidate target is the tracking target. In this way, the accuracy of target tracking by the electronic device can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of terminals and image processing, and in particular to a target tracking method and electronic equipment. Background Art

[0002] With the continuous development of image processing technology, target tracking is widely used in fields such as intelligent video surveillance, autonomous driving, and unmanned supermarkets. Generally, an electronic device can determine the tracking target in a captured image frame. The electronic device can save the characteristics of the tracking target. The electronic device then performs feature matching between the characteristics of the candidate target in the current image frame and the characteristics of the saved tracking target. If the characteristics of the candidate target match the characteristics of the saved tracking target, the electronic device determines that the candidate target in the current image frame is the tracking target.

[0003] However, if the target being tracked in the current image frame is partially blocked, disappears, reappears, or deformed, the electronic device may not be able to determine the target being tracked in the current image frame. This can easily cause target tracking failures in the electronic device, resulting in low target tracking accuracy.

[0004] Therefore, when the tracked target is partially blocked, disappears and reappears, or deformed, how to improve the accuracy of target tracking is an urgent problem to be solved. Summary of the Invention

[0005] An embodiment of the present application provides a target tracking method and an electronic device, which can improve the accuracy of target tracking by using the method during target tracking.

[0006] In a first aspect, the present application provides a target tracking method, which includes: an electronic device displays a first user interface, wherein the first user interface displays an M-th frame image, and the M-th frame image contains a tracking target; the electronic device obtains multiple tracking target features of the tracking target; the electronic device displays a second user interface, wherein the second user interface displays a K-th frame image, and the K-th frame image contains a first candidate target; the K-th frame image is an image frame after the M-th frame image; the electronic device obtains multiple candidate target features of the first candidate target; when a first candidate target feature among the multiple candidate target features matches a first tracking target feature among the multiple tracking target features, the electronic device determines that the first candidate target is the tracking target.

[0007] In this way, when the tracking target and the candidate target are the same object, but the parts containing the object are not exactly the same, the electronic device can also accurately track the target and confirm that the candidate target is the tracking target. For example, when the tracking target is the whole body of person A, the electronic device can save the features of the image corresponding to the body part above the shoulders of person A, the features of the image corresponding to the body part above the hips of person A, the features of the image corresponding to the body part above the knees of person A, and the features of the image corresponding to the whole body part of person A from head to toe. When the candidate target is person A with only the body part above the shoulders, the electronic device can match the features of the image corresponding to the body part above the shoulders of person A in the two frames of images. Instead of matching the features of the image corresponding to the body part above the hips of person A in the previous frame of image with the features of the image corresponding to the body part above the shoulders of person A in the next frame of image. In this way, the accuracy of target tracking by the electronic device can be improved.

[0008] In conjunction with the first aspect, in one possible implementation, the electronic device obtains multiple tracking target features of a tracking target, specifically including: the electronic device obtains multiple tracking target images based on the tracking target, where the tracking target images include part or all of the tracking target; and the electronic device extracts features from the multiple tracking target images to obtain multiple tracking target features, where the number of the multiple tracking target features is equal to the number of the multiple tracking target images. In this way, the electronic device can store the multiple tracking target features.

[0009] In conjunction with the first aspect, in one possible implementation, the electronic device divides the tracking target into multiple tracking target images. Specifically, the electronic device obtains multiple tracking target images of the tracking target based on key points of the tracking target, where the tracking target images include one or more key points of the tracking target. In this way, the electronic device can obtain multiple tracking target images of the tracking target based on the key points of the tracking target.

[0010] In combination with the first aspect, in a possible implementation, the multiple tracking target images include a first tracking target image and a second tracking target image, wherein: the first tracking target image and the second tracking target image contain the same key points of the tracking target, and the second tracking target image contains more key points of the tracking target than the first tracking target image; or the key points of the tracking target contained in the first tracking target image are different from the key points of the tracking target contained in the second tracking target image.

[0011] In this way, the multiple tracking target images obtained by the electronic device may contain some of the same key points. Or the multiple tracking target images obtained by the electronic device do not contain the same key points. For example, based on the complete person A, the electronic device can obtain an image of the body part above the shoulders of person A, an image of the body part above the hips of person A, an image of the body part above the knees of person A, and an image of the whole body part from head to feet of person A. Or, based on the complete person A, the electronic device can obtain an image containing only the head of person A, an image containing only the upper body of person A (i.e., the body part from the shoulders to the hips, excluding the hips), an image containing only the lower body of person A (i.e., the body part from the hips to the feet, excluding the feet), and an image containing only the feet of person A.

[0012] In conjunction with the first aspect, in one possible implementation, an electronic device obtains multiple candidate target features of a first candidate target, including: the electronic device obtains multiple candidate target images based on the first candidate target; the multiple candidate target images contain part or all of the first candidate target; and the electronic device performs feature extraction on each of the multiple candidate target images to obtain multiple candidate target features, where the number of the multiple candidate target features is equal to the number of the multiple candidate target features. In this way, the electronic device can obtain multiple features corresponding to the first candidate target.

[0013] In conjunction with the first aspect, in one possible implementation, the electronic device divides the first candidate target into multiple candidate target images, specifically including: the electronic device obtains multiple candidate target images based on key points of the first candidate target, where the candidate target images include one or more key points of the first candidate target. In this way, the electronic device can obtain multiple candidate target images based on the key points of the first candidate target.

[0014] In conjunction with the first aspect, in one possible implementation, the multiple candidate target images include a first candidate target image and a second candidate target image, wherein: the first candidate target image and the second candidate target image contain the same key points of the first candidate target, and the second candidate target image contains more key points of the first candidate target than the first candidate target image; or the key points of the first candidate target contained in the first candidate target image are different from the key points of the first candidate target contained in the second candidate target image. In this way, the multiple candidate target images can contain the same key points. Alternatively, the multiple candidate target images do not contain the same key points.

[0015] In conjunction with the first aspect, in one possible implementation, the multiple tracking target images include a first tracking target image and a second tracking target image, and the multiple tracking target images are obtained from the tracking target; the multiple candidate target images include a first candidate target image and a second candidate target image, and the multiple candidate target images are obtained from the first candidate target; the number of key points of the tracking target contained in the first tracking target image is the same as the number of key points of the first candidate target contained in the first candidate target image; the number of key points of the tracking target contained in the second tracking target image is the same as the number of key points of the first candidate target contained in the second candidate target image; the number of key points of the tracking target contained in the first tracking target image is greater than the number of key points of the tracking target contained in the second tracking target image; the first tracking target feature is extracted from the first tracking target image, and the first candidate target feature is extracted from the first candidate target image. In this way, the electronic device can more accurately determine that the first candidate target feature matches the first tracking target feature. Consequently, the electronic device can more accurately determine that the first candidate target is the tracking target.

[0016] In combination with the first aspect, in a possible implementation, when a first candidate target feature among multiple candidate target features matches a first tracking target feature among multiple tracking target features, after the electronic device determines that the first candidate target is the tracking target, the method also includes: the electronic device saves the second candidate target to a feature library that stores multiple tracking target features; the second candidate feature is extracted by the electronic device from a third candidate target image among the multiple candidate target images; the number of key points of the first candidate target contained in the third candidate target image is greater than the number of key points of the tracking target contained in the multiple tracking target images.

[0017] The electronic device can add the features corresponding to the first candidate target to the feature library of the tracking target. In this way, the features corresponding to the tracking target are increased. For example, suppose the tracking target is person A with only the upper body appearing. The first candidate target is person A with the whole body, and the electronic device can add the features corresponding to the lower body of person A to the feature library of the tracking target. When only the lower body of person A appears in subsequent image frames (for example, the upper body of person A is blocked), the electronic device can also match only the features of the image corresponding to the lower body of person A. The electronic device can also accurately track the target. In this way, the accuracy of the electronic device in target tracking in subsequent image frames is improved.

[0018] In conjunction with the first aspect, in one possible implementation, when a first candidate target feature among multiple candidate target features matches a first tracking target feature among multiple tracking target features, after the electronic device determines that the first candidate target is the tracking target, the method further includes: if the difference between M and K is equal to a preset threshold, the electronic device saves the first candidate target feature to a feature library that stores multiple tracking target features. That is, after every image frame with a preset threshold, the electronic device updates the saved tracking target feature. In this way, the electronic device can more accurately track the target in subsequent image frames.

[0019] In conjunction with the first aspect, in one possible implementation, the electronic device saves the first candidate target feature in a feature library storing multiple tracking target features. Specifically, the electronic device replaces the first tracking target feature in the feature library storing the multiple tracking target features with the first candidate target feature. The electronic device may update the saved tracking target feature. In this way, the electronic device can more accurately track the target in subsequent image frames.

[0020] In conjunction with the first aspect, in one possible implementation, the method further includes: the electronic device detecting a tracking target in the Mth image frame; the electronic device displaying a detection frame in a first user interface, the detection frame being used to enclose the tracking target; the electronic device receiving a first user operation, the first operation being used to select the tracking target in the first user interface; and the electronic device determining the tracking target in response to the first user operation. In this way, the electronic device can accurately determine the tracking target specified by the user.

[0021] In conjunction with the first aspect, in one possible implementation, the method further includes: the electronic device detecting one or more candidate targets in the Kth image frame, the one or more candidate targets including the first candidate target; and the one or more candidate targets having the same attributes as the tracked target. This ensures that the candidate targets detected by the electronic device and the tracked target are of the same species.

[0022] In a second aspect, an embodiment of the present application provides a target tracking method, which includes: an electronic device displays a first user interface, wherein the Nth frame image is displayed in the first user interface, and the Nth frame image contains a tracking target; the electronic device determines the tracking target and the first pose of the tracking target in the first user interface; the electronic device extracts features of the tracking target, obtains and saves tracking target features corresponding to the tracking target; the electronic device displays a second user interface, wherein the N+1th frame image is displayed in the second user interface, and the N+1th frame image contains one or more candidate targets; the electronic device determines the second pose of one or more candidate targets and extracts features of the candidate targets to obtain candidate target features; the electronic device determines that the first candidate target among the one or more candidate targets is the tracking target, and if the first pose and the second pose are different, the electronic device saves the candidate target features to a feature library corresponding to the tracking target, and the feature library saves the tracking target features.

[0023] By implementing the embodiments of the present application, the electronic device can track a specified object (such as a person), and the tracking target may be deformed. For example, when the tracking target is a person, the person can change from a squatting position to a sitting position or a standing position. That is, the position of the tracking target will change in consecutive image frames. Because the electronic device can save the corresponding features of the tracking target in different positions into the feature library corresponding to the tracking target. When the position of the same person in the Nth frame image and the N+1th frame image changes, the electronic device can also accurately detect the tracking target specified by the user in the Nth frame image in the N+1th frame image. In this way, the accuracy of the electronic device in target tracking is improved.

[0024] In conjunction with the second aspect, in one possible implementation, if the first pose of the tracking target and the second pose of the first candidate target are the same, the electronic device may perform feature matching between the tracking target features and the candidate target features. If the candidate target features match the tracking target features, the electronic device determines that the first candidate target is the tracking target.

[0025] In conjunction with the second aspect, in one possible implementation, if there are multiple candidate targets in the (N+1)th frame image, the electronic device may obtain the position of the tracking target in the (N)th frame image, for example, the center of the tracking target is at the first position in the (N)th frame image. The electronic device may also obtain the position of the first candidate target in the (N+1)th frame image, for example, the center of the first candidate target is at the second position in the (N+1)th frame image. If a preset distance between the first position and the second position is less than a preset distance, the electronic device may determine that the first candidate target is the tracking target.

[0026] In combination with the second aspect, in one possible implementation, if the electronic device saves the features of the tracking target in different postures, the electronic device can perform feature matching on the feature vectors in the saved tracking target features that have the same posture as the second posture of the first candidate target, and the candidate target features. If they match, the electronic device determines that the first candidate target is the tracking target.

[0027] In a third aspect, an electronic device is provided. The electronic device may include: a display screen, a processor, and a memory; the memory is coupled to the processor; and the display screen is coupled to the processor, wherein:

[0028] The display screen is used to display a first user interface, wherein the first user interface displays an M-th frame image, wherein the M-th frame image contains the tracking target; and display a second user interface, wherein the second user interface displays a K-th frame image, wherein the K-th frame image contains the first candidate target;

[0029] The processor is configured to obtain a plurality of tracking target features of a tracking target; obtain a plurality of candidate target features of a first candidate target; and determine that the first candidate target is a tracking target when a first candidate target feature among the plurality of candidate target features matches a first tracking target feature among the plurality of tracking target features;

[0030] The memory is used to store the multiple tracking target features.

[0031] In this way, when the tracking target and the candidate target are the same object, but the parts containing the object are not exactly the same, the electronic device can also accurately track the target and confirm that the candidate target is the tracking target. For example, when the tracking target is the whole body of person A, the electronic device can save the features of the image corresponding to the body part above the shoulders of person A, the features of the image corresponding to the body part above the hips of person A, the features of the image corresponding to the body part above the knees of person A, and the features of the image corresponding to the whole body part of person A from head to toe. When the candidate target is person A with only the body part above the shoulders, the electronic device can match the features of the image corresponding to the body part above the shoulders of person A in the two frames of images. Instead of matching the features of the image corresponding to the body part above the hips of person A in the previous frame of image with the features of the image corresponding to the body part above the shoulders of person A in the next frame of image. In this way, the accuracy of target tracking by the electronic device can be improved.

[0032] In conjunction with the third aspect, in one possible implementation, the processor is specifically configured to: obtain multiple tracking target images based on the tracking target, where the tracking target images contain part or all of the tracking target; and perform feature extraction on the multiple tracking target images to obtain multiple tracking target features, where the number of the multiple tracking target features is equal to the number of the multiple tracking target images. In this way, the electronic device can store multiple features of the tracking target.

[0033] In conjunction with the third aspect, in one possible implementation, the processor is specifically configured to obtain multiple tracking target images based on the key points of the tracking target, where the tracking target images contain one or more key points of the tracking target. In this way, the electronic device can obtain multiple candidate target images based on the key points of the candidate targets.

[0034] In combination with the third aspect, in one possible implementation, the multiple tracking target images include a first tracking target image and a second tracking target image, wherein: the first tracking target image and the second tracking target image contain the same key points of the tracking target, and the second tracking target image contains more key points of the tracking target than the first tracking target image; or the key points of the tracking target contained in the first tracking target image are different from the key points of the tracking target contained in the second tracking target image.

[0035] In this way, the multiple tracking target images obtained by the electronic device may contain some of the same key points. Or the multiple tracking target images obtained by the electronic device do not contain the same key points. For example, based on the complete person A, the electronic device can obtain an image of the body part above the shoulders of person A, an image of the body part above the hips of person A, an image of the body part above the knees of person A, and an image of the whole body part from head to feet of person A. Or, based on the complete person A, the electronic device can obtain an image containing only the head of person A, an image containing only the upper body of person A (i.e., the body part from the shoulders to the hips, excluding the hips), an image containing only the lower body of person A (i.e., the body part from the hips to the feet, excluding the feet), and an image containing only the feet of person A.

[0036] In conjunction with the third aspect, in one possible implementation, the processor is configured to: obtain multiple candidate target images based on a first candidate target; the multiple candidate target images include part or all of the first candidate target; and perform feature extraction on each of the multiple candidate target images to obtain multiple candidate target features, where the number of the multiple candidate target features is equal to the number of the multiple candidate target features. In this way, the electronic device can obtain multiple features corresponding to the first candidate target.

[0037] In conjunction with the third aspect, in one possible implementation, the processor is specifically configured to obtain multiple candidate target images based on key points of the first candidate target, where the candidate target images include one or more key points of the first candidate target. In this way, the electronic device can obtain multiple candidate target images based on the key points of the first candidate target.

[0038] In conjunction with the third aspect, in one possible implementation, the multiple candidate target images include a first candidate target image and a second candidate target image, wherein: the first candidate target image and the second candidate target image contain the same key points of the first candidate target, and the second candidate target image contains more key points of the first candidate target than the first candidate target image; or the key points of the first candidate target contained in the first candidate target image are different from the key points of the first candidate target contained in the second candidate target image. In this way, the multiple candidate target images can contain the same key points. Alternatively, the multiple candidate target images do not contain the same key points.

[0039] In conjunction with the third aspect, in one possible implementation, the multiple tracking target images include a first tracking target image and a second tracking target image, and the multiple tracking target images are obtained from the tracking target; the multiple candidate target images include a first candidate target image and a second candidate target image, and the multiple candidate target images are obtained from the first candidate target; the number of key points of the tracking target contained in the first tracking target image is the same as the number of key points of the first candidate target contained in the first candidate target image; the number of key points of the tracking target contained in the second tracking target image is the same as the number of key points of the first candidate target contained in the second candidate target image; the number of key points of the tracking target contained in the first tracking target image is greater than the number of key points of the tracking target contained in the second tracking target image; the first tracking target feature is extracted from the first tracking target image, and the first candidate target feature is extracted from the first candidate target image. In this way, the electronic device can more accurately determine that the first candidate target feature matches the first tracking target feature. Consequently, the electronic device can more accurately determine that the first candidate target is the tracking target.

[0040] In combination with the third aspect, in one possible implementation, the memory is used to: save the second candidate target into a feature library that stores multiple tracking target features; the second candidate feature is extracted by the processor from a third candidate target image among multiple candidate target images; the number of key points of the first candidate target contained in the third candidate target image is greater than the number of key points of the tracking target contained in the multiple tracking target images.

[0041] The electronic device can add the features corresponding to the first candidate target to the feature library of the tracking target. In this way, the features corresponding to the tracking target are increased. For example, suppose the tracking target is person A with only the upper body appearing. The first candidate target is person A with the whole body, and the electronic device can add the features corresponding to the lower body of person A to the feature library of the tracking target. When only the lower body of person A appears in subsequent image frames (for example, the upper body of person A is blocked), the electronic device can also match only the features of the image corresponding to the lower body of person A. The electronic device can also accurately track the target. In this way, the accuracy of the electronic device in target tracking in subsequent image frames is improved.

[0042] In conjunction with the third aspect, in one possible implementation, the memory is configured to: if the difference between M and K is equal to a preset threshold, save the first candidate target feature to a feature library storing multiple tracking target features. That is, the electronic device updates the stored tracking target features after every image frame that exceeds the preset threshold. This allows the electronic device to more accurately track the target in subsequent image frames.

[0043] In conjunction with the third aspect, in one possible implementation, the memory is configured to replace a first tracking target feature in a feature library storing multiple tracking target features with a first candidate target feature. The electronic device may update the stored tracking target feature. This allows the electronic device to more accurately track the target in subsequent image frames.

[0044] In a fourth aspect, an electronic device is provided, comprising one or more touch screens, one or more storage modules, and one or more processing modules; wherein the one or more storage modules store one or more programs; when the one or more processing modules execute the one or more programs, the electronic device implements the method described in any possible implementation method of the first aspect or the second aspect.

[0045] In a fifth aspect, an electronic device is provided, comprising one or more touch screens, one or more memories, and one or more processors; wherein the one or more memories store one or more programs; and when the one or more processors execute the one or more programs, the electronic device implements the method described in any possible implementation method of the first aspect or the second aspect.

[0046] In a sixth aspect, a computer-readable storage medium is provided, comprising instructions, characterized in that when the above instructions are run on an electronic device, the electronic device executes any possible implementation method in the first aspect or the second aspect.

[0047] In a seventh aspect, a computer product is provided. When the computer program product is run on a computer, the computer is caused to execute any possible implementation method in the first aspect or the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a schematic diagram of a user interface of a tablet computer 10 provided in an embodiment of the present application;

[0049] Figure 2 is a schematic diagram of a user interface of a tablet computer 10 provided in an embodiment of the present application;

[0050] Figure 3 is a schematic diagram of a user interface of a tablet computer 10 provided in an embodiment of the present application;

[0051] Figure 4 is a schematic diagram of a user interface of a tablet computer 10 provided in an embodiment of the present application;

[0052] Figure 5A-5B is a schematic diagram of a user interface of a tablet computer 10 provided in an embodiment of the present application;

[0053] Figure 6 is a schematic diagram of a user interface of a tablet computer 10 provided in an embodiment of the present application;

[0054] Figure 7 is a schematic diagram of a user interface of a tablet computer 10 provided in an embodiment of the present application;

[0055] Figure 8 is a schematic diagram of a user interface of a tablet computer 10 provided in an embodiment of the present application;

[0056] Figure 9 is a schematic diagram of a user interface of a tablet computer 10 provided in an embodiment of the present application;

[0057] Figure 10 is a schematic diagram of a user interface of a tablet computer 10 provided in an embodiment of the present application;

[0058] Figure 11 is a schematic diagram of a tracking target and corresponding features of the tracking target provided in an embodiment of the present application;

[0059] Figure 12 is a schematic diagram of a user interface when the tablet computer 10 provided by an embodiment of the present application fails to detect a tracking target;

[0060] Figure 13 Schematic diagram of person A and corresponding features of person A in the N+1th frame image provided in an embodiment of the present application;

[0061] Figure 14Schematic diagram of person A and corresponding features of person A in the N+1th frame image provided in an embodiment of the present application;

[0062] Figure 15 A schematic diagram of key points of the human body provided in the embodiments of this application;

[0063] Figure 16 Schematic diagram of a tracking target, an input image corresponding to the tracking target, and features corresponding to the input image provided in an embodiment of the present application;

[0064] Figure 17 A schematic diagram of a user interface of the tablet computer 10 provided in an embodiment of the present application detecting a tracking target;

[0065] Figure 18 Schematic diagram of the candidate target person A, the input image corresponding to the candidate target person A, and the corresponding features of the input image provided in an embodiment of the present application;

[0066] Figure 19 Schematic diagram of the candidate target person B, the input image corresponding to the candidate target person B, and the features corresponding to the input image provided in an embodiment of the present application;

[0067] Figure 20 A flowchart of a target tracking method provided in an embodiment of the present application;

[0068] Figure 21 A schematic diagram of a user interface for selecting a tracking target provided in an embodiment of the present application;

[0069] Figure 22 Schematic diagram of a tracking target, an input image corresponding to the tracking target, and features corresponding to the input image provided in an embodiment of the present application;

[0070] Figure 23 A flowchart of a target tracking method provided in an embodiment of the present application;

[0071] Figure 24 This is a flowchart of images of person A in different postures and corresponding features of person A in different postures according to an embodiment of the present application;

[0072] Figure 25 A schematic diagram of the tracking target and corresponding features of the tracking target provided in an embodiment of the present application;

[0073] Figure 26 A schematic diagram of candidate targets and corresponding features provided in an embodiment of the present application;

[0074] Figure 27 A schematic diagram of the tracking target and corresponding features of the tracking target provided in an embodiment of the present application;

[0075] Figure 28 A schematic diagram of the tracking target and corresponding features of the tracking target provided in an embodiment of the present application;

[0076] Figure 29 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0077] Figure 30 A schematic diagram of the software framework of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0078] The terms used in the following examples of the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular expressions "a," "an," "said," "above," "the," and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and encompasses any or all possible combinations of one or more of the listed items.

[0079] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0080] The term "user interface (UI)" in the specification, claims and drawings of this application refers to the media interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The user interface of an application is a source code written in a specific computer language such as Java and Extensible Markup Language (XML). The interface source code is parsed and rendered on the terminal device, and finally presented as content that the user can recognize, such as images, text, buttons and other controls. Controls, also known as widgets, are the basic elements of the user interface. Typical controls include toolbars, menu bars, input boxes, buttons, scroll bars, images and text. The properties and contents of controls in the interface are defined by tags or nodes, such as XML through <textview> 、 <imgview> 、 <videoview>The controls contained in the interface are specified by nodes such as <head> and <body>. A node corresponds to a control or attribute in the interface, and the node is presented as user-visible content after parsing and rendering. In addition, many applications, such as hybrid applications, usually also contain web pages in their interfaces. A web page, also called a page, can be understood as a special control embedded in the application interface. A web page is a source code written in a specific computer language, such as hypertext markup language (HTML), cascading style sheets (CSS), JavaScript (JS), etc. The web page source code can be loaded and displayed as user-recognizable content by a browser or a web page display component with similar functions to a browser. The specific content contained in a web page is also defined by tags or nodes in the web page source code, such as HTML through <body>. 、 、 <video> 、 <canvas>To define the elements and attributes of a web page.

[0081] A common form of user interface is a graphical user interface (GUI), which refers to a user interface related to computer operations that uses graphics to display. It can be a window, control, or other interface element displayed on the display of an electronic device.

[0082] To facilitate understanding, the following first introduces the relevant terms and concepts involved in the embodiments of this application.

[0083] (1) Target tracking

[0084] A display screen of an electronic device displays continuous image frames. The image frames may be acquired by a camera in the electronic device. Or the image frames may be sent to the electronic device by a video surveillance device. The user may select a tracking target in the Nth frame of the continuous image frames displayed by the electronic device. Target tracking in an embodiment of the present application means that the electronic device may determine the tracking target (person or object) selected by the user in the image frame after the Nth frame (for example, the N+1th frame) of the continuous image frames. The electronic device may mark the tracking target in the image frame.

[0085] Here, an electronic device is a device that can display continuous image frames, such as a mobile phone, tablet computer, computer, television, etc. This embodiment of the application does not limit the electronic device. A video surveillance device is a device that can capture images, such as a webcam, infrared camera, etc. This embodiment of the application does not limit the video surveillance device.

[0086] In an embodiment of the present application, a frame of image displayed by an electronic device may be referred to as an image frame or an Nth frame of image.

[0087] (2) Tracking target

[0088] In the embodiment of the present application, the object (eg, a person, a plant, or a car, etc.) specified by the user for tracking detection in the Nth frame image is referred to as a tracking target.

[0089] (3) Candidate targets

[0090] In the embodiment of the present application, objects that appear in the image frames subsequent to the Nth frame and belong to the same category as the tracking target are referred to as candidate targets. The candidate targets have the same attributes as the tracking target. Here, the attributes may refer to the type of object, the color of the object, the gender of the object, the height of the object, etc., which are not limited here. For example, if the tracking target is a person, then the people that appear in the image frames subsequent to the Nth frame are all candidate targets. If the tracking target is a car, then the cars that appear in the subsequent image frames or video frames in the Nth frame are all candidate targets. It is understandable that the electronic device can detect that the tracking target is a person. The electronic device can then detect people in the images subsequent to the Nth frame, and the detected people are all candidate objects. How the electronic device detects people will be introduced below and will not be elaborated here.

[0091] The following briefly describes the specific process of target tracking using a tablet as an example. Figures 1-14 , Figures 1-14 The process of target tracking by the tablet computer 10 is exemplarily shown.

[0092] Figure 1 The user interface 100 of the tablet computer 10 is shown as an example. The user interface 100 may include icons of some applications. For example, a settings icon 101, a mall icon 102, a memo icon 103, a camera icon 104, a file management icon 105, an email icon 106, a music icon 107, and a video surveillance icon 108. In some embodiments, the user interface 100 may include icons of more or fewer applications. In some embodiments, the user interface 100 may include icons of some applications related to the user interface 101. Figure 1 The application shown is an icon of different application programs, such as a video icon, an instant messaging application icon, etc. This is not limited here.

[0093] like Figure 2 The user interface 100 of the tablet computer 10 is shown. The user can click on the video monitoring icon 108. In response to the user operation, the tablet computer can display the user interface 200.

[0094] Figure 3 The user interface 200 of the tablet computer 10 is shown as an example. A frame of image is currently displayed in the user interface 200. The image frame displayed in the user interface 200 may include person A and person B. This frame of image may be captured by a camera in the tablet computer 10. This frame of image may also be sent to the tablet computer 10 by a video surveillance device.

[0095] It is understandable that if the image frames displayed on the user interface 200 of the tablet computer 10 are acquired by the video surveillance device and sent to the tablet computer 10, then before the tablet computer 10 can display the image frames acquired by the video surveillance device, the user can establish a communication connection between the tablet computer 10 and the video surveillance device. In one possible implementation, the user can establish a communication connection between the tablet computer 10 and the video surveillance device in the video surveillance application. Optionally, the tablet computer 10 can establish a communication connection with the video surveillance device via Bluetooth. Optionally, the tablet computer 10 can also establish a communication connection with the video surveillance device via a wireless local area network (WLAN). In this way, the tablet computer 10 can display the image frames acquired by the video surveillance device in real time. It is understandable that the manner in which the tablet computer 10 establishes a communication connection with the video surveillance device is not limited.

[0096] like Figure 4 As shown in the user interface 200 of the tablet computer 10, the user can swipe downward from the top of the user interface 200. In response to the user operation, the tablet computer 10 can display the user interface 300.

[0097] Figure 5A The user interface 300 of the tablet computer 10 is shown as an example. The user interface 300 may include an image frame (the image frame may include character A and character B) and a status bar 301. The status bar 301 may include controls 302. The status bar 301 may include more controls, such as a pause control (not shown) for controlling the pause playback of the image frame, a brightness adjustment control (not shown) for adjusting the brightness of the image frame, and so on. The embodiment of the present application does not limit the number of controls in the status bar 301, the specific controls, and the position of the status bar 301 in the user interface 300. The user can use the controls 302 to turn on or off the target tracking function of the tablet computer 10.

[0098] It is understood that the user interface of the tablet computer 10 for opening or closing the tablet computer 10 may be in various forms, not limited to Figure 5A The user interface 300 is shown. For example, Figure 5B The user interface 300B of the tablet computer is shown. The user interface 300B may include an image frame (the image frame may include person A and person B) and a control 303. The control 303 is used to turn on or off the target tracking function of the tablet computer 10. The position of the control 303 in the user interface 300B is not limited. Figure 6 The user interface 300 of the tablet computer 10 is shown. The user can click on the control 302 to start the target tracking function of the tablet computer 10. Figure 7 A user interface 300 of the tablet computer 10 is shown. Figure 7 The state of the control 302 in the user interface 300 indicates that the target tracking function of the tablet computer 10 has been turned on. In response to the user's operation of turning on the target tracking function of the tablet computer 10, the tablet computer 10 displays the user interface 400.

[0099] Figure 8 The user interface 400 of the tablet computer 10 is shown as an example. The user interface 400 displays the Nth frame image acquired by the camera or video surveillance device of the tablet computer 10 at time T0. The Nth frame image may include person A and person B. The user interface 400 may also include text 401. The text 401 is used to prompt the user to select the tracking target in the Nth frame image. The specific text content of the text 401 may be "Please select the tracking target" or other, which is not limited here. The tablet computer 10 can detect person A and person B in the Nth frame image. The tablet computer 10 displays a detection box 402 and a detection box 403 in the user interface 400. The detection box 402 is used to circle person A, indicating that the tablet computer 10 has detected person A in the user interface 400. The detection box 403 is used to circle person B, indicating that the tablet computer 10 has detected person B in the user interface 400. As shown in FIG. Figure 8 As shown, the tablet computer 10 can display a continuous sequence of image frames from time T0 to time Tn. That is, the tablet computer displays the Nth frame image at time T0, the N+1th frame image at time T1, the N+2th frame image at time T2, and the N+nth frame image at time Tn.

[0100] In some embodiments, the tablet computer 10 may not display the detection frame 402 and the detection frame 403 .

[0101] like Figure 9 The user interface 400 of the tablet computer 10 is shown. The user can select the person A in the Nth frame image as the tracking target. Figure 9 As shown, the user can select person A as a tracking target by clicking on person A. In response to the user operation, the tablet computer 10 can display a user interface 500 .

[0102] In some embodiments, when the tablet computer 10 does not display the detection frame 402 and the detection frame 403, the user draws a frame to enclose the tracking target. The tablet computer 10 uses the image enclosed by the user as the tracking target.

[0103] like Figure 10 User interface 500 of tablet computer 10 is shown. User interface 500 displays the Nth image frame. The Nth image frame may include person A and person B. User interface 500 may include an indicator frame 501 and a detection frame 502. Indicator frame 501 is used to indicate that the circled object (e.g., person A) is a tracking target. Detection frame 502 is used to circle person B, indicating that tablet computer 10 has detected person B in user interface 500. In response to the user selecting a tracking target, tablet computer 10 may extract features of the user-selected tracking target. Tablet computer 10 may also save the features of the tracking target.

[0104] Figure 11 It is shown as an example Figure 10 The tracking target and the corresponding features of the tracking target. Figure 11 As shown, in response to the user's operation of selecting a tracking target, the tablet computer 10 extracts features of person A and saves the features of person A. The tracking target, that is, the features of person A in the Nth frame image, can be represented by a feature vector Fa(N). For example, Fa(N) = [x1, x2, x3, ..., xn]. The feature vector Fa(N) can represent the color features, texture features, etc. of person A. The specific form and size of the feature vector Fa(N) of person A are not limited here. For example, Fa(N) can be a feature vector containing n numerical values ​​[0.5, 0.6, 0.8, ..., 0.9, 0.7, 0.3]. Wherein, n is an integer, which can be 128, 256, 512, etc., and the size of n is not limited. The tablet computer 10 can store the feature vector Fa(N) in the internal memory.

[0105] After the user specifies a tracking target, the tablet computer 10 tracks the tracking target in image frames subsequent to the Nth image frame.

[0106] Figure 12 The user interface 600 of the tablet computer 10 is shown as an example. The N+1 frame image is displayed in the user interface 600. The N+1 frame image may be the next frame image of the image frame displayed in the user interface 500. The N+1 frame image may include person A and person B. The user interface 600 may also include text 601, a detection frame 602, and a detection frame 603. The text 601 is used to prompt the user whether the tracking target is detected. Here, the features of person A extracted from the N frame image by the tablet computer 10 do not match the features of person A in the N+1 frame image. Please refer to the following for details. Figure 13 and Figure 14 The description in 601 will not be repeated here. The tablet computer 10 can prompt the user in the text 601 that the tracking target was not detected in the N+1 frame image. The specific content of the text 601 can be "No tracking target was detected". The specific content of the text 601 can also be other, such as "Target tracking failed", "Target tracking successful, etc.", and the specific content of the text 601 is not limited here. The detection box 602 is used to circle the person A, which indicates that the tablet computer 10 has detected the person A in the user interface 600. The detection box 603 is used to circle the person B, which indicates that the tablet computer 10 has detected the person B in the user interface 600.

[0107] Tablet computer 10 may extract features of person A and person B detected in user interface 600. Tablet computer 10 may then match the extracted features of person A and person B with the features of a stored tracking target. If the features of person A and person B match the features of a stored tracking target, tablet computer 10 may notify the user that tracking was successful and indicate the tracking target. If the features of person A or person B do not match the features of a stored tracking target, tablet computer 10 may notify the user that the tracking target was not detected or that tracking failed.

[0108] Figure 13 The user interface 600 shows a character A and features corresponding to the character A. In the user interface 600, only the upper body of the character A is shown. The tablet computer 10 can extract features of the character A to obtain features of the character A. Figure 13 The candidate target tracking target features shown in FIG can be represented by a feature vector Fa(N+1). The feature vector Fa(N+1) can include texture features, color features, and other features of person A, which are not limited here.

[0109] Figure 14 The user interface 600 shows a character B and features corresponding to the character B. The tablet computer 10 can extract features of the character B to obtain features of the character B. Figure 14 The candidate target features shown in FIG) can be represented by a feature vector Fb(N+1). The feature vector Fb(N+1) can include texture features, color features, and other features of person B, which are not limited here.

[0110] Tablet computer 10 can perform feature matching of the saved tracking target's features Fa(N) with features Fa(N+1) of person A and features Fb(N+1) of person B in user interface 600. If features Fa(N+1) match features Fa(N), tablet computer 10 successfully detects the tracking target. Otherwise, tablet computer 10 displays a message indicating that the tracking target was not detected. Person A (i.e., the tracking target) in user interface 400 and person A in user interface 600 are actually the same person. However, because person A in user interface 400 appears in full body, tablet computer 10 can extract and save full-body features of person A. However, because person A in user interface 600 appears only half-body, tablet computer 10 can only extract features of half-body. Therefore, when tablet computer 10 performs feature matching, the features of person A in user interface 400 and user interface 600 do not match. Consequently, tablet computer 10 cannot accurately track the tracking target when the tracking target is obscured (e.g., half of the body is obscured).

[0111] In addition, when the tracking target is deformed (different from the tracking target's posture), the tablet computer 10 only extracts features of the tracking target specified by the user in the Nth frame image and saves the extracted features. If the tracking target specified in the Nth frame image is complete (for example, standing person A), the tracking target in the N+1th frame image is deformed (for example, squatting person A). The tablet computer 10 only saves the features of the standing person A as the features of the tracking target. The features extracted by the tablet computer 10 for the standing person A and the squatting person A are different. The tablet computer 10 cannot correctly determine that the standing person A and the squatting person A are the same person. In this way, the tablet computer 10 is prone to target tracking failure, resulting in low target tracking accuracy.

[0112] When the tracking target disappears and reappears, for example, the full body of person A appears in the first frame image, person A disappears in the second frame image, and half of person A appears in the third frame image, the tablet computer 10 is prone to target tracking failure, resulting in low target tracking accuracy.

[0113] In order to improve the accuracy of target tracking by an electronic device, an embodiment of the present application proposes a target tracking method. The method includes: the electronic device displays the Nth frame image and determines the tracking target in the Nth frame image. Then the electronic device divides the tracking target into multiple tracking target images based on a first preset rule. The electronic device extracts features from each of the multiple tracking target images to obtain multiple tracking target features of the tracking target. The electronic device saves the multiple tracking target features of the tracking target. Then, the electronic device displays the N+1th frame image. The electronic device detects the candidate target in the N+1th frame image. The electronic device divides the candidate target into multiple candidate target images based on the first preset rule. The electronic device extracts features from each candidate target image to obtain candidate target features of the candidate target. The electronic device selects a first candidate target feature from the multiple candidate target features and performs feature matching with a first tracking target from the multiple tracking target features. If the first candidate target feature matches the first tracking target feature, the electronic device determines that the candidate target is the tracking target.

[0114] In the embodiment of the present application, the Kth frame image is an image frame subsequent to the Mth frame image. The Mth frame image may be the Nth frame image in the embodiment of the present application, and the Kth frame image may be the N+1th frame image, the N+2th frame image, or the N+nth frame image in the embodiment of the present application.

[0115] The following takes a tablet computer as an example to introduce a specific process of target tracking performed by the tablet computer according to a target tracking method provided in an embodiment of the present application.

[0116] 1. Enable target tracking.

[0117] When the tablet computer is tracking a target, it is necessary to first enable the target tracking function. The specific process of enabling the tablet computer can be referred to the above description. Figures 1-6 The description is not repeated here.

[0118] 2. Determine the tracking target in the Nth frame image.

[0119] The process of the tablet computer 10 determining the tracking target can be specifically referred to the above Figure 8-Figure 9 The description of , will not be repeated here. Figure 9 As shown, the user has selected person A in the Nth frame image as the tracking target. In response to the user operation, the tablet computer 10 determines that person A in the Nth frame image is the tracking target.

[0120] 3. Extract and save tracking target features.

[0121] The tablet computer 10 can divide the person A in the user interface 500 into multiple images according to certain preset rules (such as key points of the human body). The tablet computer 10 extracts features from the multiple images respectively, obtains features corresponding to each image, and saves them.

[0122] Figure 15 The schematic diagram of the key points of the human body is shown as an example. Figure 15 As shown, the key points of the human body of person A may include 14 key points, namely key point A1 to key point A14. Key point A1 to key point A14 may be basic skeletal points of the human body. For example, key point A1 may be the skeletal point of the head of the human body. Key point A2 may be the skeletal point of the neck of the human body. Key point A3 may be the skeletal point of the left shoulder of the human body. Key point A4 may be the skeletal point of the right shoulder of the human body. Key point A5 may be the skeletal point of the left elbow of the human body. Key point A6 may be the skeletal point of the right elbow of the human body. Key point A7 may be the skeletal point of the left hand of the human body. Key point A8 may be the skeletal point of the right hand of the human body. Key point A9 may be the skeletal point of the left hip of the human body. Key point A10 may be the skeletal point of the right hip of the human body. Key point A11 may be the skeletal point of the left knee of the human body. Key point A12 may be the skeletal point of the right knee of the human body. Key point A13 may be the skeletal point of the left foot of the human body. Key point A14 may be the skeletal point of the right foot of the human body.

[0123] The tablet computer 10 can identify key points of the human body. The following content will introduce how the tablet computer identifies key points of the human body, so I will not go into details here.

[0124] The tablet computer 10 can divide the image of the tracking target specified by the user into multiple input images for feature extraction according to the key points of the human body. The tablet computer 10 can extract features from each input image to obtain features of the input image and save the features of the input image.

[0125] Figure 16 The following example shows the input image of the feature extraction corresponding to the tracking target and the features corresponding to the input image. Figure 16 As shown, person A (i.e., the tracking target) in user interface 500 can be divided into four input images for feature extraction based on key points. That is, the tracking target can be divided into input image 1601, input image 1602, input image 1603, and input image 1604 based on key points. Specifically, the body portion of the tracking target containing key points A1-A4 (i.e., the body portion of the tracking target above the shoulders) can be divided into one image for feature extraction (i.e., input image 1601). The body portion of the tracking target containing key points A1-A10 (i.e., the body portion of the tracking target above the hips) can be divided into one image for feature extraction (i.e., input image 1602). The body portion of the tracking target containing key points A1-A12 (i.e., the body portion of the tracking target above the knees) can be divided into one image for feature extraction (i.e., input image 1603). The body portion of the tracking target containing key points A1-A14 (i.e., the entire body portion of the tracking target from head to toe) can be divided into one image for feature extraction (i.e., input image 1604).

[0126] Input images 1601 through 1604 contain a gradually increasing number of keypoints. That is, input image 1602 contains more keypoints than input image 1601. Input image 1602 may contain keypoints from input image 1601. That is, input image 1602 and input image 1601 share the same body part (i.e., the body part above the shoulders) as the tracking target. Input image 1603 contains more keypoints than input image 1602. Input image 1603 may contain keypoints from input image 1602. That is, input image 1603 and input image 1602 share the same body part (i.e., the body part above the hips) as the tracking target. Input image 1604 contains more keypoints than input image 1603. Input image 1604 may contain keypoints from input image 1603. That is, input image 1604 and input image 1603 share the same body part (i.e., the body part above the knees) as the tracking target.

[0127] The tablet computer 10 performs feature extraction on the input image 1601 to obtain features of the input image 1601. For example, the features of the input image 1601 can be represented by a feature vector F1. F1 can represent one or more of the color features, texture features, transformation features, and other features of the input image 1601, without limitation. The color features, texture features, transformation features, and other features of the input image 1601 can be specifically described in the prior art and will not be elaborated upon here. The tablet computer 10 performs feature extraction on the input image 1602 to obtain features of the input image 1602. The features of the input image 1602 can be represented by a feature vector F2. F2 can represent one or more of the color features, texture features, transformation features, and other features of the input image 1602, without limitation. The tablet computer 10 performs feature extraction on the input image 1603 to obtain features of the input image 1603. The features of the input image 1603 can be represented by a feature vector F3. F3 can represent one or more of the color features, texture features, transformation features, and other features of the input image 1603, without limitation. Tablet computer 10 performs feature extraction on input image 1604 to obtain features of input image 1604. The features of input image 1604 can be represented by feature vector F4. F4 can represent one or more features of input image 1604, such as color features, texture features, and transformation features, without limitation.

[0128] In this way, the tablet computer 10 can save the characteristics of the tracking target in multiple dimensions. Figure 16 The tracking target is divided into four input images, and the features of the four input images are obtained and saved respectively. However, the embodiment of the present application does not limit how the tracking target is specifically divided into the input images for feature extraction. For example, the tablet computer 10 can divide the tracking target into more or fewer input images based on key points, which is not limited here. The embodiment of the present application does not limit the feature extraction method of the input image. It is understandable that during the process of target tracking by the tablet computer 10, when extracting features from the input image corresponding to the tracking target and the input image corresponding to the candidate target, the tablet computer 10 always uses the same feature extraction method.

[0129] 4. Determine whether the candidate target is the tracking target in the N+1th frame image

[0130] Figure 17 The user interface 700 of the tablet computer 10 is shown as an example. The user interface 700 may display the N+1 frame image. The user interface 700 may include candidate target person A and candidate target person B. The user interface 700 may include prompt text 701, an indicator box 702, and a detection box 703. The prompt text 701 is used to indicate that the tracking target has been detected in the N+1 frame image. The indicator box 702 is used to indicate that the circled object (e.g., person A) is the tracking target. The detection box 703 is used to circle person B, indicating that the tablet computer 10 has detected person B in the user interface 700.

[0131] After tablet computer 10 detects candidate targets (person A and person B) in the N+1th frame, it extracts features from the candidate targets. Tablet computer 10 then matches the features of candidate targets A and B with the features of the tracking target. Tablet computer 10 determines that the features of target person A match the features of the tracking target. Tablet computer 10 displays prompt text 701 and an indicator box 702 in the user interface.

[0132] Figure 18 The input image for feature extraction corresponding to the candidate target person A and the features corresponding to the input image are exemplarily shown. The candidate target person A can be divided into two input images for feature extraction based on key points. That is, the candidate target person A can be divided into input image 1801 and input image 1802 based on key points. Specifically, the body part of the candidate target person A containing key points A1 to key points A4 (that is, the body part above the tracking target shoulder) can be divided into one image for feature extraction (that is, input image 1801). The body part of the candidate target person A containing key points A1 to key points A10 (that is, the body part above the tracking target hip) can be divided into one image for feature extraction (that is, input image 1802).

[0133] The tablet computer 10 performs feature extraction on the input image 1801 to obtain features of the input image 1801. For example, the features of the input image 1801 can be represented by a feature vector F1'. The feature vector F1' can represent one or more of the color features, texture features, transformation features, and the like of the input image 1801, without limitation. The tablet computer 10 performs feature extraction on the input image 1802 to obtain features of the input image 1802. The features of the input image 1802 can be represented by a feature vector F2'. The feature vector F2' can represent one or more of the color features, texture features, transformation features, and the like of the input image 1801, without limitation.

[0134] Figure 19 The following example illustrates an input image for feature extraction corresponding to candidate target person B and features corresponding to the input image. Candidate target person B can be divided into a single input image for feature extraction based on key points. That is, candidate target person B can be divided into input image 1901 based on key points. Specifically, the body portion of candidate target person B containing key points A1-A4 (i.e., the body portion above the tracking target's shoulders) can be divided into a single image for feature extraction (i.e., input image 1901).

[0135] Tablet computer 10 performs feature extraction on input image 1901 to obtain features of input image 1901. For example, the features of input image 1901 can be represented by feature vector V1. Feature vector V1 can represent one or more features of input image 1901, such as color features, texture features, and transformation features, without limitation.

[0136] In this way, the tablet computer 10 can match the features of a candidate target consisting solely of body parts above the hips with the features of a tracking target consisting solely of body parts above the hips. If the tracking target and the candidate target are the same person, the tablet computer 10 can determine that the candidate target in the next frame is the user-specified tracking target in the previous frame. Similarly, if the tracking target and the candidate target are the same person, but the person's body parts appear differently, or if part of the candidate target's body is obscured, the tablet computer 10 can still accurately determine that the candidate target is the user-specified tracking target.

[0137] Based on the above content, the application scenario of the target tracking method proposed in the embodiment of the present application is introduced in combination with UI. The following is a detailed description of a target tracking method proposed in the embodiment of the present application in combination with the accompanying drawings. Figure 20 As shown, the method may include:

[0138] S100: The electronic device displays a user interface A, wherein the user interface A displays an Nth frame image, and the Nth frame image includes a tracking target.

[0139] An electronic device is a device that can display a user interface and has a target tracking function, such as Figure 1 The electronic device may also be a smart phone, a television, etc., which is not limited here.

[0140] The electronic device can display a user interface A. The user interface A is used to display the Nth frame image. The Nth frame image contains the tracking target. The Nth frame image is an image captured by the camera of the electronic device, or an image or a frame of video captured by a video surveillance device. For example, the user interface A can be as follows Figure 8 The user interface 400 is shown. The person A in the user interface 400 may be a tracking target. In the embodiment of the present application, the user interface A may be referred to as a first user interface.

[0141] It is understandable that before the electronic device displays the user interface A, the electronic device turns on the target tracking function. There are many ways for the electronic device to turn on the target tracking function. For example, the user can turn on the target tracking function in a specific application of the electronic device (for example Figure 1 In the user interface of the video surveillance application shown (eg Figure 5A User interface 300 shown) clicks on a control for turning on the target tracking function (eg Figure 5A In response to the user operation, the electronic device starts the target tracking function. Figure 1-Figure 7 The process of enabling target tracking on the tablet computer 10 shown in FIG. For another example, a user may input a voice command (e.g., "Please enable target tracking") to the electronic device to enable the target tracking function. In response to the user's voice instruction, the electronic device enables the target tracking function. The embodiments of the present application do not limit the manner in which the electronic device enables the target tracking function.

[0142] S101: The electronic device determines a tracking target in user interface A.

[0143] There are many ways for the electronic device to determine the tracking target in the user interface A.

[0144] For example, in one possible implementation, the electronic device may detect multiple objects (e.g., people, plants, vehicles, etc.) in the Nth frame image. The electronic device may display a detection frame to encircle the object detected by the electronic device. The user may click on the object encircled by the detection frame. In response to the user operation, the electronic device may use the image portion encircled by the detection frame as a tracking target. Figures 8-10 The process of determining the tracking target by the electronic device described in . For example, Figure 9 The detection frame 402 shown encircles the person A, which is the tracking target. The electronic device can capture the image portion encircled by the detection frame 402 from the Nth frame image as the image corresponding to the tracking target (person A).

[0145] It is understandable that the electronic device can detect a person in an image with a person and a background. Specifically, the electronic device can detect objects (such as people, plants, vehicles, etc.) in the image through a target detection model. The input of the target detection model can be an image. For example, Figure 7 The output of the object detection model can be an image with the objects in the image labeled. For example, Figure 8 The user interface 400 shows an image in which a person A is encircled by a detection frame 402 and a person B is encircled by a detection frame 403. How the electronic device detects objects in the image can be referred to the existing technology and will not be described in detail here.

[0146] Optionally, the user can draw an indicator frame in the user interface for selecting a tracking target. In response to the user operation, the electronic device determines that the object in the indicator frame is the tracking target. Figure 21 As shown, a user can draw an indicator frame 802 in user interface 800. The object within indicator frame 802 is the tracking target. The image portion enclosed by indicator frame 802 is the image corresponding to the tracking target. The indicator frame can be a rectangular frame, a square frame, a diamond frame, etc., and the shape of the indicator frame is not limited here.

[0147] It is understandable that the embodiments of the present application do not limit the manner in which the electronic device determines the tracking target.

[0148] S102: The electronic device obtains a plurality of tracking target images for feature extraction according to the tracking target, where the plurality of tracking target images include a part or all of the tracking target.

[0149] The electronic device can divide the image corresponding to the tracking target into multiple tracking target images for feature extraction. The multiple tracking target images can include part or all of the tracking target. In this way, the electronic device can obtain tracking target images corresponding to the tracking target in multiple dimensions for feature extraction.

[0150] In a possible implementation, the electronic device can divide the image corresponding to the tracking target into multiple tracking target images according to the key points. For example, if the tracking target is a person, Figure 15 As shown, the electronic device can detect 14 key points of the human body, namely key points A1 to A14. Then, the electronic device divides the image corresponding to the tracking target into multiple tracking target images for feature extraction according to the key points of the tracking target.

[0151] Specifically, if Figure 16 As shown, the electronic device can divide the image corresponding to the tracking target into an input image containing only the key points A1 to A4 of the tracking target (e.g., input image 1601), an input image containing only the key points A1 to A10 of the tracking target (e.g., input image 1602), an input image containing only the key points A1 to A12 of the tracking target (e.g., input image 1603), and an input image containing the key points A1 to A14 of the tracking target (e.g., input image 1604). Figure 16 In this way, when the tracking target in the next frame image only appears above the shoulders, or only appears above the hips, or only appears above the knees, the electronic device can also accurately determine the tracking target.

[0152] Alternatively, as Figure 22 As shown, the electronic device can divide the image corresponding to the tracking target into an input image (e.g., input image 2201) that only includes key points A1 and A2 of the tracking target, an input image (e.g., input image 2202) that only includes key points A3-A8 of the tracking target, an input image (e.g., input image 2203) that only includes key points A9-A12 of the tracking target, and an input image (e.g., input image 2204) that includes key points A13-A14 of the tracking target. Input images 2201, 2202, 2203, and 2204 are only portions of the tracking target. That is, input image 2201 only includes the head of the tracking target. Input image 2202 only includes the upper body of the tracking target (i.e., the body portion from the shoulders to the hips, excluding the hips). Input image 2203 only includes the lower body of the tracking target (i.e., the body portion from the hips to the feet, excluding the feet). Input image 2204 only includes the feet of the tracking target. There is no repeated body part between any two input images of input image 2201, input image 2202, input image 2203, and input image 2204. For example, there is no repeated body part between input image 2201 and input image 2202. There is no repeated body part between input image 2201 and input image 2203.

[0153] Optionally, the electronic device may divide the image corresponding to the tracking target into an input image containing only the key points A1-key points A4 of the tracking target's body part (e.g., input image 1601), an input image containing only the key points A1-key points A10 of the tracking target's body part (e.g., input image 1602), an input image containing only the key points A1-key points A12 of the tracking target's body part (e.g., input image 1603), an input image containing the key points A1-key points A14 of the tracking target's body part (e.g., input image 1604), an input image containing only the key points A1 and A2 of the tracking target's body part (e.g., input image 2201), an input image containing only the key points A3-A8 of the tracking target's body part (e.g., input image 2202), an input image containing only the key points A9-A12 of the tracking target's body part (e.g., input image 2203), and an input image containing the key points A13-A14 of the tracking target's body part (e.g., input image 2204). For example, the electronic device may Figure 16 The tracking targets shown are divided into Figure 16 Input image 1601, input image 1602, input image 1603, input image 1604, and Figure 22 The input images 2201, 2202, 2203 and 2204 are shown. Here, Figure 16 Input image 1601, input image 1602, input image 1603, input image 1604, and Figure 22 The input image 2201 , input image 2202 , input image 2203 , and input image 2204 shown may all be referred to as tracking target images in the embodiments of the present application.

[0154] It is understandable that the electronic device can divide the image corresponding to the tracking target into Figure 16 or Figure 22 More or fewer tracking target images are shown. In the embodiment of the present application, the electronic device can divide the tracking target into multiple tracking target images based on the key points. The embodiment of the present application does not limit which key points of the tracking target are included in the divided tracking target images, nor the specific number of tracking target images.

[0155] It is understandable that the electronic device can detect the key points of the human body (such as bone points). The electronic device can identify the key points of the person in the image frame displayed in the electronic device through the human key point recognition algorithm. Here, identifying the key points can refer to determining the position information of the key points. The position information of the identified key points can be used to separate the tracking target image from the corresponding image of the tracking target. Among them, the input of the human key point recognition algorithm can be a human body image. The output of the human key point recognition algorithm can be the position information of the human key points (for example, two-dimensional coordinates). The electronic device can identify the key points such as Figure 15 The key points A1 to A14 shown in the figure can correspond to the basic skeleton points of the human body, such as the head skeleton point, neck skeleton point, left shoulder skeleton point, right shoulder skeleton point, left elbow skeleton point, right elbow skeleton point, left hand skeleton point, right hand skeleton point, left hip skeleton point, right hip skeleton point, left knee skeleton point, right knee skeleton point, left foot skeleton point, and right foot skeleton point. Figure 15 The description of Figure 15 As shown, the electronic device can identify more or fewer key points of the human body, and this embodiment of the present application is not limited to this.

[0156] S103: The electronic device extracts features from the multiple tracking target images, and obtains and saves tracking target features corresponding to the multiple tracking target images.

[0157] Specifically, the electronic device can perform feature extraction on multiple tracking target images through a feature extraction algorithm to obtain and save the features of the multiple tracking target images (for example, one or more of the color features, texture features, transformation features, etc.). A feature vector obtained by the electronic device through feature extraction on the tracking target image is the tracking target feature corresponding to the tracking target image. The electronic device can save the tracking target feature corresponding to the tracking target image. It is understandable that one tracking target image can correspond to one tracking target feature. If there are multiple tracking target images, the electronic device can obtain and save multiple tracking target features.

[0158] The input of the feature extraction algorithm may be a tracking target image. The output of the feature extraction algorithm may be a feature vector of the tracking target image. It is understood that the feature vectors corresponding to different input images have the same dimension, but different values ​​in the feature vectors.

[0159] like Figure 16 As shown, the tracking target can be divided into input image 1601, input image 1602, input image 1603 and input image 1604. Input image 1601 may include key points A1 to A4. The corresponding features of input image 1601 may be one or more of the color features, texture features, transformation features, etc. of input image 1601. For example Figure 16 The feature vector F1 shown in FIG. 1 may represent one or more of the color features, texture features, transformation features, etc. of the input image 1601, which are not limited here. The specific form of the feature vector F1 is not limited in this embodiment of the application.

[0160] The corresponding feature of the input image 1602 may be one or more of the color feature, texture feature, transformation feature, etc. of the input image 1602. For example Figure 16 The feature vector F2 shown in FIG. 1 may represent one or more of the color features, texture features, transformation features, etc. of the input image 1602, which is not limited here. The corresponding features of the input image 1603 may be one or more of the color features, texture features, transformation features, etc. of the input image 1603, which is not limited here. For example Figure 16 The feature vector F3 shown in FIG. 1 may represent one or more of the color features, texture features, transformation features, etc. of the input image 1603, which is not limited here. The corresponding feature of the input image 1604 may be one or more of the color features, texture features, transformation features, etc. of the input image 1604, which is not limited here. For example Figure 16 The feature vector F4 shown in FIG. 4 may represent one or more of the color features, texture features, transformation features, etc. of the input image 1604, which are not limited here. The electronic device extracts features from the input images 1601, 1602, 1603, and 1604 and saves the features. The electronic device may save the features corresponding to the input images in the following form: Figure 16 As shown. The electronic device can use a table to save the tracking target, the input image obtained by dividing the tracking target, and the features corresponding to the input image. Figure 16 In

[15] , the feature vector F1, the feature vector F2, the feature vector F3 and the feature vector F4 are all mapped to the tracking target. That is, the feature vector F1, the feature vector F2, the feature vector F3 and the feature vector F4 can all represent the tracking target.

[0161] like Figure 22 As shown, the tracking target can be divided into input image 2201, input image 2202, input image 2203 and input image 2204. Input image 2201 may include key point A1-key point A2. The corresponding features of input image 2201 may be one or more of the color features, texture features, transformation features, etc. of input image 2201. For example Figure 22 The feature vector F5 shown in FIG. 5 may represent one or more of the color features, texture features, transformation features, etc. of the input image 2201, which is not limited here. The corresponding features of the input image 2202 may be one or more of the color features, texture features, transformation features, etc. of the input image 2202. For example, Figure 22 The feature vector F6 shown in FIG. 2 may represent one or more of the color, texture, or transformation characteristics of the input image 2202, without limitation. Reference may be made to the description of feature vector F1 above for feature vector F6. It should be understood that the feature vector F6 corresponding to the input image 2202 is merely an example. The specific form of feature vector F6 is not limited in this embodiment of the present application.

[0162] The corresponding feature of the input image 2203 may be one or more of the color feature, texture feature, transformation feature, etc. of the input image 2203, which is not limited here. Figure 22 The feature vector F7 shown in FIG. 2 may represent one or more of the color, texture, or transformation features of the input image 2203, without limitation. It should be understood that the feature vector F7 corresponding to the input image 2203 is merely an example. The specific form of the feature vector F7 is not limited in this embodiment of the present application.

[0163] The corresponding feature of the input image 2204 may be one or more of the color feature, texture feature, transformation feature, etc. of the input image 2204, which is not limited here. Figure 22 The feature vector F8 shown in FIG. 1 may represent one or more of the color, texture, or transformation characteristics of the input image 2204, without limitation. Reference may be made to the description of feature vector F1 above for feature vector F8. It should be understood that the feature vector F8 corresponding to the input image 2204 is merely an example. The specific form of feature vector F8 is not limited in this embodiment of the present application.

[0164] The electronic device extracts features from input image 2201, input image 2202, input image 2203, and input image 2204 and saves the features. The electronic device can save the tracking target, the input image obtained by dividing the tracking target, and the features corresponding to the input image. Figure 22 In FIG, the feature vector F5, the feature vector F6, the feature vector F7 and the feature vector F8 are all mapped to the tracking target. That is, the feature vector F5, the feature vector F6, the feature vector F7 and the feature vector F8 can all represent the tracking target.

[0165] In one possible implementation, for Figure 10 The electronic device shown in FIG 1 determines the tracking target (person A), and performs feature extraction. The features obtained and stored may be feature vector F1, feature vector F2, feature vector F3, and feature vector F4. That is, the electronic device divides the tracking target into input image 1601, input image 1602, input image 1603, and input image 1604. The electronic device performs feature extraction on input image 1601, input image 1602, input image 1603, and input image 1604, respectively, and may obtain feature vector F1, feature vector F2, feature vector F3, and feature vector F4, respectively. The electronic device may store feature vector F1, feature vector F2, feature vector F3, and feature vector F4.

[0166] Optionally, for Figure 10 The electronic device shown in FIG2 determines the tracking target (person A), and performs feature extraction. The features obtained and stored may be feature vector F5, feature vector F6, feature vector F7, and feature vector F8. That is, the electronic device divides the tracking target into input image 2201, input image 2202, input image 2203, and input image 2204. The electronic device performs feature extraction on input image 2201, input image 2202, input image 2203, and input image 2204, respectively, and may obtain feature vector F5, feature vector F6, feature vector F7, and feature vector F8, respectively. The electronic device may store feature vector F5, feature vector F6, feature vector F7, and feature vector F8.

[0167] Optionally, for Figure 10 The electronic device determines the tracking target (person A) shown in FIG, and performs feature extraction on the tracking target. The features obtained and stored may be feature vector F1, feature vector F2, feature vector F3, feature vector F4, feature vector F5, feature vector F6, feature vector F7, and feature vector F8. That is, the electronic device divides the tracking target into input image 1601, input image 1602, input image 1603, and input image 1604, and input image 2201, input image 2202, input image 2203, and input image 2204. The electronic device performs feature extraction on input image 1601, input image 1602, input image 1603, input image 1604, and input image 2201, input image 2202, input image 2203, and input image 2204, respectively, and may obtain feature vector F1, feature vector F2, feature vector F3, feature vector F4, feature vector F5, feature vector F6, feature vector F7, and feature vector F8, respectively. The electronic device may store the feature vector F1 , the feature vector F2 , the feature vector F3 , the feature vector F4 , the feature vector F5 , the feature vector F6 , the feature vector F7 , and the feature vector F8 .

[0168] It is understood that feature vector F1, feature vector F2, feature vector F3, feature vector F4, feature vector F5, feature vector F6, feature vector F7, and feature vector F8 can all be referred to as tracking target features in the embodiments of the present application. The embodiments of the present application do not limit the number of tracking target features that can be extracted by the electronic device.

[0169] It is understandable that for different tracking targets, the number of tracking target features extracted by the electronic device may be different.

[0170] It is understood that there may be multiple tracking targets in the embodiments of the present application. When there are multiple tracking targets, the electronic device may, according to steps S102-S103, divide each tracking target into tracking target images, perform feature extraction on the tracking target images, and store tracking target features. The electronic device may store the tracking target images corresponding to the multiple tracking targets and the tracking target features corresponding to the tracking target images.

[0171] S104: The electronic device displays a user interface B, where the user interface B displays the N+1th frame image, which includes one or more candidate targets.

[0172] The electronic device can display a user interface B, which can be Figure 17 User interface 700 is shown. User interface B displays the N+1th frame image. The N+1th frame image is the next frame image after the Nth frame image. The N+1th frame image may contain one or more candidate targets. As shown in user interface 700, person A and person B are both candidate targets. In this embodiment of the present application, user interface B may be referred to as a second user interface.

[0173] In a possible implementation, if the electronic device does not detect the candidate target in the N+1th frame image, the electronic device may display the N+2th frame image and then perform steps S104 to S107 on the N+2th frame image.

[0174] It is understood that the electronic device may not detect the candidate target in the N+1th frame image, but the candidate target appears in the N+1th frame but the electronic device does not detect it. Alternatively, the candidate target may not appear in the N+1th frame image. For example, the candidate target does not appear in the shooting range of the electronic device's camera, and the electronic device does not display the candidate target in the N+1th frame image. In this case, the electronic device may display the N+2th frame image. Then, steps S104-S107 are performed on the N+2th frame image.

[0175] S105: The electronic device obtains multiple candidate target images corresponding to each candidate target in the one or more candidate targets, where the multiple candidate target images include a part or all of the candidate targets.

[0176] The electronic device detects that there are one or more candidate targets in the N+1 frame image. Then, the electronic device can divide each of the one or more candidate targets into a candidate target image for feature extraction. The candidate target image corresponding to each candidate target contains part or all of the candidate target. For example, Figure 17 The N+1th frame image is displayed in the user interface 700. The N+1th frame image includes a candidate target person A and a candidate target person B.

[0177] like Figure 18 As shown, the candidate target person A can be divided into input image 1801 and input image 1802. Figure 19 As shown, the candidate target person B can be divided into input image 1901. Here, input image 1801, input image 1802 and input image 1901 can all be referred to as candidate target images in the embodiment of the present application.

[0178] It is understood that the method for dividing a candidate target into multiple candidate target images is consistent with the method for dividing a tracking target into multiple tracking target images. That is, if the tracking target is divided into an input image that only includes the tracking target's key points A1-A4 body part, an input image that only includes the tracking target's key points A1-A10 body part, an input image that only includes the tracking target's key points A1-A12 body part, and an input image that includes the tracking target's key points A1-A14 body part. If the candidate target is a full-body image that includes key points A1-A14, then the candidate target should also be divided into an input image that only includes the tracking target's key points A1-A4 body part, an input image that only includes the tracking target's key points A1-A10 body part, an input image that only includes the tracking target's key points A1-A12 body part, and an input image that includes the tracking target's key points A1-A14 body part. If the candidate target is a half-body image that includes key points A1-A10 (e.g., person A shown in user interface 700). Then the candidate targets are divided into an input image containing only the body part of the tracking target's key points A1 to A4 and an input image containing only the body part of the tracking target's key points A1 to A10.

[0179] S106 : The electronic device performs feature extraction on the multiple candidate target images to obtain candidate target features corresponding to the multiple candidate target images.

[0180] Specifically, the electronic device can perform feature extraction on multiple candidate target images using a feature extraction algorithm to obtain multiple candidate target image features (e.g., one or more of color features, texture features, transformation features, etc.). A feature vector obtained by the electronic device through feature extraction of the candidate target image is the candidate target feature corresponding to the candidate target image. Reference can be made to the description of feature extraction of the tracking target image in the above steps, which will not be repeated here.

[0181] like Figure 18 As shown, the candidate target person A can be divided into input image 1801 and input image 1802. Input image 1801 may include key points A1 to A4. The corresponding features of input image 1801 may be one or more of the color features, texture features, transformation features, etc. of input image 1801. For example Figure 18 The feature vector F1' shown in FIG. 1 is used to represent one or more of the color, texture, and transformation features of the input image 1801, which are not limited here. The feature vector F1' corresponding to the input image 1801 is merely an example. The specific form of the feature vector F1' is not limited in this embodiment of the present application.

[0182] The corresponding features of the input image 1802 may be one or more of the color features, texture features, transformation features, etc. of the input image 1802. For example Figure 18 The feature vector F2' shown in FIG. 1 may represent one or more of the color, texture, or transformation characteristics of input image 1802, without limitation. Reference may be made to the description of feature vector F1 above for feature vector F2'. It should be understood that the feature vector F2' corresponding to input image 1802 is merely an example. The specific form of feature vector F2' is not limited in this embodiment of the present application.

[0183] like Figure 19 As shown, the candidate target person A can be divided into an input image 1901. The input image 1901 may include key points A1 to A4 of the human body. The corresponding features of the input image 1901 may be one or more of the color features, texture features, transformation features, etc. of the input image 1901. For example Figure 19 The feature vector V1 shown in FIG. 1 may represent one or more of the color features, texture features, transformation features, and other features of the input image 1901, without limitation. The feature vector V1 corresponding to the input image 1901 is merely an example. The specific form of the feature vector V1 is not limited in this embodiment of the present application.

[0184] Alternatively, the electronic device may also perform feature extraction on only candidate target images containing more key points. For example, Figure 18 1 and 1802 are shown in FIG. Input image 1802 contains more key points, so the electronic device can only perform feature extraction on input image 1802. In this way, the efficiency of the electronic device can be improved.

[0185] S107: The electronic device performs feature matching on the first candidate target feature and the first tracking target feature. If the first candidate target feature matches the first tracking target feature, the electronic device determines that the candidate target is a tracking target.

[0186] The electronic device may perform feature matching of the first candidate target feature with the first tracking target feature. The electronic device may perform feature matching of the first candidate target feature corresponding to the first candidate target image containing the most key points of the human body among the candidate target features with the first tracking target feature. The multiple candidate target images include the first candidate target image. The multiple tracking target images include the first tracking target image. The first tracking target feature is obtained by the electronic device from the first tracking target image. For example, for Figure 18 For the candidate target person A shown, the electronic device can perform feature matching between the candidate target image containing more key points of candidate target person A, that is, the feature vector F2' corresponding to the input image 1802, and the tracking target feature (i.e., feature vector F2) corresponding to the candidate target image containing the same number of key points in the tracking target (i.e., input image 1602). Here, input image 1802 is the first candidate target image, and feature vector F2' is the first candidate target feature. Input image 1602 is the first tracking target image, and feature vector F2 is the first tracking target feature.

[0187] In one possible implementation, the electronic device may calculate the Euclidean distance D1 between feature vector F2' and feature vector F2. If the Euclidean distance D1 between feature vector F2' and feature vector F2 is less than a preset Euclidean distance D, the electronic device determines that feature vector F2' matches feature vector F2. Furthermore, the electronic device determines that candidate target person A is the tracking target. It will be appreciated that the preset Euclidean distance D may be configured by the electronic device's system.

[0188] In a possible implementation, the multiple tracking target images include a first tracking target image and a second tracking target image, and the multiple tracking target images are obtained from the tracking target; the multiple candidate target images include a first candidate target image and a second candidate target image, and the multiple candidate target images are obtained from the first candidate target; the number of key points of the tracking target contained in the first tracking target image is the same as the number of key points of the first candidate target contained in the first candidate target image; the number of key points of the tracking target contained in the second tracking target image is the same as the number of key points of the first candidate target contained in the second candidate target image; the number of tracking target key points contained in the first tracking target image is greater than the number of tracking target key points contained in the second tracking target image; the first tracking target feature is extracted from the first tracking target image, and the first candidate target feature is extracted from the first candidate target image.

[0189] In one possible implementation, the electronic device selects the first candidate target feature and the first tracking target feature that contain the same number of key points and the largest number of key points from the tracking target feature of the tracking target and the candidate target features of the candidate targets for feature matching. Here, the first candidate target feature that contains the largest number of key points means that the first candidate target image corresponding to the first candidate target is the candidate target image that contains the largest number of key points among multiple candidate target images. The first tracking target feature that contains the largest number of key points means that the first tracking target image corresponding to the first tracking target is the tracking target image that contains the largest number of key points among multiple tracking target images. Here, the first candidate target image corresponding to the first candidate target feature means that the electronic device extracts the first candidate target feature from the first candidate target image. The first tracking target image corresponding to the first tracking target feature means that the electronic device can extract the first tracking target feature from the first tracking target image. For example, if the multiple tracking target features corresponding to the tracking target include a feature vector F1 containing feature information of key points A1 to A4 of the tracking target, a feature vector F2 containing feature information of key points A1 to A10 of the tracking target, a feature vector F3 containing feature information of key points A1 to A12 of the tracking target, and a feature vector F4 containing feature information of key points A1 to A14 of the tracking target, and the multiple candidate target features corresponding to the candidate target include a feature vector F1' containing feature information of key points A1 to A4 of the candidate target, and a feature vector F2' containing feature information of key points A1 to A10 of the candidate target, then the electronic device can perform feature matching on the feature vector F2' containing the same number of key points with the feature vector F2.

[0190] Furthermore, when the candidate target feature with the largest number of key points does not match the tracking target feature, the electronic device may perform feature matching on the candidate target feature with the smaller number of key points and the tracking target feature containing the same number of key points. For example, if feature vector F2' does not match feature vector F2, the electronic device may perform feature matching on feature vector F1' with feature vector F1.

[0191] Optionally, when the candidate target feature with the largest number of key points does not match the tracking target feature, the electronic device displays the N+2 frame image and performs steps S104 to S107 on the N+2 frame image.

[0192] In one possible implementation, when the electronic device determines that candidate target A in the N+1th frame is the tracking target selected by the user in the Nth frame, the electronic device may add the candidate target features corresponding to candidate target A in the N+1th frame to the tracking target features corresponding to the tracking target stored by the electronic device. That is, the features corresponding to the tracking target stored by the electronic device become tracking target features and candidate target features. For example, the tracking target features corresponding to the tracking target may only include feature vector F1 and feature vector F2. The candidate target features corresponding to candidate target A in the N+1th frame include feature vector F1', feature vector F2', feature vector F3' containing feature information of key points A1-A12 of candidate target A, and feature vector F4' containing feature information of key points A1-A14 of candidate target A. The electronic device may add feature vector F3' and feature vector F4' to the tracking target features corresponding to the stored tracking target. Alternatively, if the tracking target is a standing person A and the candidate target is a squatting or walking person A, the electronic device is more likely to extract different features for the same body part when the person is in different postures. The electronic device can add all the feature vectors F1', F2', F3', and F4' to the stored tracking target features corresponding to the tracking target. In this way, the electronic device can also store the features of each body part of the same person in different postures.

[0193] Optionally, when the electronic device determines that candidate target A in the N+1th frame image is the tracking target selected by the user in the Nth frame image, the electronic device may update the tracking target features corresponding to the tracking target stored in the electronic device to the candidate target features corresponding to candidate target A in the N+1th frame image. For example, the tracking target features corresponding to the tracking target only include feature vector F1 and feature vector F2. The candidate target features corresponding to candidate target A in the N+1th frame include feature vector F1' and feature vector F2'. The electronic device updates the feature vector F1 and feature vector F2 corresponding to the stored tracking target to feature vector F1' and feature vector F2'.

[0194] Specifically, the electronic device may update the tracking target features corresponding to the tracking target stored in the electronic device every preset number of frames. For example, the preset number of frames may be 10, meaning that the electronic device updates the tracking target features corresponding to the tracking target stored in the electronic device every 10 frames. The preset number of frames may be configured by the electronic device system, and the specific value of the preset number of frames may be 5, 10, or other values. The embodiment of the present application does not limit the value of the preset number of frames.

[0195] After the electronic device completes step S107, it displays the N+2 frame image and continues to execute steps S104 to S107. The electronic device stops executing steps S104 to S107 until the target tracking function is stopped or turned off.

[0196] Optionally, when the electronic device does not detect the tracking target within a preset time period, the electronic device can stop or turn off the target tracking function. It is understood that the preset time period can be configured by the electronic device system. The specific value of the preset time period is not limited in this embodiment of the application.

[0197] In implementing a target tracking method provided by an embodiment of the present application, an electronic device may divide a tracked target identified in an Nth frame of image into multiple tracked target images for feature extraction. Thus, the electronic device extracts features from the multiple tracked target images to obtain and store multiple dimensions of target features. The electronic device then detects a candidate target in the N+1th frame of image and divides the candidate target into multiple candidate target images for feature extraction in the same manner as the tracked target image was divided. The electronic device extracts features from the candidate target images to obtain multiple dimensions of candidate target features. The electronic device may select tracking target features and candidate target features that contain the same number of key point feature information for matching. The electronic device only uses features of the body parts common to the tracking target and candidate targets for matching. Thus, when the tracking target is a full-body person A and the candidate target is a half-body person A, the electronic device can accurately detect that the candidate target is the tracking target. Even if the tracked object is obscured, partially displayed, or deformed, the electronic device can accurately track the target as long as the user-selected tracking target and the candidate target share common parts. This improves the accuracy of target tracking by the electronic device.

[0198] Electronic devices can track designated objects (such as people), and the tracking target may deform. For example, when the tracking target is a person, the person may change from a squatting position to a sitting position or a standing position. That is, the position of the tracking target will change in consecutive image frames. When the position of the same person in the Nth frame image and the N+1th frame image changes, the electronic device may not be able to accurately detect the tracking target specified by the user in the Nth frame image in the N+1th frame image. This will cause the electronic device to fail in target tracking.

[0199] In order to improve the target tracking accuracy of the electronic device when the target is deformed, the embodiment of the present application provides a target tracking method. Figure 23 As shown, the method may specifically include:

[0200] S200: The electronic device displays a user interface A, wherein the user interface A displays an Nth frame image, and the Nth frame image includes a tracking target.

[0201] Step S200 may refer to step S100 and will not be described in detail here.

[0202] S201: The electronic device determines a tracking target and a first position of the tracking target in the user interface A.

[0203] How the electronic device determines the tracking target in user interface A can be referred to as described in step S201. The electronic device can identify the posture of the tracking target. The first posture can be any one of squatting, sitting, walking, and standing. In the embodiments of the present application, user interface A can be referred to as the first user interface.

[0204] In one possible implementation, the electronic device can identify the posture of a human body through human key points and human posture recognition technology. The input of the human posture recognition technology is the human key point information (such as the position information of the key points). The output of the human posture recognition technology is the posture of the human body, such as squatting, sitting, walking, standing, etc. Different postures of the person A correspond to different feature vectors, such as Figure 24 The features corresponding to the person A in different postures are shown in FIG. Figure 24 In the image 2401, person A is squatting. The feature vector corresponding to image 2401 is F11. F11 may include information about each key point of person A in the squatting position. Image 2402 is person A in the sitting position. The feature vector corresponding to image 2401 is F22. F22 may include information about each key point of person A in the sitting position. Image 2403 is person A in the squatting position. The feature vector corresponding to image 2403 is F33. F33 may include information about each key point of person A in the walking position. Image 2404 is person A in the standing position. The feature vector corresponding to image 2404 is F44. F44 may include information about each key point of person A in the standing position.

[0205] It is understandable that the posture of the human body is not limited to Figure 24 The squatting, sitting, walking, and standing shown in the figure. The human body posture can also include half squatting, running, jumping, etc., which are not limited in the embodiments of this application. The posture recognition technology of electronic devices is not limited to recognizing human posture based on key points of the human body, and the embodiments of this application do not limit the method of posture recognition of electronic devices.

[0206] S202: The electronic device extracts features of the tracking target to obtain and save tracking target features corresponding to the tracking target.

[0207] The electronic device can extract the features of the tracking target with a determined posture, obtain and save the tracking target features. Figure 25 As shown, the tracking target specified by the user in the Nth frame image can be a person A in a squatting position, that is, Figure 25 The electronic device extracts features of the tracking target, obtains a first feature, and saves the tracking target features. Here, reference may be made to the description of the feature extraction of the tracking target in step S103 above, which will not be repeated here.

[0208] S203: The electronic device displays a user interface B, where the user interface B displays the N+1th frame image, which includes one or more candidate targets.

[0209] Step S203 may refer to the description of step S104 and will not be described in detail here. In the embodiment of the present application, user interface B may be referred to as a second user interface.

[0210] S204: The electronic device determines a second posture of one or more candidate targets and performs feature extraction on the candidate targets to obtain candidate target features.

[0211] The electronic device can detect one or more candidate targets in the N+1 frame image. The electronic device can also determine the second posture of the one or more candidate targets. The second posture can be any one of squatting, sitting, walking, and standing. The electronic device can extract features of the candidate targets in the second posture. For example, Figure 26 As shown, the candidate target in the N+1th frame image is image 2601, that is, person A in a sitting position. The electronic device extracts features of person A in a sitting position to obtain candidate target features, such as feature vector F22.

[0212] S205. The electronic device determines that a first candidate target among one or more candidate targets is a tracking target. If the first pose and the second pose are different, the electronic device saves the candidate target feature into a feature library corresponding to the tracking target, where the tracking target feature is stored.

[0213] The N+1th frame image may include one or more candidate targets, one or more of which may include the first target. The electronic device may determine whether the first candidate target in the N+1th frame is the tracking target specified by the user in the Nth frame. The electronic device may determine that the first candidate target is the tracking target in various ways.

[0214] In one possible implementation, if the Nth frame image contains only one object, namely the tracking target, and the N+1th frame image contains only one object, namely the first candidate target, the electronic device may directly determine the first candidate target as the tracking target.

[0215] In one possible implementation, if the first pose of the tracking target is the same as the second pose of the first candidate target, the electronic device may perform feature matching between the tracking target feature and the candidate target feature. If the candidate target feature matches the tracking target feature, the electronic device determines that the first candidate target is the tracking target.

[0216] In one possible implementation, if there are multiple candidate targets in the N+1 frame image, the electronic device can obtain the position of the tracking target in the N frame image, for example, the center of the tracking target is at the first position in the N frame image. The electronic device can obtain the position of the first candidate target in the N+1 frame image, for example, the center of the first candidate target is at the second position in the N+1 frame image. If the preset distance between the first position and the second position is less than the preset distance, the electronic device can determine that the first candidate target is the tracking target. It can be understood that the preset distance can be configured by the system of the electronic device, and the embodiments of the present application do not specifically limit the preset distance.

[0217] In one possible implementation, if the electronic device saves the features of the tracking target at different postures, the electronic device can perform feature matching on the feature vector of the saved tracking target features whose posture is the same as the second posture of the first candidate target and the candidate target features. If there is a match, the electronic device determines that the first candidate target is the tracking target.

[0218] After the electronic device determines that the first candidate target is the tracking target, if the first posture of the tracking target is different from the second posture of the first candidate, the electronic device can save the candidate target features to the feature library corresponding to the tracking target, and the feature library stores the tracking target features. Figure 25 As shown in , the tracking target is person A who is squatting. Figure 26 As shown, the first candidate target is person A who is sitting. After the electronic device determines that the first candidate target is the tracking target, the electronic device can save the candidate target features corresponding to the first candidate target into the feature library corresponding to the target. Figure 27 As shown, the electronic device saves the feature vector F22 corresponding to the first candidate target into the feature library corresponding to the tracking target. That is, the features corresponding to the tracking target are increased from the feature vector F11 to the feature vector F11 and the feature vector F22.

[0219] It is understandable that if in subsequent image frames, such as the N+2 frame, the N+3 frame, and so on, the posture of person A changes to walking or standing, the electronic device can save the feature vectors corresponding to the different postures of person A into the feature library corresponding to the tracking target. Figure 28 As shown, the feature vectors corresponding to the tracking target increase from feature vector F11 to feature vector F11, feature vector F22, feature vector F33 and feature vector F44.

[0220] By implementing the embodiments of the present application, the electronic device can track a specified object (such as a person), and the tracking target may be deformed. For example, when the tracking target is a person, the person can change from a squatting position to a sitting position or a standing position. That is, the position of the tracking target will change in consecutive image frames. Because the electronic device can save the corresponding features of the tracking target in different positions into the feature library corresponding to the tracking target. When the position of the same person in the Nth frame image and the N+1th frame image changes, the electronic device can also accurately detect the tracking target specified by the user in the Nth frame image in the N+1th frame image. In this way, the accuracy of the electronic device in target tracking is improved.

[0221] The following introduces an exemplary electronic device 100 provided in an embodiment of the present application.

[0222] Figure 29 1 is a schematic structural diagram of an electronic device 100 provided in an embodiment of the present application.

[0223] The following embodiments are described in detail using electronic device 100 as an example. It should be understood that electronic device 100 may have more or fewer components than shown in the figure, may combine two or more components, or may have a different component configuration. The various components shown in the figure may be implemented in hardware, including one or more signal processing and / or application-specific integrated circuits, software, or a combination of hardware and software.

[0224] The electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0225] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0226] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0227] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0228] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0229] In an embodiment of the present application, the processor 110 can also be used to obtain multiple candidate target features of the first candidate target; when the first candidate target feature among the multiple candidate target features matches the first tracking target feature among the multiple tracking target features, the first candidate target is determined to be the tracking target; the tracking target is determined by the processor in the Mth frame image; the multiple tracking target features are features obtained by the processor from the tracking target; M is less than K.

[0230] In an embodiment of the present application, the processor 110 can also be used to obtain multiple tracking target images based on the tracking target, where the tracking target image contains part or all of the tracking target; perform feature extraction on the multiple tracking target images to obtain multiple tracking target features, where the number of the multiple tracking target features is equal to the number of the multiple tracking target images.

[0231] In an embodiment of the present application, the processor 110 can also be used to obtain multiple candidate target images based on the first candidate target; the multiple candidate target images contain part or all of the first candidate target; and feature extraction is performed on the multiple candidate target images respectively to obtain multiple candidate target features, and the number of the multiple candidate target features is equal to the number of the multiple candidate target features.

[0232] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0233] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C bus lines. The processor 110 may be coupled to the touch sensor 180K, the charger, the flash, the camera 193, and the like via different I2C bus interfaces. For example, the processor 110 may be coupled to the touch sensor 180K via the I2C interface, enabling communication between the processor 110 and the touch sensor 180K via the I2C bus interface, thereby implementing the touch function of the electronic device 100.

[0234] The I2S interface can be used for audio communication.

[0235] The PCM interface can also be used for audio communication to sample, quantize and encode analog signals.

[0236] The UART interface is a universal serial data bus used for asynchronous communication.

[0237] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display 194 and the camera 193. MIPI interfaces include the camera serial interface (CSI) and the display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to implement the camera function of the electronic device 100. The processor 110 and the display 194 communicate via the DSI interface to implement the display function of the electronic device 100.

[0238] The GPIO interface can be configured via software. The GPIO interface can be configured as either a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, display 194, wireless communication module 160, audio module 170, sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0239] The SIM interface can be used to communicate with the SIM card interface 195 to implement the function of transmitting data to the SIM card or reading data in the SIM card.

[0240] The USB interface 130 is an interface that complies with USB standard specifications, and specifically may be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc.

[0241] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present invention is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0242] The charging management module 140 is configured to receive charging input from a charger, which may be a wireless charger or a wired charger.

[0243] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to provide power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160.

[0244] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0245] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals.

[0246] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc., applied to the electronic device 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc.

[0247] The modem processor includes a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a medium- or high-frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is passed to the application processor.

[0248] The wireless communication module 160 can provide wireless communication solutions for application on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication technology (NFC), infrared technology (IR), etc.

[0249] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0250] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than one.

[0251] In this embodiment, the display screen can be used to display the first user interface and the second user interface. The display screen can display the Nth frame image, the N+1th frame image, and so on.

[0252] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.

[0253] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.

[0254] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.

[0255] In this embodiment, the camera 193 can also be used to obtain the Nth frame image, the N+1th frame image, and so on.

[0256] Digital signal processors are used to process digital signals. In addition to processing digital image signals, they can also process other digital signals.

[0257] Video codecs are used to compress or decompress digital video.

[0258] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension.

[0259] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0260] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, applications required for at least one function (such as face recognition function, fingerprint recognition function, mobile payment function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as face information template data, fingerprint information template, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0261] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0262] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals.

[0263] Speaker 170A, also called "horn", is used to convert audio electrical signals into sound signals.

[0264] The receiver 170B, also called the "earpiece", is used to convert audio electrical signals into sound signals.

[0265] Microphone 170C, also known as a "microphone" or "speaker," converts sound signals into electrical signals. When making a call or sending a voice message, a user can place their mouth close to microphone 170C and speak, inputting the sound signal into microphone 170C. Headphone jack 170D is used to connect a wired headset. Headphone jack 170D can be a USB port 130, or a 3.5mm Open Mobile Terminal Platform (OMTP) standard port, or a Cellular Telecommunications Industry Association of the USA (CTIA) standard port.

[0266] The pressure sensor 180A is used to sense pressure signals and convert the pressure signals into electrical signals.

[0267] The gyro sensor 180B may be used to determine the motion posture of the electronic device 100 .

[0268] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates the altitude using the air pressure value measured by the air pressure sensor 180C to assist in positioning and navigation.

[0269] The magnetic sensor 180D includes a Hall sensor, and the electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip leather case.

[0270] Accelerometer 180E can detect the magnitude of acceleration of electronic device 100 in all directions (generally three axes). It can also detect the magnitude and direction of gravity when electronic device 100 is stationary. It can also be used to identify the electronic device's posture, enabling applications such as switching between landscape and portrait modes and pedometers.

[0271] The distance sensor 180F is used to measure distance. The electronic device 100 can measure distance by infrared or laser.

[0272] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared light emitting diode. The electronic device 100 emits infrared light outward through the LED.

[0273] Ambient light sensor 180L is used to sense ambient light brightness. Electronic device 100 can adaptively adjust the brightness of display screen 194 based on the perceived ambient light. Ambient light sensor 180L can also be used to automatically adjust white balance when taking photos. Ambient light sensor 180L can also work with proximity light sensor 180G to detect whether electronic device 100 is in a pocket to prevent accidental touches.

[0274] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, fingerprint photography, fingerprint call answering, etc.

[0275] The temperature sensor 180J is used to detect temperature.

[0276] The touch sensor 180K is also called a "touch panel." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen." The touch sensor 180K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, in a location different from that of the display screen 194.

[0277] The buttons 190 include a power button, a volume button, and the like. The buttons 190 may be mechanical buttons or touch buttons. The electronic device 100 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 100.

[0278] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, motor 191 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0279] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.

[0280] The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to and disconnected from the electronic device 100 by inserting it into or removing it from the SIM card interface 195. The electronic device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communications.

[0281] Figure 30 It is a software structure block diagram of the electronic device 100 according to an embodiment of the present application.

[0282] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other via software interfaces. In some embodiments, the system is divided into four layers: application layer, application framework layer, runtime and system libraries, and kernel layer.

[0283] The application layer can include a series of application packages.

[0284] like Figure 30 As shown, the application package may include camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message and other applications (also referred to as applications).

[0285] The application framework layer provides an application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions.

[0286] like Figure 30 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.

[0287] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.

[0288] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.

[0289] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.

[0290] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including answering, hanging up, etc.).

[0291] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0292] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically, without requiring user interaction. For example, the Notification Manager can be used to notify users of completed downloads, message reminders, and so on. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog interfaces on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.

[0293] The runtime includes the core library and the virtual machine. The runtime is responsible for the scheduling and management of the system.

[0294] The core library consists of two parts: one part is the function that the programming language (for example, Java language) needs to call, and the other part is the core library of the system.

[0295] The application layer and application framework layer run in a virtual machine. The virtual machine executes the application layer and application framework layer programming files (for example, Java files) as binary files. The virtual machine performs functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0296] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.

[0297] The surface manager is used to manage the display subsystem and provide the fusion of two-dimensional (2D) and three-dimensional (3D) layers for multiple applications.

[0298] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0299] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0300] A 2D graphics engine is a drawing engine for 2D drawings.

[0301] The kernel layer is the layer between hardware and software. The kernel layer contains at least display driver, camera driver, audio driver, sensor driver, and virtual card driver.

[0302] The following describes the workflow of the software and hardware of the electronic device 100 in conjunction with capturing a photo scene.

[0303] When the touch sensor 180K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, touch operation timestamp, and other information). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. For example, if the touch operation is a touch single-click operation and the control corresponding to the single-click operation is the control of the camera application icon, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a still image or video through the camera 193.

[0304] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0305] As used in the above embodiments, the term “when…” may be interpreted to mean “if…” or “after…” or “in response to determining…” or “in response to detecting…”, depending on the context. Similarly, the phrases “upon determining…” or “if (stated condition or event) is detected” may be interpreted to mean “if determining…” or “in response to determining…” or “upon detecting (stated condition or event)” or “in response to detecting (stated condition or event)”, depending on the context.

[0306] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk).

[0307] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.< / canvas> < / video> < / videoview> < / imgview> < / textview>

Claims

1. A target tracking method, characterized in that: include: The electronic device displays a first user interface, wherein the first user interface displays the Mth frame image; The electronic device determines a tracking target in the M-th frame image according to the received user operation; The electronic device acquires a plurality of tracking target features of the tracking target; The electronic device displays a second user interface, wherein the second user interface displays a K-th image frame, where the K-th image frame is an image frame subsequent to the M-th image frame; The electronic device detects a first candidate target having the same attribute as the tracking target in the Kth frame image; The electronic device acquires a plurality of candidate target features of the first candidate target; The electronic device selects, from the multiple tracking target features and the multiple candidate target features, a first candidate target feature and a first tracking target feature that contain the same number of key points and contain the largest number of key points, the first candidate target feature belonging to the multiple candidate target features, and the first tracking target feature belonging to the multiple tracking target features; Performing feature matching on the first candidate target feature and the first tracking target feature; If the first candidate target feature matches the first tracking target feature, the electronic device determines that the first candidate target is a tracking target.

2. The method according to claim 1, characterized in that The method further comprises: If the first candidate target feature and the first tracking target feature do not match, the electronic device selects, from the multiple tracking target features and the multiple candidate target features, a fourth candidate target feature and a fourth tracking target feature that contain the same number of key points and the second largest number of key points; performing feature matching on the fourth candidate target feature and the fourth tracking target feature; If the fourth candidate target feature matches the fourth tracking target feature, the electronic device determines the first candidate target as the tracking target.

3. The method according to claim 1, characterized in that The electronic device obtains a plurality of tracking target features of the tracking target, specifically including: The electronic device acquires a plurality of tracking target images according to the tracking target, wherein the tracking target images include part or all of the tracking target; The electronic device performs feature extraction on the multiple tracking target images to obtain the multiple tracking target features, wherein the number of the multiple tracking target features is equal to the number of the multiple tracking target images.

4. The method according to claim 3, characterized in that The electronic device divides the tracking target into a plurality of tracking target images, specifically including: The electronic device obtains a plurality of tracking target images of the tracking target according to the key points of the tracking target, where the tracking target images contain one or more key points of the tracking target.

5. The method according to claim 4, characterized in that The multiple tracking target images include a first tracking target image and a second tracking target image, wherein: The first tracking target image and the second tracking target image contain the same key points of the tracking target, and the second tracking target image contains more key points of the tracking target than the first tracking target image; Alternatively, the key points of the tracking target included in the first tracking target image are different from the key points of the tracking target included in the second tracking target image.

6. The method according to claim 1, characterized in that The electronic device acquires a plurality of candidate target features of the first candidate target, including: The electronic device acquires a plurality of candidate target images according to the first candidate target; the plurality of candidate target images include part or all of the first candidate target; The electronic device performs feature extraction on the multiple candidate target images respectively to obtain the multiple candidate target features, and the number of the multiple candidate target features is equal to the number of the multiple candidate target features.

7. The method according to claim 6, characterized in that The electronic device divides the first candidate target into a plurality of candidate target images, specifically including: The electronic device obtains the multiple candidate target images according to the key points of the first candidate target, and the candidate target images contain one or more key points of the first candidate target.

8. The method according to claim 7, characterized in that The plurality of candidate target images include the first candidate target image and the second candidate target image, wherein: The first candidate target image and the second candidate target image contain the same key points of the first candidate target, and the second candidate target image contains more key points of the first candidate target than the first candidate target image; Alternatively, the key points of the first candidate target included in the first candidate target image are different from the key points of the first candidate target included in the second candidate target image.

9. The method according to claim 1, characterized in that The plurality of tracking target images include a first tracking target image and a second tracking target image, and the plurality of tracking target images are acquired from the tracking target; the plurality of candidate target images include a first candidate target image and a second candidate target image, and the plurality of candidate target images are acquired from the first candidate target; The number of key points of the tracking target contained in the first tracking target image is the same as the number of key points of the first candidate target contained in the first candidate target image; The number of key points of the tracking target contained in the second tracking target image is the same as the number of key points of the first candidate target contained in the second candidate target image; The number of the tracking target key points included in the first tracking target image is greater than the number of the tracking target key points included in the second tracking target image; The first tracking target feature is extracted from the first tracking target image, and the first candidate target feature is extracted from the first candidate target image.

10. The method according to any one of claims 1 to 9, characterized in that When a first candidate target feature among the multiple candidate target features matches a first tracking target feature among the multiple tracking target features, after the electronic device determines that the first candidate target is a tracking target, the method further includes: The electronic device saves the second candidate target into a feature library that stores the multiple tracking target features; the second candidate feature is extracted by the electronic device from a third candidate target image among the multiple candidate target images; the number of key points of the first candidate target contained in the third candidate target image is greater than the number of key points of the tracking target contained in the multiple tracking target images.

11. The method according to claim 9, characterized in that When a first candidate target feature among the multiple candidate target features matches a first tracking target feature among the multiple tracking target features, after the electronic device determines that the first candidate target is a tracking target, the method further includes: If the difference between M and K is equal to a preset threshold, the electronic device saves the first candidate target feature into a feature library storing the plurality of tracking target features.

12. The method according to claim 11, characterized in that The electronic device saving the first candidate target feature into a feature library storing the plurality of tracking target features specifically includes: The electronic device replaces the first tracking target feature in the feature library storing the plurality of tracking target features with the first candidate target feature.

13. The method according to claim 1, wherein The method further comprises: The electronic device detects the tracking target in the M-th frame image; The electronic device displays a detection frame in the first user interface, where the detection frame is used to encircle the tracking target; The electronic device receives a first user operation, where the first user operation is used to select the tracking target in the first user interface.

14. An electronic device, characterized in that: include: Display, processor, memory; The memory is coupled to the processor; The display screen is coupled to the processor, wherein: The display screen is used to display a first user interface, in which the Mth frame image is displayed; and to display a second user interface, in which the Kth frame image is displayed; The processor is configured to determine a tracking target in the M-th frame image according to a received user operation; detect a first candidate target in the K-th frame image, the first candidate target having the same attribute as the tracking target; obtain multiple tracking target features of the tracking target; obtain multiple candidate target features of the first candidate target; select a first candidate target feature and a first tracking target feature that contain the same number of key points and the largest number of key points from the multiple tracking target features and the multiple candidate target features, the first candidate target feature belonging to the multiple candidate target features, and the first tracking target feature belonging to the multiple tracking target features; perform feature matching on the first candidate target feature and the first tracking target feature; and determine that the first candidate target is the tracking target if the first candidate target feature and the first tracking target feature match. The memory is used to store the multiple tracking target features.

15. The electronic device according to claim 14, characterized in that The processor is further configured to: If the first candidate target feature and the first tracking target feature do not match, selecting a fourth candidate target feature and a fourth tracking target feature that contain the same number of key points and the second largest number of key points from the multiple tracking target features and the multiple candidate target features; performing feature matching on the fourth candidate target feature and the fourth tracking target feature; If the fourth candidate target feature matches the fourth tracking target feature, the first candidate target is determined as the tracking target.

16. The electronic device according to claim 14, characterized in that The processor is specifically configured to: Acquire a plurality of tracking target images according to the tracking target, wherein the tracking target images include part or all of the tracking target; Feature extraction is performed on the multiple tracking target images to obtain the multiple tracking target features, wherein the number of the multiple tracking target features is equal to the number of the multiple tracking target images.

17. The electronic device according to claim 16, wherein: The processor is specifically configured to: The plurality of tracking target images are acquired according to the key points of the tracking target, wherein the tracking target images contain one or more key points of the tracking target.

18. The electronic device according to claim 17, wherein: The multiple tracking target images include a first tracking target image and a second tracking target image, wherein: The first tracking target image and the second tracking target image contain the same key points of the tracking target, and the second tracking target image contains more key points of the tracking target than the first tracking target image; Alternatively, the key points of the tracking target included in the first tracking target image are different from the key points of the tracking target included in the second tracking target image.

19. The electronic device according to claim 14, wherein: The processor is configured to: Acquire multiple candidate target images according to the first candidate target; the multiple candidate target images include part or all of the first candidate target; Feature extraction is performed on each of the plurality of candidate target images to obtain the plurality of candidate target features, where the number of the plurality of candidate target features is equal to the number of the plurality of candidate target features.

20. The electronic device according to claim 19, wherein The processor is specifically configured to: The plurality of candidate target images are acquired according to the key points of the first candidate target, where the candidate target images contain one or more key points of the first candidate target.

21. The electronic device according to claim 20, characterized in that The plurality of candidate target images include the first candidate target image and the second candidate target image, wherein: The first candidate target image and the second candidate target image contain the same key points of the first candidate target, and the second candidate target image contains more key points of the first candidate target than the first candidate target image; Alternatively, the key points of the first candidate target included in the first candidate target image are different from the key points of the first candidate target included in the second candidate target image.

22. The electronic device according to claim 14, wherein: The plurality of tracking target images include a first tracking target image and a second tracking target image, and the plurality of tracking target images are acquired from the tracking target; the plurality of candidate target images include a first candidate target image and a second candidate target image, and the plurality of candidate target images are acquired from the first candidate target; The number of key points of the tracking target contained in the first tracking target image is the same as the number of key points of the first candidate target contained in the first candidate target image; The number of key points of the tracking target contained in the second tracking target image is the same as the number of key points of the first candidate target contained in the second candidate target image; The number of the tracking target key points included in the first tracking target image is greater than the number of the tracking target key points included in the second tracking target image; The first tracking target feature is extracted from the first tracking target image, and the first candidate target feature is extracted from the first candidate target image.

23. The electronic device according to any one of claims 14 to 22, characterized in that: The memory is used for: The second candidate target is saved in a feature library that stores the multiple tracking target features; the second candidate feature is extracted by the processor from a third candidate target image among the multiple candidate target images; the number of key points of the first candidate target contained in the third candidate target image is greater than the number of key points of the tracking target contained in the multiple tracking target images.

24. The electronic device according to claim 22, wherein: The memory is used for: If the difference between the M and the K is equal to a preset threshold, the first candidate target feature is saved in a feature library storing the plurality of tracking target features.

25. The electronic device according to claim 24, characterized in that The memory is used for: The first tracking target feature in the feature library storing the plurality of tracking target features is replaced with the first candidate target feature.

26. A computer storage medium, characterized in that The method comprises computer instructions, which, when executed on an electronic device, cause the electronic device to execute the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Subblock weight Mean-Shift tracking method with improved level set target extraction

    CN103903280A

  • Target tracking method based on self-adaptive blocks of video

    CN104820996A

  • Multi-target tracking method and system based on semantic segmentation

    CN110298248A

Cited By

  • Target tracking method and electronic device

    WO2022068522A1