Target Tracking Method and Apparatus

By using the depth information of the video stream, determining the changing areas between image frames as the tracking target, the problem of difficult to track small moving objects in the prior art is solved, and high-accurate target tracking is achieved.

CN115147451BActive Publication Date: 2025-06-27HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110336639.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-29
Publication Date
2025-06-27
Estimated Expiration
2041-03-29

AI Technical Summary

Technical Problem

Existing target tracking techniques are difficult to effectively select and track too small moving objects, and tracking failures are likely to occur during the movement of objects.

Method used

By acquiring the depth information of the image frames in the video stream, it is determined that the change area between adjacent image frames is the object to be detected, and a tracking target is selected based on the displacement value and displacement direction of the depth information.

Benefits of technology

It realizes convenient selection of tracking targets, can effectively track too small objects in motion, and improves the accuracy and stability of target tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147451B_ABST
    Figure CN115147451B_ABST
Patent Text Reader

Abstract

The present application discloses a target tracking method and apparatus thereof, relating to the field of big data, and is used to conveniently select a tracking target. The target tracking method includes: obtaining depth information of an image frame in a video stream; determining a change region between first adjacent image frames in the video stream as an object to be detected, the first adjacent image frames including a first image frame and a second image frame, the first image frame being the image frame in front of the second image frame, and the change region being a difference region of the depth information; determining a displacement value and a displacement direction between the position of the object to be detected in the second image frame and the position of the object to be detected in the first image frame, the position being the position of the depth information, and the displacement value and the displacement direction being the displacement value and the displacement direction of the depth information; and selecting a tracking target, where the tracking target is an object to be detected with a displacement value greater than a first preset value and a displacement direction being a direction close to the focal plane. The embodiments of the present application are applied to data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing, and in particular, to a target tracking method and apparatus thereof. Background Art

[0002] Target tracking is an important research topic in the field of computer vision and is currently widely applied in related fields such as video live broadcast, security monitoring, robotics, and human-computer interaction. Target tracking takes an image sequence as input according to a selected tracking target and outputs the size, position, etc. of the selected tracking target in each frame of the image sequence. The accuracy of target tracking depends on the selected tracking target. Therefore, the selection of the tracking target is a key step in triggering target tracking. To select a tracking target, currently, a target detection model (such as the Yolo model) can be used to identify multiple objects in an image and output detection frames to mark the positions of the objects. Then, the object within the target frame selected by the user's click is selected as the tracking target. However, the Yolo model cannot detect too small objects in the image, which may lead to failure in positioning and tracking the target. Although it is also possible to select a tracking target by manually drawing an image frame on the image to mark the object, for a moving object, the object may be at a first position in a certain frame (such as the first frame) of the image sequence when starting to draw the image frame. As the drawing progresses and the object moves, the object has moved out of the original position in another frame (such as the tenth frame) of the image sequence, which also results in failure in positioning and tracking the target. Summary of the Invention

[0003] In view of the above, embodiments of this application provide a target tracking method and apparatus thereof, which can conveniently select a tracking target.

[0004] In a first aspect, an embodiment of this application provides a target tracking method, the method including: obtaining depth information of an image frame in a video stream; determining a change region between a first adjacent image frame in the video stream as an object to be detected, the first adjacent image frame including a first image frame and a second image frame, the first image frame being an image frame in front of the second image frame, the change region being a difference region of depth information; determining a displacement value and a displacement direction between the position of the object to be detected in the second image frame and the position of the object to be detected in the first image frame, the position being the position of depth information, the displacement value and the displacement direction being the displacement value and the displacement direction of depth information; selecting a tracking target, the tracking target being an object to be detected with a displacement value greater than a first preset value and a displacement direction being a direction close to the focal plane.

[0005] This application determines the object to be detected as a difference region of depth information between adjacent image frames, and when the position of the depth information of the object to be detected significantly moves forward, selects the object to be detected as the tracking target, which can conveniently select the tracking target.

[0006] According to some embodiments of the present application, the method further includes: during tracking, determining the position of the tracking target in the second adjacent image frames, where the second adjacent image frames include a third image frame and a fourth image frame, the third image frame is the image frame in front of the fourth image frame, and the position is the position of the depth information; determining the displacement value and displacement direction between the position of the tracking target in the fourth image frame and the position of the tracking target in the third image frame, where the displacement value and the displacement direction are the displacement value and displacement direction of the depth information; if the displacement value is greater than a second preset value and the displacement direction is away from the focal plane, exiting the tracking.

[0007] When the position of the depth information of the tracking target significantly moves backward, the present application exits the tracking, so that the tracking can be conveniently exited.

[0008] According to some embodiments of the present application, the method further includes: detecting human body key points in the image frames of the video stream; where the object to be detected is the change region connected to the first parameter of the human body key points between the first adjacent image frames in the video stream; the tracking target is the object to be detected with a displacement value greater than a first preset value, a displacement direction towards the focal plane, and a depth information difference greater than a third preset value between the second parameter of the human body key points in the image frame.

[0009] When the position of the depth information of the object to be detected significantly moves forward and is connected to the first parameter of the human body key points, the present application selects the object to be detected as the tracking target, so that the tracking target can be conveniently selected.

[0010] According to some embodiments of the present application, the method further includes: during tracking, determining the position of the tracking target in the second adjacent image frames, where the second adjacent image frames include a third image frame and a fourth image frame, the third image frame is the image frame in front of the fourth image frame, and the position is the position of the depth information; determining the displacement value and displacement direction between the position of the tracking target in the fourth image frame and the position of the tracking target in the third image frame, where the displacement value and the displacement direction are the displacement value and displacement direction of the depth information; if the displacement value is greater than a second preset value and the displacement direction is away from the focal plane, or the tracking target is not connected to the first parameter of the human body key points, exiting the tracking.

[0011] When the position of the depth information of the tracking target significantly moves backward or is no longer connected to the first parameter of the human body key points, the present application exits the tracking, so that the tracking can be conveniently exited.

[0012] According to some embodiments of the present application, the displacement value of the depth information is the absolute value of the average depth change value of the pixel points.

[0013] In a second aspect, an embodiment of the present application provides a target tracking device, the device comprising: an acquisition unit configured to acquire depth information of an image frame in a video stream; a determination unit configured to determine a change region between first adjacent image frames in the video stream as an object to be detected, the first adjacent image frames including a first image frame and a second image frame, the first image frame being an image frame in front of the second image frame, and the change region being a difference region of depth information; the determination unit is further configured to determine a displacement value and a displacement direction between the position of the object to be detected in the second image frame and the position of the object to be detected in the first image frame, the position being the position of depth information, and the displacement value and the displacement direction being the displacement value and the displacement direction of depth information; the determination unit is further configured to select a tracking target, the tracking target being an object to be detected with a displacement value greater than a first preset value and a displacement direction towards the focal plane.

[0014] According to some embodiments of the present application, the determination unit is further configured to, during tracking, determine the position of the tracking target in second adjacent image frames, the second adjacent image frames including a third image frame and a fourth image frame, the third image frame being an image frame in front of the fourth image frame, and the position being the position of depth information; the determination unit is further configured to determine a displacement value and a displacement direction between the position of the tracking target in the fourth image frame and the position of the tracking target in the third image frame, the displacement value and the displacement direction being the displacement value and the displacement direction of depth information; the determination unit is further configured to exit tracking if the displacement value is greater than a second preset value and the displacement direction is away from the focal plane.

[0015] According to some embodiments of the present application, the determination unit is further configured to detect human key points in an image frame of the video stream; the determination unit is further configured to determine a change region connected to a first parameter of the human key points between first adjacent image frames in the video stream as an object to be detected; the determination unit is further configured to select a tracking target, the tracking target being an object to be detected with a displacement value greater than a first preset value, a displacement direction towards the focal plane, and a depth information difference greater than a third preset value between the object to be detected and a second parameter of the human key points in the image frame.

[0016] According to some embodiments of the present application, the determination unit is further configured to, during tracking, determine the position of the tracking target in second adjacent image frames, the second adjacent image frames including a third image frame and a fourth image frame, the third image frame being an image frame in front of the fourth image frame, and the position being the position of depth information; the determination unit is further configured to determine a displacement value and a displacement direction between the position of the tracking target in the fourth image frame and the position of the tracking target in the third image frame, the displacement value and the displacement direction being the displacement value and the displacement direction of depth information; the determination unit is further configured to exit tracking if the displacement value is greater than a second preset value and the displacement direction is away from the focal plane, or if the tracking target is not connected to the first parameter of the human key points.

[0017] According to some embodiments of the present application, the displacement value of the depth information is the absolute value of the average depth change value of the pixel points.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory is used to store program instructions. When the processor calls the program instructions, the target tracking method described in any one of the above is implemented.

[0019] In a fourth aspect, an embodiment of the present application provides a server, which includes a processor and a memory. The memory is used to store program instructions. When the processor calls the program instructions, the target tracking method described in any one of the above is implemented.

[0020] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a program. The program enables a computer device to implement the target tracking method described in any one of the above.

[0021] In a sixth aspect, an embodiment of the present application provides a computer program product, which includes computer-executable instructions. The computer-executable instructions are stored in a computer-readable storage medium; at least one processor of the device can read the computer-executable instructions from the computer-readable storage medium, and the at least one processor executes the computer-executable instructions to enable the device to execute the target tracking method described in any one of the above.

[0022] For the beneficial effects of the second to sixth aspects and their various implementation manners, reference can be made to the analysis of the beneficial effects in the first aspect and its various implementation manners, which will not be elaborated here. Description of the Drawings

[0023] Figure 1 It is a schematic diagram of the tracking system according to the embodiment of the present application.

[0024] Figure 2 It is a schematic diagram of the hardware structure of the electronic device according to the embodiment of the present application.

[0025] Figure 3 It is a block diagram of the software structure of the electronic device according to the embodiment of the present application.

[0026] Figure 4 It is a flowchart of the target tracking method according to the embodiment of the present application.

[0027] Figures 5A - 5D It is a human-computer interaction interface diagram provided by the embodiment of the present application.

[0028] Figures 6A - 6B It is some other human-computer interaction interface diagrams provided by the embodiment of the present application.

[0029] Figures 7A - 7D These are some other human-computer interaction interface diagrams provided by the embodiments of the present application.

[0030] Figures 8A - 8B These are some other human-computer interaction interface diagrams provided by the embodiments of the present application.

[0031] Figures 9A - 9B These are some schematic diagrams provided by the embodiments of the present application.

[0032] Figures 10A - 10E This is the user interface provided by the embodiments of the present application.

[0033] Figures 11A - 11B These are some other user interfaces provided by the embodiments of the present application.

[0034] Figure 12 This is the schematic diagram of the human key points provided by the embodiments of the present application.

[0035] Figures 13A - 13B These are some other user interfaces provided by the embodiments of the present application.

[0036] Figure 14 This is the schematic diagram of the hardware structure of the server provided by the embodiments of the present application.

[0037] Figure 15 This is the schematic diagram of the structure of the target tracking device of the present application. Detailed implementation manners

[0038] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the embodiments of the present application, words such as "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "for example" aims to present relevant concepts in a specific manner.

[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the description of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. It should be understood that unless otherwise stated in this application, "a plurality of" means two or more than two.

[0040] Reference Figure 1As shown in the figure, it is a schematic diagram of the tracking system according to an embodiment of the present application. The tracking system 10 may include an electronic device 11 and a server 12. In this embodiment, the electronic device 11 may be an electronic device such as a smart phone, a tablet computer, a PDA (Personal Digital Assistant), a smart camera device, or a wearable device with an image capture function. A network connection may be established between the electronic device 11 and the server 12. The network connection may be a wired or wireless connection. The electronic device 11 may include an imaging module 111. The imaging module 111 may be an imaging module such as a binocular camera, a structured light camera, a TOF (Time of flight) camera, or an ordinary monocular camera. The imaging module 111 is used to capture an image of a scene. The image can be used to obtain the depth information of the object being photographed. If the imaging module 111 is a binocular camera, a structured light camera, or a TOF (Time of flight) camera, the depth information of the object being photographed is included in the image, and the depth information of the object being photographed in the image can be directly obtained subsequently. If the imaging module 111 is an ordinary monocular camera, a monocular depth estimation algorithm can be used subsequently to obtain the depth information of the object being photographed in the image. The imaging module 111 captures the image at a fixed frequency, for example, 30 frames per second. The imaging module 111 may be fixed to capture images within the same scene, or may be driven to move to track an object. The electronic device 11 includes a client 112. The client 112 may be an application program with an imaging function running on the electronic device 11, such as a camera application APP, an APP for live streaming with goods, an APP for providing video calls, or an APP for monitoring applications. The client 112 may call the camera application APP through an application programming interface (API) to request permission to call the imaging module 111, and after obtaining the permission, may control the call of the imaging module 111. The electronic device 11 may obtain the video stream captured by the imaging module 111 and send the video stream to the server 12 through the client 112. The server 12 may store the video stream at a storage location associated with the live channel identifier for a playback end to play the video stream or send the video stream to other electronic devices for video calls.

[0041] In this embodiment, the electronic device 11 can obtain the depth information of the image frames in the video stream; determine the change region between adjacent image frames in the video stream as the object to be detected, where the adjacent image frames include a first image frame and a second image frame, the first image frame is the image frame in front of the second image frame, and the change region is the depth information difference region; determine the displacement value and displacement direction between the position of the object to be detected in the second image frame and the position of the object to be detected in the first image frame, where the displacement value and the displacement direction are the displacement value and displacement direction of the depth information; select a tracking target, where the tracking target is the object to be detected with a displacement value greater than a first preset value and a displacement direction towards the focal plane. The electronic device 11 also tracks the tracking target and then sends the processed video stream to the server 12.

[0042] To reduce the computational load of the electronic device, the processing of the video stream can also be performed by the server 12. Specifically, after the server receives the video stream sent by the electronic device through the client, it can also obtain the depth information of the image frames in the video stream; determine the change region between adjacent image frames in the video stream as the object to be detected, where the adjacent image frames include a first image frame and a second image frame, the first image frame is the image frame in front of the second image frame, and the change region is the depth information change region; determine the displacement value and displacement direction between the position of the object to be detected in the second image frame and the position of the object to be detected in the first image frame, where the displacement value and the displacement direction are the displacement value and displacement direction of the depth information; select a tracking target, where the tracking target is the object to be detected with a displacement value greater than a first preset value and a displacement direction towards the focal plane. The server 12 can also track the tracking target through the electronic device and then send the processed video stream to other electronic devices through the client. That is, in the embodiments of the present application, selecting a tracking target and tracking the tracking target can be implemented in the electronic device 11 or in the server 12, and this is not limited herein.

[0043] Reference Figure 2 As shown, it is a schematic diagram of the hardware structure of the electronic device according to the embodiments of the present application. The electronic device 100 may include at least one of a mobile phone with an image capture function, a foldable electronic device, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), a wearable device, a vehicle-mounted device, or a smart home device. The embodiments of the present application do not impose any special restrictions on the specific type of the electronic device 100.

[0044] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) connector 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0045] It can be understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0046] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0047] The processor may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0048] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 may be a cache memory. This memory may store instructions or data that have been used by the processor 110 or are used frequently. If the processor 110 needs to use this instruction or data, it can be directly called from this memory. This avoids repeated accesses and reduces the waiting time of the processor 110, thus improving the efficiency of the system.

[0049] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc. The processor 110 may be connected to modules such as a touch sensor, an audio module, a wireless communication module, a display, a camera, etc. through at least one of the above interfaces.

[0050] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0051] The USB connector 130 is an interface that complies with the USB standard specifications and can be used to connect the electronic device 100 and peripheral devices. Specifically, it can be a Mini USB connector, a Micro USB connector, a USB Type C connector, etc. The USB connector 130 can be used to connect a charger to charge the electronic device 100, or to connect other electronic devices to transfer data between the electronic device 100 and other electronic devices. It can also be used to connect headphones to output the audio stored in the electronic device through the headphones. This connector can also be used to connect other electronic devices, such as VR devices, etc. In some embodiments, the standard specifications of the Universal Serial Bus can be USB1.x, USB2.0, USB3.x, and USB4.

[0052] The charging management module 140 is used to receive the charging input from the charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 can receive the charging input from the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 can receive the wireless charging input through the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device through the power management module 141.

[0053] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives the input from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be set in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.

[0054] The wireless communication function of the electronic device 100 can be implemented through antenna 1, antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0055] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: Antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0056] The mobile communication module 150 may provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 may receive electromagnetic waves through the antenna 1, filter, amplify, and perform other processing on the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 may also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 150 may be provided in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be provided in the same device.

[0057] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, receiver 170B, etc.), or displays an image or video through the display screen 194. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 110 and be provided in the same device as the mobile communication module 150 or other functional modules.

[0058] The wireless communication module 160 may provide wireless communication solutions applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), Bluetooth low energy (BLE), ultra wide band (UWB), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive the signals to be sent from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation.

[0059] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, such that electronic device 100 can communicate with a network and other electronic devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).

[0060] Electronic device 100 may implement a display function through a GPU, display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, and is connected to display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0061] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or more display screens 194.

[0062] The electronic device 100 can implement the camera function through the camera module 193, ISP, video codec, GPU, display screen 194, application processor AP, neural network processor NPU, etc.

[0063] The camera module 193 can be used to collect color image data and depth data of the photographed object. The ISP can be used to process the color image data collected by the camera module 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera sensor. The light signal is converted into an electrical signal, and the camera sensor transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin color of the image through algorithms. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera module 193.

[0064] In some embodiments, the camera module 193 can be composed of a color camera module and a 3D sensing module.

[0065] In some embodiments, the photosensitive element of the camera of the color camera module can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats.

[0066] In some embodiments, the 3D sensing module may be a (time of flight, TOF) 3D sensing module or a structured light 3D sensing module. Among them, structured light 3D sensing is an active depth sensing technology. The basic components of the structured light 3D sensing module may include an infrared (Infrared) emitter, an IR camera module, etc. The working principle of the structured light 3D sensing module is to first emit a light spot with a specific pattern to the object to be photographed, then receive the light spot pattern coding on the surface of the object, and then compare the similarities and differences with the original projected light spot, and use the triangulation principle to calculate the three-dimensional coordinates of the object. The three-dimensional coordinates include the distance between the electronic device 100 and the object to be photographed. Among them, TOF 3D sensing can be an active depth sensing technology. The basic components of the TOF 3D sensing module may include an infrared (Infrared) emitter, an IR camera module, etc. The working principle of the TOF 3D sensing module is to calculate the distance (i.e., depth) between the TOF 3D sensing module and the object to be photographed by the time of infrared return, so as to obtain a 3D depth of field map.

[0067] The structured light 3D sensing module can also be applied to fields such as face recognition, somatosensory game consoles, and industrial machine vision detection. The TOF 3D sensing module can also be applied to fields such as game consoles, augmented reality (AR) / virtual reality (VR), etc.

[0068] In some other embodiments, the camera module 193 may also be composed of two or more cameras. These two or more cameras may include a color camera, and the color camera can be used to collect color image data of the object to be photographed. These two or more cameras may use stereo vision technology to collect depth data of the object to be photographed. Stereo vision technology is based on the principle of human eye parallax. Under natural light, images of the same object are taken from different angles through two or more cameras, and then operations such as triangulation are performed to obtain the distance information between the electronic device 100 and the object to be photographed, that is, depth information.

[0069] In some other embodiments, the camera module 193 may also be composed of one camera. This camera takes an RGB image from one or the only perspective. The GPU in the processor 110 can estimate the distance of each pixel in the image relative to the camera module 193 according to the monocular depth estimation algorithm, that is, depth information.

[0070] In some embodiments, the camera module 193 can be fixed to capture images of the same scene and the same perspective, or can be driven to capture images of different scenes. The camera module 193 can be fixed before a tracking target is selected; after the tracking target is selected, it can be driven to perform target tracking.

[0071] In some embodiments, the electronic device 100 may include one or more camera modules 193. Specifically, the electronic device 100 may include one front camera module 193 and one rear camera module 193. Among them, the front camera module 193 is generally used to capture the color image data and depth data of the photographer himself facing the display screen 194, and the rear camera module is used to capture the color image data and depth data of the photographed object (such as a person, a landscape, etc.) faced by the photographer.

[0072] In some embodiments, the CPU, GPU, or NPU in the processor 110 can process the color image data and depth data captured by the camera module 193. In some embodiments, the NPU can identify the color image data captured by the camera module 193 (specifically, the color camera module) through a neural network algorithm based on the skeleton point recognition technology, such as the convolutional neural network algorithm (CNN), to determine the skeleton points of the photographed person. The CPU or GPU can also run the neural network algorithm to determine the skeleton points of the photographed person according to the color image data. In some embodiments, the CPU, GPU, or NPU can also be used to confirm the figure of the photographed person (such as body proportions, the fatness or thinness of the body parts between the skeleton points) based on the depth data captured by the camera module 193 (which can be a 3D sensing module) and the identified skeleton points, and can further determine the body beautification parameters for the photographed person. Finally, the captured image of the photographed person is processed according to the body beautification parameters so that the figure of the photographed person in the captured image is beautified. How to perform body beautification processing on the image of the photographed person based on the color image data and depth data captured by the camera module 193 will be introduced in detail in subsequent embodiments and will not be elaborated here.

[0073] The digital signal processor is used to process digital signals and can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0074] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0075] The NPU is a neural-network (NN) computing processor. By drawing on the structure of the biological neural network, such as the transmission pattern between human brain neurons, it can quickly process the input information and can also continuously learn by itself. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.

[0076] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to achieve the data storage function. For example, files such as music and videos are saved in the external memory card. Or files such as music and videos are transferred from the electronic device to the external memory card.

[0077] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.). The data storage area can store the data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional methods or data processing of the electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.

[0078] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. Such as music playback, recording, etc.

[0079] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0080] The speaker 170A, also called the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music through the speaker 170A, or output the audio signal of a hands-free call.

[0081] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the voice can be received by bringing the receiver 170B close to the human ear.

[0082] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 may be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 may be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 may also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, implement a directional recording function, etc.

[0083] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D may be a USB interface 130, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0084] The pressure sensor 180A is used to sense a pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A may be disposed on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor may include at least two parallel plates with conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 194, the electronic device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 180A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities may correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.

[0085] The gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake during shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 180B detects the shaking angle of the electronic device 100, calculates the distance that the lens module needs to compensate according to the angle, and controls the lens to move in the opposite direction to offset the shaking of the electronic device 100, thereby achieving anti-shake. The gyroscope sensor 180B can also be used for navigation and somatosensory game scenarios.

[0086] The barometric pressure sensor 180C is used to measure the barometric pressure. In some embodiments, the electronic device 100 calculates the altitude according to the barometric pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.

[0087] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip leather case. When the electronic device is a foldable electronic device, the magnetic sensor 180D can be used to detect the folding or unfolding of the electronic device, or the folding angle. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip according to the magnetic sensor 180D. Furthermore, according to the detected opening and closing state of the leather case or the flip, features such as automatic flip unlocking can be set.

[0088] The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.

[0089] The distance sensor 180F is used to measure the distance. The electronic device 100 can measure the distance by infrared or laser. In some embodiments, in the shooting scene, the electronic device 100 can use the distance sensor 180F to measure the distance to achieve rapid focusing.

[0090] The proximity light sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The light-emitting diode may be an infrared light-emitting diode. The electronic device 100 emits infrared light outward through the light-emitting diode. The electronic device 100 uses the photodiode to detect the infrared reflected light from a nearby object. When the intensity of the detected reflected light is greater than a threshold value, it may be determined that there is an object near the electronic device 100. When the intensity of the detected reflected light is less than the threshold value, the electronic device 100 may determine that there is no object near the electronic device 100. The electronic device 100 may use the proximity light sensor 180G to detect that the user holds the electronic device 100 close to the ear during a call, so as to automatically turn off the screen to achieve the purpose of power saving. The proximity light sensor 180G can also be used in the holster mode and the pocket mode for automatic unlocking and locking of the screen.

[0091] The ambient light sensor 180L can be used to sense the ambient light brightness. The electronic device 100 can adaptively adjust the brightness of the display screen 194 according to the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance during photography. The ambient light sensor 180L can also cooperate with the proximity light sensor 180G to detect whether the electronic device 100 is blocked, for example, when the electronic device is in the pocket. When it is detected that the electronic device is blocked or in the pocket, some functions (such as the touch function) can be disabled to prevent accidental operations.

[0092] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access to the application lock, fingerprint photography, fingerprint answering of incoming calls, etc.

[0093] The temperature sensor 180J is used to detect the temperature. In some embodiments, the electronic device 100 executes a temperature processing strategy using the temperature detected by the temperature sensor 180J. For example, when the temperature detected by the temperature sensor 180J exceeds a threshold value, the electronic device 100 reduces the performance of the processor in order to reduce the power consumption of the electronic device to implement thermal protection. In other embodiments, when the temperature detected by the temperature sensor 180J is lower than another threshold value, the electronic device 100 heats the battery 142. In still other embodiments, when the temperature is lower than yet another threshold value, the electronic device 100 can boost the output voltage of the battery 142.

[0094] The touch sensor 180K, also known as the "touch control device". The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 together form a touch screen, also known as the "touch control screen". The touch sensor 180K is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a different position from that of the display screen 194.

[0095] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire the vibration signals of the vibrating bone mass of the human vocal part. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure pulsation signals. In some embodiments, the bone conduction sensor 180M can also be disposed in the earphone to form a bone conduction earphone. The audio module 170 can parse out voice signals based on the vibration signals of the vibrating bone mass of the vocal part acquired by the bone conduction sensor 180M to implement the voice function. The application processor can parse out heart rate information based on the blood pressure pulsation signals acquired by the bone conduction sensor 180M to implement the heart rate detection function.

[0096] The button 190 can include a power-on button, a volume button, etc. The button 190 can be a mechanical button or a touch button. The electronic device 100 can receive button inputs and generate key signal inputs related to the user settings and function controls of the electronic device 100.

[0097] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, playing audio, etc.) can correspond to different vibration feedback effects. Touch operations acting on different areas of the display screen 194 can also correspond to different vibration feedback effects for the motor 191. Different application scenarios (such as time reminder, receiving information, alarm clock, game, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0098] The indicator 192 can be an indicator light and can be used to indicate the charging state, the change in battery power, and can also be used to indicate messages, missed calls, notifications, etc.

[0099] The SIM card interface 195 is used to connect to a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact with and separation from the electronic device 100. The electronic device 100 can support one or more SIM card interfaces. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The types of multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0100] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of this application, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device 100.

[0101] Figure 3 It is the software structure block diagram of the electronic device 100 in the embodiments of this application.

[0102] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom, namely the application layer, the application framework layer, Android runtime (ART) and native C / C++ libraries, the Hardware Abstract Layer (HAL), and the kernel layer.

[0103] The application layer can include a series of application packages.

[0104] As Figure 3 shown, the application packages can include applications such as the camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and short message.

[0105] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0106] As Figure 3As shown, the application framework layer may include a window manager, a content provider, a view system, a resource manager, a notification manager, an activity manager, an input manager, etc.

[0107] The window manager provides the Window Manager Service (WMS). The WMS can be used for window management, window animation management, surface management, and as a transfer station for the input system.

[0108] The content provider is used to store and obtain data, and make this data accessible to applications. This data can include videos, images, audio, incoming and outgoing calls, browsing history and bookmarks, phone books, etc.

[0109] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon can include a view for displaying text and a view for displaying pictures.

[0110] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc.

[0111] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is complete, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background-running application, or a notification that appears in the form of a dialog window on the screen. For example, it can prompt text information in the status bar, emit a prompt sound, vibrate the electronic device, blink the indicator light, etc.

[0112] The activity manager can provide the Activity Manager Service (AMS). The AMS can be used for the startup, switching, scheduling of system components (such as activities, services, content providers, broadcast receivers), and the management and scheduling of application processes.

[0113] The input manager can provide the Input Manager Service (IMS). The IMS can be used to manage system inputs, such as touch screen input, key input, sensor input, etc. The IMS retrieves events from input device nodes and distributes the events to appropriate windows through interaction with the WMS.

[0114] The Android Runtime includes the core libraries and the Android Runtime. The Android Runtime is responsible for converting source code into machine code. The Android Runtime mainly includes the Ahead-of-Time (AOT) compilation technology and the Just-in-Time (JIT) compilation technology.

[0115] The core libraries are mainly used to provide the functions of basic Java class libraries, such as libraries for basic data structures, mathematics, IO, tools, databases, networks, etc. The core libraries provide APIs for users to develop Android applications.

[0116] The native C / C++ libraries can include multiple functional modules. For example: Surface Manager, Media Framework, libc, OpenGL ES, SQLite, Webkit, etc.

[0117] Among them, the Surface Manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications. The Media Framework supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media libraries can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc. OpenGL ES provides the drawing and operation of 2D and 3D graphics in applications. SQLite provides a lightweight relational database for the applications of the electronic device 100.

[0118] The Hardware Abstraction Layer runs in the user space, encapsulates the kernel layer drivers, and provides call interfaces to the upper layer.

[0119] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.

[0120] Next, in combination with the scenario of capturing a photo, the working processes of the software and hardware of the electronic device will be exemplarily described.

[0121] When the touch sensor 180K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking the touch operation as a touch click operation and the control corresponding to the click operation as the control of the camera application icon as an example, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a static image or video through the camera.

[0122] Please refer toFigure 4 , which is a flowchart of a target tracking method according to an embodiment of the present application. The target tracking method is applied to an electronic device and selects and tracks a tracking target according to changes in depth information. The target tracking method includes:

[0123] S401: The electronic device receives an operation to turn on the client on the electronic device.

[0124] For ease of description, the present application will be described below taking a mobile phone as an example. For example, the user can click on the application icon of the "video call" client or the "live streaming with goods" client on the mobile phone to turn on the "video call" client or the "live streaming with goods" client, then the camera module starts to collect a video stream and obtains an image of the scene in front of the camera module. When the "video call" client or the "live streaming with goods" client is turned on, the tracking mode in the "video call" client or the "live streaming with goods" client has been turned on. The turning on of the tracking mode in the "video call" client can be achieved through a control in the "video call" client. The turning on of the tracking mode in the "live streaming with goods" client can be achieved through a control in the "live streaming with goods" client.

[0125] To turn on the tracking mode in the "video call" client, in a specific implementation, as Figure 5A shown, the user can click on the "video call" client on the home screen of the mobile phone. The client is not limited to the "video call" client, but can also be other clients including video call functions or other clients with similar video call functions. The present application does not make any restrictions on this. After the mobile phone detects that the user clicks on the "video call" client, it can display the user interface of the video call, as Figure 5B shown. In Figure 5B , the user interface of the video call may include: a video display area 51, a hang-up control 52, a camera switching control 53, a more options control 54, a status bar 55, a settings control 56, and a tracking mode switch control 57. It can be understood that Figure 5B is an example of the user interface of the video call. The user interface of the video call may also include a window shrinking control, etc. The present application does not limit the content and form of the user interface of the video call.

[0126] The video display area 51 is used to display the video stream captured by the camera module of the mobile phone of the video contact. The hang-up control 52 is used to interrupt the video call. The mobile phone can detect a touch operation acting on the hang-up control 52 (such as a click operation on the hang-up control 52), and interrupt the video call in response to the operation. The camera switching control 53 is used to switch the camera. The mobile phone can detect a touch operation acting on the camera switching control 53 (such as a click operation on the camera switching control 53), and in response to the operation, switch the camera module of the mobile phone from the front camera to the rear camera, or switch the camera module of the mobile phone from the rear camera to the front camera. The more options control 54 may include a window switching control, etc. The mobile phone can detect a touch operation acting on the more options control 54 (such as a click operation on the more options control 54), and display the window switching control. The window switching control is used to display the video stream captured by the camera module of the mobile phone and switch the video window. The mobile phone can detect a touch operation acting on the window switching control (such as a click operation on the window switching control), and in response to the operation, switch the content displayed by the window switching control and the video display area 51. The status bar 55 may include network, signal strength, battery status, and time, etc.

[0127] The setting control 56 is used to receive the setting instruction input by the user. As Figure 5B shown, the user can click the setting control 56. After the mobile phone detects the selected setting control 56 by the user, it displays the setting interface, as Figure 5C shown. The setting interface may include a default turn-on intelligent tracking control and a status bar. Figure 5C is an instance of the setting interface. The setting interface of the video call may also include more or fewer controls than Figure 5C shown. The present application does not limit the content and form of the user interface of the video call. The default turn-on intelligent tracking control includes a default turn-on intelligent tracking font and a corresponding selection control. The selection control can have two states: on and off. When the user operates on the default turn-on intelligent tracking font or the selection control, the state of the tracking mode can be switched. For example, when the tracking mode is in the off state, the user selects the default turn-on intelligent tracking font or the selection control (as Figure 5C shown), then the mobile phone switches to the on state of the tracking mode (as Figure 5D shown). In Figure 5C , the selection control is in the off state. In Figure 5D , the selection control is in the on state; when the tracking mode is in the on state, the user selects the default turn-on intelligent tracking font or the selection control, then the mobile phone switches to the off state of the tracking mode.

[0128] The tracking mode switch control 57 is used to receive the instruction for turning on or off the tracking mode input by the user. As Figure 6AAs shown, the tracking mode is in the off state, and the user can click on the tracking mode switch control 57. After the mobile phone detects that the user selects the tracking mode switch control 57, the tracking mode is turned on, as Figure 6B shown. In Figure 6A , the tracking mode switch control includes a flag 61, indicating that the tracking mode is in the off state; in Figure 6B , the tracking mode switch control does not include a flag, indicating that the tracking mode is in the on state. If the tracking mode is in the on state, the user can also click on the tracking mode switch control. After the mobile phone detects that the user selects the tracking mode switch control, the tracking mode is turned off.

[0129] To turn on the tracking mode in the "Live Commerce" client, in a specific implementation, as Figure 7A shown, the user can click on the "Live Commerce" client on the home screen of the mobile phone. The client is not limited to the "Live Commerce" client, and can also be clients such as Kuaishou, Taobao, etc. that include live commerce functions or other similar live commerce function clients. This application does not make any restrictions on this. After the mobile phone detects that the user clicks on the "Live Commerce" client, it can display the user interface of live commerce, as Figure 7B shown. In Figure 7B , the user interface of live commerce can include: a video display area 71, a play / pause control 72, a status bar 73, a settings control 74, and a tracking mode switch control 75. It can be understood that Figure 7B is an example of the user interface of live commerce. The user interface of live commerce can include more or fewer controls. This application does not limit the content and form of the user interface of live commerce.

[0130] The video display area 71 is used to display the video stream collected by the camera module of the mobile phone. The play / pause control 72 is used to pause / play the live broadcast. The mobile phone can detect a touch operation on the play / pause control 72 (such as a click operation on the play / pause control 72) and respond to the operation to pause / play the live broadcast. For example, if the play / pause control 72 is in the play state, the user can select the play / pause control 72, and the mobile phone responds to the selection operation to pause the live broadcast; if the play / pause control 72 is in the pause state, the user can select the play / pause control 72, and the mobile phone responds to the selection operation to play the live broadcast. The status bar 73 can include network, signal strength, battery status, and time, etc.

[0131] The settings control 74 is used to receive the setting instructions input by the user. As Figure 7B shown, the user can click on the settings control 74. After the mobile phone detects the settings control 74 selected by the user, it displays the settings interface, as Figure 7C shown. The settings interface can include a default on intelligent tracking control and a status bar.Figure 7C is an example of a setting interface, and the setting interface of the video call may also include Figure 7C More or fewer controls as shown. This application does not limit the content and form of the user interface for video calls. The default smart tracking control includes a default smart tracking font and a corresponding selection control. The selection control can have two states, on and off. When the user operates the default smart tracking font or the selection control, the state of the tracking mode can be switched. For example, when the tracking mode is off, the user selects the default smart tracking font or the selection control (such as Figure 7C ), the phone switches to tracking mode on (as shown Figure 7D shown), in Figure 7C , the selection controls are closed. Figure 7D In the tracking mode, the selection control is in the on state; when the tracking mode is in the on state, the user chooses to turn on the smart tracking font by default or selects the control, the phone switches to the tracking mode off state.

[0132] The tracking mode switch control 75 is used to receive a tracking mode on or off instruction input by a user. Figure 8A As shown, the tracking mode is in the off state, and the user can click the tracking mode switch control 75. After the mobile phone detects that the user selects the tracking mode switch control 75, the tracking mode is turned on. Figure 8B As shown. Figure 8A In the example, the tracking mode switch control includes a mark 81, indicating that the tracking mode is in the off state; in the example Figure 8B In the example, the tracking mode switch control does not include a sign, indicating that the tracking mode is in the on state. If the tracking mode is in the on state, the user can also click the tracking mode switch control, and the mobile phone turns off the tracking mode after detecting that the user selects the tracking mode switch control.

[0133] To facilitate subsequent description, the application is explained below using the client as a "video call" client as an example.

[0134] S402, the electronic device obtains a video stream captured by a camera module, where the video stream includes image frames.

[0135] After the client is turned on, in the tracking mode, the electronic device first selects the tracking target, and then tracks the tracking target. Before selecting the tracking target, the camera module is fixed to collect image frames of the same scene and the same perspective. The camera collects video streams at a fixed frequency, such as 30 frames per second. The video stream includes multiple frames of image frames sorted in chronological order. The electronic device obtains the video stream collected by the camera module in real time. For example, at time t1, the electronic device obtains image frame 1 (such as Figure 9Aas shown); at time t2, the electronic device obtains the image frame 2 in the video stream collected by the camera module (such as Figure 9B as shown). It can be understood that although the electronic device displays the image frame 1 and the image frame 2 in Figure 9A and Figure 9B , this does not prevent the image frames collected by the camera module from being considered as the image frames in Figure 9A and Figure 9B .

[0136] S403, obtain the depth information of the image frames in the video stream.

[0137] If the camera module is a binocular camera, a structured light camera, or a TOF (Time of flight) camera, the depth information of the photographed object is included in the image frame, and the depth information of the photographed object in the image frame can be directly obtained. If the camera module is an ordinary monocular camera and the depth information of the photographed object is not included in the image frame, a monocular depth estimation algorithm can be used to obtain the depth information of the photographed object in the image frame. The electronic device obtains the depth information of the image frames in the video stream in real time. Continuing with the above Figure 9A and Figure 9B as an example to illustrate the present application, Figure 9A includes objects: people, tables, cups, razors, etc. When the electronic device obtains the Figure 9A shown image frame 1, it obtains the depth information of people, tables, cups, razors, etc. in the image frame 1; Figure 9B also includes objects: people, tables, cups, razors, etc. When the electronic device obtains the Figure 9B shown image frame 2, it obtains the depth information of people, tables, cups, razors, etc. in the image frame 2. Specifically, the electronic device can obtain the depth information of all pixel points of people, the depth information of all pixel points of the table, the depth information of all pixel points of the cup, and the depth information of all pixel points of the razor, etc. in the Figure 9A shown image frame 1; the electronic device can obtain the depth information of all pixel points of people, the depth information of all pixel points of the table, the depth information of all pixel points of the cup, and the depth information of all pixel points of the razor, etc. in the Figure 9B shown image frame 2.

[0138] S404, determine that the change region between the first adjacent image frames in the video stream is the object to be detected. The first adjacent image frames include the first image frame and the second image frame, the first image frame is the image frame in front of the second image frame, and the change region is the difference region of the depth information.

[0139] The video stream captured by the camera module is continuous in time, and the position of the object being photographed does not change suddenly. If there are no moving objects in the scene, the change between the first adjacent image frames is very small. If there are moving objects in the scene, the change between adjacent image frames will exceed the threshold. In this embodiment, determining the change region between the first adjacent image frames in the video stream as the object to be detected includes:

[0140] Determine the similarity between the first adjacent image frames in the video stream, where the similarity is the similarity of depth information; if the similarity between the first adjacent image frames in the video stream is less than the threshold, determine the change region between the first adjacent image frames in the video stream, and the change region is the difference region of depth information; determine the change region as the object to be detected. Among them, when determining the change region, a preset rule can be set. For example, exclude people from the change region, or determine the object as the change region only when the overall similarity of the objects in the first adjacent image frames in the video stream is less than the threshold. For example, determine a person as the change region only when the similarity of the whole person in the first adjacent image frames in the video stream is less than the threshold.

[0141] The first adjacent image frames can be one first adjacent image frame or multiple first adjacent image frames. Continuing with the above Figure 9A and Figure 9B as an example to illustrate the present application, one first adjacent image frame can include Figure 9A the image frame 1 shown in Figure 9B and Figure 9A the image frame 2 shown in Figure 9B Among them, Figure 9A the person, table, cup, and razor in the image frame 1 shown in Figure 9B and Figure 9A the cup, human arm, and eyes in the image frame 1 shown in Figure 9B and

[0142] the similarity of the depth information between the person, table, cup, and razor in the image frame 2 shown in Figure 9A the image frame 1 shown in Figure 9B and Figure 9A the similarity of the depth information between the cup, human arm, and eyes in the image frame 1 shown in Figure 9B andFigure 9B The image frame 2 shown, the image frame 3, and the image frame 4 are three adjacent image frames. The process of determining the object to be detected from multiple first adjacent image frames is similar to the process of determining the object to be detected from one first adjacent image frame, which will not be elaborated here. Among them, in the process of determining the object to be detected from multiple first adjacent image frames, if the similarity between the objects (such as cups) in one of the first adjacent image frames is greater than the threshold, this first adjacent image frame can be ignored.

[0143] The similarity between the first adjacent image frames is the similarity between the pixel points of the first adjacent image frames. Specifically, the device can compare the similarity between the pixel points of the first adjacent image frames; determine the pixel points whose change value of the depth information between the adjacent image frames exceeds the threshold as the first pixel points; determine the depth change values of the first pixel points in the adjacent image frames; determine the depth change image according to the depth change values of the first pixel points; calculate the minimum depth change value between each pixel point in the depth change image and its spatially adjacent pixel points to form a distance difference image; perform threshold binary processing on the distance difference image to obtain a binary image; in the binary image, perform connected component labeling to determine the connected components; determine the object to be detected according to the connected components. The connected component is the above-mentioned change region.

[0144] Continuing with the above Figure 9A and Figure 9B as an example to illustrate how to determine the object to be detected according to the pixel points in the case of one first adjacent image frame, Figure 9A When comparing the image frame 1 shown with Figure 9B the image frame 2 shown, the change values of the depth information of the pixel points a, b, c, d, f, g, h between the image frame 1 and the image frame 2, which are 0.21 meters, 0.25 meters, 0.29 meters, 0.3 meters, 0.21 meters, 0.25 meters, 0.27 meters, exceed the threshold of 0.2 meters. Then, the first pixel points are determined as the pixel points a, b, c, d, f, g, h, and the depth change values of the first pixel points a, b, c, d, f, g, h in the first adjacent image frame are determined to be 0.21 meters, 0.25 meters, 0.29 meters, 0.3 meters, 0.21 meters, 0.25 meters, 0.27 meters respectively.

[0145] The connected component refers to a region composed of pixels with the same pixel value and adjacent positions in the image frame. There is a certain similarity between each pixel point in the connected component and its adjacent pixel points in space. Then, the depth difference between each pixel point in the connected component and the adjacent pixel points will not change suddenly, that is, the absolute value of the depth difference between each pixel point in the connected component and the adjacent pixel points is less than a certain depth difference. The image data of the depth change image is as Figure 10AAs shown. The depth change image is a single-channel image. The value represented by each pixel point in the depth change image is the depth change value of the pixel point. For example, in Figure 10A , the resolution of the depth change image is 100x100. The pixel points with a value of 0 in the depth change image indicate that the change value of the depth information of the pixel points in the first adjacent image frame is less than the threshold, and the pixel points with a non-zero value indicate that the change value of the depth information of the pixel points in the first adjacent image frame exceeds the threshold. The pixel point and its spatially adjacent pixel points can be as Figure 10B or as Figure 10C shown. Figure 10B And Figure 10C take the pixel point f in the above-mentioned first pixel point as an example for illustration. In Figure 10B , the pixel point f has 4 pixel points adjacent to it spatially, namely pixel point g, pixel point h, pixel point i, and pixel point j. The pixel point g, the pixel point h, the pixel point i, and the pixel point j are respectively located directly above, directly below, directly to the left, and directly to the right of the pixel point f. In Figure 10C , the pixel point f has 8 pixel points adjacent to it spatially, namely pixel point g, pixel point h, pixel point i, pixel point j, pixel point k, pixel point l, pixel point m, and pixel point n. The pixel point g, the pixel point h, the pixel point i, the pixel point j, the pixel point k, the pixel point l, the pixel point m, and the pixel point n are respectively located directly above, directly below, directly to the left, directly to the right, upper left corner, upper right corner, lower left corner, and lower right corner of the pixel point f.

[0146] For the convenience of description, continue to take the pixel point f in the above-mentioned first pixel point and the case where the pixel point f has 4 pixel points adjacent to it spatially as an example to illustrate how to calculate the minimum depth change value between each pixel point in the depth change image and its spatially adjacent pixel points. As Figure 10B shown, the depth change value of the pixel point f is 0.21 meters, the depth change value of the pixel point g is 0.25 meters, the depth change value of the pixel point h is 0.27 meters, the depth change value of the pixel point i is 0.2 meters, the depth change value of the pixel point j is 0.3 meters. The depth change values between the pixel point f and its spatially adjacent pixel points g, h, i, j are: 0.04 meters, 0.06 meters, 0.01 meters, 0.09 meters. Then the minimum depth change value between the pixel point f and its spatially adjacent pixel points is 0.01 meters. The image data of the distance difference image is as Figure 10D shown. The distance difference image is a single-channel image. The value represented by each pixel point in the distance difference image is the minimum depth change value between the pixel point and its spatially adjacent pixel points.

[0147] The threshold binary processing of the distance difference image can be as follows: if the value represented by a pixel point is less than a preset value (such as 0.03, etc.), the value represented by the pixel point after threshold binary processing is 1; if the value represented by a pixel point is greater than the preset value (such as 0.03, etc.), the value represented by the pixel point after threshold binary processing is 0. Continuing to illustrate the present application with the pixel point f in the first pixel point above as an example, the value 0.01 represented by the pixel point f is less than the preset value 0.03, so the value represented by the pixel point f after threshold binary processing is 1, that is, the pixel value of the pixel point f in the binary image is 1. It can be understood that according to the above process of determining the pixel value of the pixel point f in the binary image, the pixel values of other first pixel points a, b, c, d, g, h in the binary image can also be determined. Then, according to the above Figure 9A and Figure 9B The image data of the obtained binary image is as Figure 10E shown, and the binary image is a single-channel image. The pixel value represented by each pixel point in the binary image is the value after threshold binary processing of the minimum depth change value.

[0148] In the binary image, connected component labeling is performed to determine the connected components. The number of connected components can be one or more. After performing connected component labeling on the binary image shown above Figure 10E connected components H, I, and J are determined. Connected component H is a cup, connected component I is a person's arm, and connected component J is a person's eye. Connected component H includes pixel points a, b, c, d, connected component I includes pixel points f, g, and connected component J includes h. After determining the connected components, according to the above preset rules: excluding people from the changing area, or determining an object as a changing area only when the similarity between the overall objects in the first adjacent image frames of the video stream is less than a threshold, connected components I and J can be excluded, thereby completing the determination of the connected components. Then, the connected component with the largest area and / or the largest depth change value in the first adjacent image frames can be determined as the object to be detected. For example, taking another example to illustrate, the number of connected components after exclusion is two, namely connected component K and connected component L. If the area of connected component K is larger than the area of connected component L, the object to be detected is determined as connected component K, or if the depth change value of connected component K in the first adjacent image frames is greater than the depth change value of connected component L, the object to be detected is determined as connected component K. Thus, when there are multiple connected component changes, some connected components can be excluded, that is, some noises in the image frames can be excluded. Among them, when according to Figure 9A and Figure 9BAfter obtaining the connected component H, it can be determined that the connected component H is the object to be detected, that is, the cup is the object to be detected. It can be understood that in this application, when determining the first pixel point, or determining the depth change value of the first pixel point, or determining the depth change image, or forming the distance difference image, or obtaining the binary image, the human arm and human eyes can be excluded from the change.

[0149] In the case of multiple first adjacent image frames, the process of determining the object to be detected based on pixel points is similar to the process of determining the object to be detected based on pixel points in the case of one first adjacent image frame, and will not be elaborated here. Among them, in the process of determining the object to be detected based on pixel points in the case of multiple first adjacent image frames, the average value of the change values of the depth information of the pixel points between all first adjacent image frames is used to determine the first pixel point and the depth change value of the first pixel point. For example, Figure 9A the shown image frame 1 is compared with Figure 9B the shown image frame 2. The change values of the depth information of the pixel points a, b, c, d, f, g, h in the image frame, namely 0.21 m, 0.25 m, 0.29 m, 0.3 m, 0.21 m, 0.25 m, 0.27 m, exceed the threshold of 0.2 m. For the convenience of description, hereinafter, only the pixel point a will be used as an example to illustrate whether the pixel point a is the first pixel point, and if the pixel point a is the first pixel point, to determine the depth change value of the first pixel point a. The image frame 2 is compared with the image frame 3, and the change value of the depth information of the pixel point a, 0.27 m, exceeds the threshold of 0.2 m. The image frame 3 is compared with the image frame 4, and the change value of the depth information of the pixel point a, 0.25 m, exceeds the threshold of 0.2 m. Then it is determined that the pixel point a is the first pixel point, and the depth change values of the first pixel point a are determined to be the average value of 0.21 m, 0.27 m and 0.25 m respectively, that is, 0.73 m / 3. According to the process of determining whether the pixel point a is the first pixel point, and if the pixel point a is the first pixel point, to determine the depth change value of the first pixel point a, it is determined whether the pixel points b, c, d, f, g, h are the first pixel points and the depth change values of each first pixel point among the pixel points b, c, d, f, g, h.

[0150] S405, determine the displacement value and displacement direction between the position of the object to be detected in the second image frame and the position of the object to be detected in the first image frame. The position is the position of the depth information, and the displacement value and displacement direction are the displacement value and displacement direction of the depth information.

[0151] The object to be detected includes multiple interconnected pixel points in the image frame. The displacement value of the depth information is the absolute value of the average depth change value of the pixel points. The displacement direction includes the direction towards the focal plane and the direction away from the focal plane. If the average depth change value of the pixel points is greater than zero, the displacement direction is away from the focal plane; if the average depth change value of the pixel points is less than zero, the displacement direction is towards the focal plane. Continuing with the above Figure 9A and Figure 9B as an example to illustrate the present application, the object to be detected, the cup, includes pixel points a, b, c, and d; the pixel points a, b, c, and d are in Figure 9B the image frame 2 shown in Figure 9A and the change values of the depth information between the image frame 1 shown in

[0152] S406, select a tracking target, where the tracking target is an object to be detected whose displacement value is greater than a first preset value and the displacement direction is towards the focal plane.

[0153] Continuing with the above Figure 9A and Figure 9B as an example to illustrate the present application, Figure 9B the position of the object to be detected, the cup, in the image frame 2 shown in Figure 9A compared to the position of the object to be detected, the cup, in the image frame 1 shown in

[0154] S407, track the selected tracking target.

[0155] When tracking the tracking target, the camera module can be driven to collect image frames of different scenes.

[0156] S408, during tracking, determine the position of the tracking target in the second adjacent image frames in the video stream. The second adjacent image frames include the third image frame and the fourth image frame, and the third image frame is the image frame in front of the fourth image frame. The position is the position of the depth information.

[0157] During tracking, the electronic device obtains the video stream collected by the camera module in real time, and also obtains the depth information of the image frames in the video stream in real time. The second adjacent image frame can be one second adjacent image frame or multiple second adjacent image frames. Hereinafter, the present application will be described by taking one second adjacent image frame as an example. For example, at time t3, the electronic device obtains image frame 7 in the video stream collected by the camera module (as shown in Figure 11A ); at time t4, the electronic device obtains image frame 8 in the video stream collected by the camera module (as shown in Figure 11B ). Image frame 7 is the third image frame, image frame 8 is the fourth image frame, and image frame 7 is the image frame in front of image frame 8. The present application determines the positions of the tracking target in image frame 7 and image frame 8 in the depth information. It can be understood that although image frame 7 and image frame 8 are shown in Figure 11A and Figure 11B by the electronic device, this does not prevent the image frames collected by the camera module from being considered as the image frames in Figure 11A and Figure 11B .

[0158] S409. Determine the displacement value and displacement direction between the position of the tracking target in the fourth image frame and the position of the tracking target in the third image frame, where the displacement value and the displacement direction are the displacement value and displacement direction of the depth information.

[0159] The tracking target includes a plurality of interconnected pixel points in the image frame. The displacement value of the depth information is the absolute value of the average depth change value of the pixel points. The displacement direction includes the direction close to the focal plane and the direction away from the focal plane. If the average depth change value of the pixel points is greater than zero, the displacement direction is away from the focal plane; if the average depth change value of the pixel points is less than zero, the displacement direction is close to the focal plane. Continuing with the above Figure 11A and Figure 11B as an example to illustrate the present application, the tracking target cup includes pixel points a, b, c, d; the change values of the depth information of pixel points a, b, c, d between image frame 8 and image frame 7 are 0.33 meters, 0.28 meters, 0.36 meters, and 0.27 meters respectively. Then, it is determined that the average depth change value of the pixel points between the position of the tracking target cup in image frame 8 and the position of the tracking target cup in image frame 7 is 0.31 meters, and it is determined that the displacement value between the position of the tracking target cup in image frame 8 and the position of the tracking target in image frame 7 is |0.31| meters, that is, 0.31 meters. The average depth change value of 0.31 meters of the pixel points between the position of the tracking target cup in image frame 8 and the position of the tracking target cup in image frame 7 is greater than zero, so it is determined that the displacement direction is the direction away from the focal plane.

[0160] S410. If the displacement value is less than the second preset value and the displacement direction is the direction away from the focal plane or the displacement direction is the direction close to the focal plane, focus and display the tracking target.

[0161] The focused display may include tracking the target while following the focus, framing and displaying the tracking target, or cropping and centering the tracking target. Framing and displaying the tracking target may be marking the tracking target in the image frame by a square, a circle, or the contour shape of the object, etc. Cropping and centering the displayed tracking target may be cropping other parts of the image except the tracking target and magnifying and centering the remaining part for display, as Figure 11A shown. In Figure 11A the cup is cropped and centered for display.

[0162] S411, if the displacement value is greater than the second preset value and the displacement direction is away from the focal plane, exit the tracking.

[0163] In this embodiment, when exiting the tracking, the tracking target is also reselected. Continuing with the above Figure 11A and Figure 11B as an example to illustrate the present application, Figure 11B the position of the tracking target cup in the image frame 8 shown is displaced by 0.31 meters compared to the position of the tracking target in the image frame 7 shown Figure 11A and the displacement value is greater than the second preset value, and the displacement direction is away from the focal plane, then the tracking of the cup is exited.

[0164] Figure 4 The target tracking method shown can be used not only in the scenario of selecting and tracking the target according to the change of depth information, but also in the scenario of selecting and tracking the target according to the change of the key points of the human body and depth information in the image frame. In the scenario of selecting and tracking the target according to the change of the key points of the human body and depth information in the image frame, the difference from the above Figure 4 scenario of selecting and tracking the target according to depth information is as follows:

[0165] After obtaining the depth information of the image frames in the video stream, the human key points in the image frames in the video stream are also detected, and the change region connected to the first parameter of the human key points between the first adjacent image frames in the video stream is determined as the object to be detected. When selecting a tracking target, the tracking target is an object to be detected with a displacement value greater than a first preset value, a displacement direction towards the focal plane, and a depth information difference greater than a third preset value between the second parameter of the human key points in the image frame. When focusing on and displaying the tracking target, if the first condition and the second condition are satisfied, the tracking target is focused on and displayed. The first condition includes that the tracking target is connected to the first parameter of the human key points in the image frame, and the second condition includes that the displacement value is less than a second preset value and the displacement direction is away from the focal plane or the displacement direction is towards the focal plane; when exiting the tracking, if the displacement value is greater than the second preset value and the displacement direction is away from the focal plane, or the tracking target is not connected to the first parameter of the human key points, the tracking is exited. When detecting the human key points in the image frames in the video stream, the positions of the human key points in the image frames are detected according to the depth information of the image frames in the video stream. The human key points are as Figure 12 shown. The human key points include a first parameter, a third parameter, a fourth parameter, a fifth parameter, and a sixth parameter. The first parameter, the third parameter, the fourth parameter, the fifth parameter, and the sixth parameter are the left and right wrists, the left and right shoulders, the neck, the head, and the left and right hips respectively. In Figure 12 although only the human key points including the first parameter, the third parameter, the fourth parameter, the fifth parameter, and the sixth parameter are shown, it is obvious that the human key points may also include parts such as a seventh parameter, an eighth parameter, and a ninth parameter. The seventh parameter, the eighth parameter, and the ninth parameter are the left and right elbows, the left and right knees, and the left and right ankles respectively.

[0166] When determining the object to be detected, as in Figure 4 the process of determining the connected region in step S404, the electronic device first determines the connected region, and determines the connected region connected to the first parameter of the human key points as the object to be detected according to the positions of the human key points in the image frame and the connected region. For example, the electronic device determines the connected region razor according to Figure 13A and Figure 13B The connected region razor includes pixel points o, p, q, and r. Figure 13A The image frame 9 shown in Figure 13B is the first image frame, and Figure 13A and Figure 13B The image frame 10 shown in is the second image frame. The image frame 9 is the image frame in front of the image frame 10. It can be understood that although the electronic device displays the image frame 9 and the image frame 10 in Figure 13A and Figure 13B this does not prevent the image frames collected by the camera module from being considered as the image frames in Figure 13BThe position of the human key points in the shown image frame 10 and the connected region are determined, and the connected region adjacent to the first parameter of the human key points, the razor, is the object to be detected.

[0167] According to the position of the human key points in the image frame and the connected region, the connected region connected to the first parameter of the human key points is determined as the object to be detected. Specifically:

[0168] Determine the position of the center of each connected region in the image frame; according to the position of the first parameter of the human key points in the image frame and the position of the center of each connected region in the image frame, determine the Euclidean distance between the first parameter of the human key points in the image frame and the center of the connected region, and determine the connected region with the Euclidean distance less than the preset value as the object to be detected. Among them, the Euclidean distance less than the preset value means that the connected region is connected to the first parameter of the human key points.

[0169] In the image frame coordinate system, the position of the human key points in the image frame includes the coordinates of the human key points in the image frame coordinate system, and the position of the center of each connected region in the image frame includes the coordinates of the center of the connected region in the image frame coordinate system. Determining the Euclidean distance between the first parameter of the human key points in the image frame and the center of the connected region according to the position of the human key points in the image frame and the position of the center of each connected region in the image frame includes: through the formula Determine the Euclidean distance between the first parameter of the human key points in the image frame and the center of the connected region according to the coordinates of the human key points in the image frame and the position of the center of each connected region in the image frame. Among them, p ji is the Euclidean distance between the j-th wrist of the first parameter of the human key points in the image frame and the center of the i-th connected region, x 1i is the abscissa of the center of the i-th connected region in the image frame, x j is the abscissa of the j-th wrist of the first parameter of the human key points in the image frame, y 1i is the ordinate of the center of the i-th connected region in the image frame, y j is the ordinate of the j-th wrist of the first parameter of the human key points in the image frame. Among them, i = 1, 2,..., n, j = 1, 2. Continuing with the above Figure 13B Taking the shown image frame 10 as an example to illustrate the present application, the electronic device determines that the position of the center of the connected region razor in the image frame 10 is (x 11 , y 11) The positions of the human key points in the image frame 10 include the position of the left wrist in the first parameter of the human key points being (x1, y1), and the position of the right wrist in the first parameter of the human key points being (x2, y2). The Euclidean distance between the left wrist in the first parameter of the human key points in the image frame 10 and the center of the connected component razor is greater than the preset value, and the Euclidean distance between the right wrist in the first parameter of the human key points in the image frame 10 and the center of the connected component is less than the preset value, then it is determined that the connected component razor is the object to be detected.

[0170] In this embodiment, the electronic device determines that the connected component with the Euclidean distance less than the preset value, the largest area, and / or the largest depth change value in the first adjacent image frame is the object to be detected. When selecting the tracking target, such as Figure 4 in the process of determining in step S406 that the displacement value is greater than the first preset value and the displacement direction is the direction close to the focal plane, the electronic device first determines the object to be detected in which the displacement value is greater than the first preset value and the displacement direction is the direction close to the focal plane among the objects to be detected, and then determines the object to be detected in which the depth information difference between the object to be detected and the second parameter of the human key points in the image frame is greater than the third preset value as the tracking target. To determine the object to be detected in which the depth information difference between the object to be detected and the second parameter of the human key points in the image frame is greater than the third preset value, specifically:

[0171] Determine the depth information of the second parameter of the human key points according to the depth information of the image frame and the positions of the human key points in the image frame, and determine the object to be detected in which the depth information difference between the depth information of the object to be detected and the depth information of the second parameter of the human key points in the image frame is greater than the third preset value.

[0172] The second parameter of the human key points is the human body trunk. In this embodiment, determine the depth information of the third parameter, the fourth parameter, the fifth parameter, and the sixth parameter of the human key points in the image frame according to the depth information of the image frame and the positions of the third parameter, the fourth parameter, the fifth parameter, and the sixth parameter of the human key points in the image frame, and determine the depth information of the second parameter of the human key points according to the depth information of the third parameter, the fourth parameter, the fifth parameter, and the sixth parameter of the human key points in the image frame. Continuing with the above Figure 13BTaking the image frame 10 shown as an example to illustrate the present application. In the image frame 10, the depth information of the third parameter of the human body key points for the left and right shoulders is 1.5 meters and 1.54 meters respectively, the depth information of the fourth parameter for the neck is 1.52 meters, the depth information of the fifth parameter for the head is 1.52 meters, and the depth information of the sixth parameter for the left and right hips is 1.51 meters and 1.53 meters respectively. Then, the depth information of the second parameter of the human body key points, the torso, in the image frame 10 is (1.5 meters + 1.54 meters + 1.52 meters + 1.52 meters + 1.51 meters + 1.53 meters) / 6, that is, 1.52 meters. The object to be detected in the image frame includes a plurality of interconnected pixel points. The depth information of the object to be detected is the average value of the depth information of the pixel points of the object to be detected, and the depth information difference between the object to be detected and the second parameter of the human body key points in the image frame is the average depth information difference between the pixel points of the object to be detected and the second parameter of the human body key points in the image frame.

[0173] Continuing with the above Figure 13A and Figure 13B as an example to illustrate the present application, Figure 13B the position of the object to be detected, the razor, in the image frame 10 shown Figure 13A is compared with the position of the object to be detected, the razor, in the image frame 9 shown. The displacement value is greater than the first preset value, and the displacement direction is towards the direction of the focal plane. Moreover, for the pixel points o, p, q, r included in the object to be detected, the razor, the average relative depth difference between the depth information in the image frame 10 and the second parameter of the human body key points, the torso, is greater than the third preset value. Then, the object to be detected, the razor, is selected as the tracking target.

[0174] When focusing on and displaying the tracking target, such as Figure 4 in the process from step S409 to step S410 in which it is determined that the displacement value is less than the second preset value and the displacement direction is away from the focal plane or the displacement direction is towards the focal plane, the electronic device determines that the displacement value is less than the second preset value and the displacement direction is away from the focal plane or the displacement direction is towards the focal plane, and then determines that the tracking target is connected to the first parameter of the human body key points in the image frame. Determining that the tracking target is connected to the first parameter of the human body key points in the image frame is similar to determining that the object to be detected is connected to the first parameter of the human body key points, and will not be elaborated here.

[0175] When exiting the tracking, such as Figure 4 in step S411 in which it is determined that the displacement value is greater than the second preset value and the displacement direction is away from the focal plane, the electronic device determines that the displacement value is greater than the second preset value and the displacement direction is away from the focal plane, or the electronic device determines that the tracking target is not connected to the first parameter of the human body key points, and exits the tracking. When the electronic device determines that the tracking target is not connected to the first parameter of the human body key points, specifically:

[0176] Determine the position of the center of the tracking target in the image frame; determine the Euclidean distance between the first parameter of the human key point in the image frame and the center of the tracking target according to the position of the human key point in the image frame and the position of the center of the tracking target in the image frame, and determine whether the Euclidean distance is less than a preset value. If the Euclidean distance is greater than the preset value, the electronic device determines that the tracking target is not connected to the first parameter of the human key point.

[0177] Determining the position of the center of the tracking target in the image frame is similar to determining the position of the center of each connected component in the image frame, which will not be elaborated here. Determining the Euclidean distance between the first parameter of the human key point in the image frame and the center of the tracking target according to the position of the human key point in the image frame and the position of the center of the tracking target in the image frame is similar to determining the Euclidean distance between the first parameter of the human key point in the image frame and the center of the connected component according to the position of the human key point in the image frame and the position of the center of each connected component in the image frame, which will not be elaborated here.

[0178] Reference Figure 14 , is a schematic hardware structure diagram of the server according to an embodiment of the present application. The server 14 includes a memory 143, a processor 144, and a communication interface 145. Those skilled in the art can understand that Figure 14 the structure shown in

[0179] does not limit the server 14. The server 14 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements.

[0180] The processor 144 may be a Central Processing Unit (CPU), a graphics processing unit (GPU), an image signal processor (ISP), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 144 may be a microprocessor or any conventional processor, etc. The processor 144 is the control center of the server 14, connecting various parts of the entire server 14 through various interfaces and lines.

[0181] The communication interface 145 may include standard wired interfaces, wireless interfaces, etc. The communication interface 145 is used for the server 14 to communicate with the electronic device.

[0182] Figure 4 The target tracking method shown can be used not only on the electronic device but also on the system composed of the electronic device and the server. The difference between applying the target tracking method to the system composed of the electronic device and the server and applying it to the electronic device is as follows:

[0183] After the electronic device obtains the video stream collected by the camera module in step S402, the electronic device also transmits the collected video stream to the server through the client. The server executes steps S403 to S406, and after selecting the tracking target, transmits the selected tracking target and the drive signal to the electronic device to control the movement of the camera module of the electronic device and track the selected tracking target. The server also executes steps S408 to S410, and when the displacement value is greater than the second preset value and the displacement direction is away from the focal plane, transmits an exit tracking signal to the electronic device to control the electronic device to exit the tracking.

[0184] Obviously, the above-mentioned target tracking method can be used not only in the scenario of selecting and tracking a tracking target according to the change of depth information, but also in the scenario of selecting and tracking a tracking target according to the change of the key points of the human body and depth information in the image frame. The system composed of the electronic device and the server can also be applied to the scenario of selecting and tracking a tracking target according to the change of the key points of the human body and depth information in the image frame. The difference is that in the scenario of selecting and tracking a tracking target according to the change of the key points of the human body and depth information in the image frame, compared with the scenario of selecting and tracking a tracking target according to depth information, the difference is mainly executed by the server. Only when the displacement value is greater than the second preset value and the displacement direction is away from the focal plane, or when the first parameter of the tracking target is not connected to the key points of the human body, a tracking exit signal is transmitted to the electronic device to control the electronic device to exit tracking.

[0185] Please refer to Figure 15 , Figure 15 FIG. is a schematic structural diagram of a target tracking device provided by an embodiment of the present invention. The target tracking device 15 may include an acquisition unit 151 and a determination unit 152.

[0186] The acquisition unit 151 is configured to acquire the depth information of the image frame in the video stream.

[0187] The determination unit 152 is configured to determine that the change region between the first adjacent image frames in the video stream is an object to be detected. The first adjacent image frames include a first image frame and a second image frame, and the first image frame is the image frame in front of the second image frame. The change region is the difference region of the depth information.

[0188] The determination unit 152 is further configured to determine the displacement value and displacement direction between the position of the object to be detected in the second image frame and the position of the object to be detected in the first image frame. The position is the position of the depth information, and the displacement value and displacement direction are the displacement value and displacement direction of the depth information.

[0189] The determination unit 152 is further configured to select a tracking target, where the tracking target is an object to be detected with a displacement value greater than a first preset value and a displacement direction close to the focal plane.

[0190] Optionally, when the determination unit 152 is used for tracking, it determines the position of the tracking target in the second adjacent image frames. The second adjacent image frames include a third image frame and a fourth image frame, and the third image frame is the image frame in front of the fourth image frame. The position is the position of the depth information. The determination unit 152 is further configured to determine the displacement value and displacement direction between the position of the tracking target in the fourth image frame and the position of the tracking target in the third image frame. The displacement value and the displacement direction are the displacement value and displacement direction of the depth information. The determination unit 152 is further configured to exit tracking if the displacement value is greater than a second preset value and the displacement direction is away from the focal plane.

[0191] Optionally, the determination unit 152 is configured to detect human key points in an image frame in a video stream. The determination unit 152 is further configured to determine a change region associated with a first parameter of the human key points between the first adjacent image frames in the video stream as an object to be detected. The determination unit 152 is further configured to select a tracking target, where the tracking target is an object to be detected with a displacement value greater than a first preset value, a displacement direction towards the focal plane, and a depth information difference greater than a third preset value between the second parameter of the human key points in the image frame.

[0192] Optionally, when the determination unit 152 is further configured to perform tracking, it determines the position of the tracking target in the second adjacent image frames, where the second adjacent image frames include a third image frame and a fourth image frame, the third image frame is the image frame in front of the fourth image frame, and the position is the position of the depth information. The determination unit 152 is further configured to determine the displacement value and the displacement direction of the position of the tracking target in the fourth image frame compared with the position of the tracking target in the third image frame, where the displacement value and the displacement direction are the displacement value and the displacement direction of the depth information. The determination unit 152 is further configured to exit the tracking if the displacement value is greater than a second preset value and the displacement direction is away from the focal plane, or the tracking target is not connected to the first parameter of the human key points.

[0193] Optionally, the displacement value of the depth information is the absolute value of the average depth change value of the pixel points.

[0194] The target tracking device described in the embodiments of the present application can be used to implement the operations performed by the electronic device or the server in the above target tracking method.

[0195] In addition to the above methods and devices, the embodiments of the present application further provide a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when they run on a processor, they implement the target tracking method.

[0196] A computer program product includes computer execution instructions, and the computer execution instructions are stored in a computer-readable storage medium; at least one processor of the device can read the computer execution instructions from the computer-readable storage medium, and the at least one processor executes the computer execution instructions to enable the device to implement the target tracking method.

[0197] Before tracking, the present application can determine a significantly forward-moving object as the selected tracking target, and can determine the position and size of the selected tracking target through a simple interaction method, without manually drawing a bounding box and can detect too small objects; during tracking, it can determine a significantly backward-moving object as the object to exit the tracking, and can make the object exit the tracking through a simple interaction method, so as to exit the tracking.

[0198] Before tracking, the present application can determine an object picked up by hand and significantly moved forward as a selected tracking target, and can determine the position and size of the selected tracking target through a simple interaction method, without manually drawing a bounding box and can detect too small objects; during tracking, it can determine an object put down by hand or significantly moved backward as an object to exit tracking, and can make the object exit tracking through a simple interaction method, so as to exit tracking.

[0199] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0200] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0201] In addition, each functional unit in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0202] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A target tracking method, characterized in that, The method includes: Obtaining the depth information of the image frames in the video stream; Determining that the change region between the first adjacent image frames in the video stream is the object to be detected. The first adjacent image frames include a first image frame and a second image frame, and the first image frame is the image frame in front of the second image frame. The change region is the difference region of the depth information; Determining the displacement value and displacement direction between the position of the object to be detected in the second image frame and the position of the object to be detected in the first image frame. The position is the position of the depth information, and the displacement value and displacement direction are the displacement value and displacement direction of the depth information; Selecting a tracking target, where the tracking target is the object to be detected with a displacement value greater than a first preset value and a displacement direction towards the focal plane; Detecting the human key points in the image frames of the video stream; Wherein, the object to be detected is the change region connected to the first part of the human key points between the first adjacent image frames in the video stream; the tracking target is the object to be detected with a displacement value greater than a first preset value, a displacement direction towards the focal plane, and a depth information difference greater than a third preset value between the second part of the human key points in the image frame; 2. The target tracking method according to claim 1, characterized in that The method further includes: During tracking, determining the position of the tracking target in the second adjacent image frames. The second adjacent image frames include a third image frame and a fourth image frame, and the third image frame is the image frame in front of the fourth image frame. The position is the position of the depth information; Determining the displacement value and displacement direction between the position of the tracking target in the fourth image frame and the position of the tracking target in the third image frame. The displacement value and the displacement direction are the displacement value and displacement direction of the depth information; If the displacement value is greater than a second preset value and the displacement direction is away from the focal plane, exit the tracking.

3. The target tracking method according to claim 1, characterized in that, The method further includes: During tracking, determining the position of the tracking target in the second adjacent image frames. The second adjacent image frames include a third image frame and a fourth image frame, and the third image frame is the image frame in front of the fourth image frame. The position is the position of the depth information; Determining the displacement value and displacement direction between the position of the tracking target in the fourth image frame and the position of the tracking target in the third image frame. The displacement value and the displacement direction are the displacement value and displacement direction of the depth information; If the displacement value is greater than a second preset value and the displacement direction is away from the focal plane, or the tracking target is not connected to the first part of the human key points, exit the tracking.

4. The target tracking method according to any one of claims 1 to 3, characterized in that: The displacement value of the depth information is the absolute value of the average depth change value of the pixel points.

5. A target tracking device, characterized in that, The device includes: An obtaining unit, configured to obtain the depth information of the image frames in the video stream; A determining unit, configured to determine that the change region between the first adjacent image frames in the video stream is the object to be detected. The first adjacent image frames include a first image frame and a second image frame, and the first image frame is the image frame in front of the second image frame. The change region is the difference region of the depth information; The determining unit is further configured to determine the displacement value and displacement direction between the position of the object to be detected in the second image frame and the position of the object to be detected in the first image frame. The position is the position of the depth information, and the displacement value and displacement direction are the displacement value and displacement direction of the depth information; The determining unit is further configured to select a tracking target, where the tracking target is a to-be-detected object with a displacement value greater than a first preset value and a displacement direction towards the focal plane; The determining unit is further configured to detect human key points in an image frame in the video stream; The determining unit is further configured to determine a change region connected to a first part of the human key points between the first adjacent image frames in the video stream as the to-be-detected object; The determining unit is further configured to select a tracking target, where the tracking target is a to-be-detected object with a displacement value greater than a first preset value, a displacement direction towards the focal plane, and a depth information difference greater than a third preset value between the to-be-detected object and a second part of the human key points in the image frame.

6. The target tracking device according to claim 5, wherein: The determining unit is further configured to, during tracking, determine the position of the tracking target in the second adjacent image frames, where the second adjacent image frames include a third image frame and a fourth image frame, the third image frame is the image frame in front of the fourth image frame, and the position is the position of the depth information; The determining unit is further configured to determine the displacement value and the displacement direction between the position of the tracking target in the fourth image frame and the position of the tracking target in the third image frame, where the displacement value and the displacement direction are the displacement value and the displacement direction of the depth information; The determining unit is further configured to, if the displacement value is greater than a second preset value and the displacement direction is away from the focal plane, exit the tracking.

7. The target tracking device according to claim 5, wherein: The determining unit is further configured to, during tracking, determine the position of the tracking target in the second adjacent image frames, where the second adjacent image frames include a third image frame and a fourth image frame, the third image frame is the image frame in front of the fourth image frame, and the position is the position of the depth information; The determining unit is further configured to determine the displacement value and the displacement direction between the position of the tracking target in the fourth image frame and the position of the tracking target in the third image frame, where the displacement value and the displacement direction are the displacement value and the displacement direction of the depth information; The determining unit is further configured to, if the displacement value is greater than a second preset value and the displacement direction is away from the focal plane, or the tracking target is not connected to the first part of the human key points, exit the tracking.

8. The target tracking device according to any one of claims 5 to 7, wherein: The displacement value of the depth information is the absolute value of the average depth change value of the pixel points.

9. An electronic device, characterized in that, The device includes a processor and a memory, where the memory is used to store program instructions, and when the processor calls the program instructions, the target tracking method according to any one of claims 1 to 4 is implemented.

10. A server, characterized in that, The server includes a processor and a memory, where the memory is used to store program instructions, and when the processor calls the program instructions, the target tracking method according to any one of claims 1 to 4 is implemented.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program, and the program enables a computer device to implement the target tracking method according to any one of claims 1 to 4.

12. A computer program product, characterized in that, The computer program product includes computer-executable instructions stored in a computer-readable storage medium; at least one processor of the device can read the computer-executable instructions from the computer-readable storage medium, and the at least one processor executes the computer-executable instructions to cause the device to perform the target tracking method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Depth information-based multi-target tracking method

    CN102063725A