Moving target detection method, electronic equipment and storage medium
By registering key points and calculating optical flow of terminal devices, the problem of out-of-focus during shooting of moving targets is solved, precise focus pursuit is achieved, and the shooting effect is improved.
Patent Information
- Application Number
- CN202410046026.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-01-10
AI Technical Summary
When users use terminal devices to shoot moving targets, they are prone to loss of focus, resulting in unclear shooting effects and reducing user experience.
By performing key point registration, optical flow calculation and inter-frame differential processing on two consecutive frames of target images, the position and speed of moving targets are accurately detected, and the focus pursuit algorithm is used to achieve accurate focus pursuit.
It improves the shooting clarity of sports targets and improves the shooting experience of users.
Smart Images

Figure CN120339338A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and in particular, to a method for detecting moving targets, an electronic device, and a storage medium. Background Art
[0002] With the continuous development of technology, terminal devices represented by mobile phones and tablet computers are increasingly used in people's lives and work, bringing great convenience to people. For example, people can take pictures, make video calls, or record videos through terminal devices.
[0003] However, currently, in scenarios where users use terminal devices to record videos or take pictures of moving targets (such as vehicles driving on the road), since the main target for focusing is constantly moving, it is very easy to have the problem of taking the moving target out of focus. Summary of the Invention
[0004] To solve the above problems, this application provides a method for detecting moving targets, an electronic device, and a storage medium, aiming to more accurately detect the position and moving speed of the moving target, so as to perform precise focus tracking based on the detected position and moving speed of the moving target, and then solve the problem of taking the moving target out of focus.
[0005] In a first aspect, this application provides a method for detecting moving targets. The method includes: The terminal device first responds to a shooting operation triggered by the user on the terminal device, obtains continuous first target images and second target images (such as two consecutive frames of images containing a moving car), and performs key point registration processing on the first target images and the second target images to obtain the registered first target images and the registered second target images. Then, based on the registered first target images and the registered second target images, a target optical flow map is calculated. And after performing binarization processing on the target optical flow map, the initial connected domain where the moving target is located in the obtained optical flow binarized image is determined. Next, inter-frame difference and binarization processing are performed on the registered first target images and the registered second target images to obtain a difference binarized image, and the previously obtained initial connected domain is projected onto the difference binarized image to filter out the filtered connected domain using a preset threshold as the position area where the moving target is located. Furthermore, the filtered connected domain can also be projected onto the target optical flow map to more accurately calculate the speed of the moving target within the filtered connected domain.
[0006] It can be seen that in the above motion target detection method, when a user holds a terminal device such as a mobile phone to capture a motion target (such as a moving car) (or the terminal device such as a mobile phone is placed under a tripod to capture the motion target), the terminal device in this embodiment first performs precise registration of key points on two consecutive target images, then calculates the optical flow connectivity domain and inter-frame difference (auxiliary information) of the two target images after precise registration, and then through comprehensive processing of the two, more accurately determines the detection result of the motion target (including position and speed). Subsequently, the focus tracking algorithm can be used to achieve precise focus tracking of the motion target according to the detection result of the motion target (including position and speed), optimizing the focus tracking result, thus solving the problem of shooting a motion target (such as a moving car) out of focus, and further improving the user's shooting experience.
[0007] In a possible implementation manner, performing key point registration processing on the first target image and the second target image to obtain the registered first target image and the registered second target image may include: using a mean filter algorithm to perform denoising processing on the first target image and the second target image to obtain the denoised first target image and the denoised second target image; using a preset key point detection algorithm to detect key points on the denoised first target image and the denoised second target image to obtain the key points included in the denoised first target image and the denoised second target image respectively; then using the key points included in the denoised first target image and the denoised second target image respectively to perform registration processing on the denoised first target image and the denoised second target image, and calculating the circumscribed rectangle according to the key points as the common area; performing cropping processing on the denoised first target image and the denoised second target image according to the common area to obtain the cropped and registered first target image and the cropped and registered second target image, thereby removing the influence of noise and improving the efficiency and accuracy of subsequent image processing.
[0008] In a possible implementation manner, the preset key point detection algorithm may be an ORB key point detection algorithm to more accurately detect the key points included in the denoised first target image and the denoised second target image respectively.
[0009] In a possible implementation manner, calculating the target optical flow map according to the registered first target image and the registered second target image may include: using the FarneBack algorithm to calculate the dense optical flow of the registered first target image and the registered second target image to obtain the target optical flow map. Thereby, the accuracy of the calculated target optical flow map can be improved.
[0010] In a possible implementation, after calculating the target optical flow map based on the registered first target image and the registered second target image, the method further includes: performing grayscale processing on the target optical flow map to obtain an optical flow grayscale map corresponding to the target optical flow map; performing binarization processing on the optical flow grayscale map by using a preset image binarization algorithm to obtain an optical flow binarized image corresponding to the optical flow grayscale map. This is to more accurately determine the initial connected region where the moving target is located in the subsequent process.
[0011] In a possible implementation, the preset image binarization algorithm can be the OTSU algorithm of the maximum inter-class variance method to obtain a more accurate optical flow binarized image.
[0012] In a possible implementation, after performing binarization processing on the target optical flow map, determining the initial connected region where the moving target is located in the obtained optical flow binarized image may include: performing morphological operations of dilation and / or erosion on the optical flow binarized image to filter out noise, obtaining a smoother optical flow binarized image after the morphological operation, and determining the initial connected region where the moving target is located according to the coordinates of the pixel points in the optical flow binarized image after the morphological operation, so as to improve the accuracy of the initial position result of the moving target.
[0013] In a possible implementation, performing inter-frame difference and binarization processing on the registered first target image and the registered second target image to obtain a difference binarized image may include: calculating the inter-frame difference between the registered first target image and the registered second target image, and statistically obtaining the grayscale histogram of the inter-frame difference; performing binarization processing on the grayscale histogram by using a preset dynamic threshold to obtain a binarized inter-frame difference image; performing morphological operations of dilation and / or erosion on the binarized inter-frame difference image to filter out noise, obtaining a smoother inter-frame difference image after the morphological operation, and using it as the difference binarized image.
[0014] In a possible implementation, projecting the initial connected region onto the difference binarized image to filter out the connected regions after filtering by using a preset threshold as the position region where the moving target is located may include: projecting the initial connected region onto the difference binarized image, calculating the density of the difference points falling within the initial connected region, and filtering out the initial connected regions with a density less than the preset threshold, and screening out the filtered connected regions as the position region where the moving target is located, thereby improving the accuracy of the detected position of the moving target.
[0015] In a possible implementation, the preset threshold can be 0.5.
[0016] In a possible implementation, projecting the filtered connected component onto the target optical flow map and calculating the speed of the moving target within the filtered connected component may include: projecting the filtered connected component onto the target optical flow map and calculating the average optical flow of all points falling within the filtered connected component as the speed of the moving target within the filtered connected component, thereby improving the accuracy of the detected moving target's speed.
[0017] In a second aspect, the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor is configured to call and execute the computer program to implement the moving target detection method described in any one of the above first aspects.
[0018] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is run by the processor of an electronic device, it is used to implement the moving target detection method described in any one of the above first aspects.
[0019] In a fourth aspect, the present application provides a computer program product. When the computer program product runs on a computer, it causes the computer to execute the moving target detection method described in any one of the above first aspects. Description of the Drawings
[0020] Figure 1 It is a schematic diagram of the scenario provided by the embodiment of the present application;
[0021] Figure 2 It is a schematic diagram of the electronic device provided by the embodiment of the present application;
[0022] Figure 3 It is a software structure block diagram of the electronic device provided by the embodiment of the present application;
[0023] Figure 4 It is a flowchart of the moving target detection method provided by the embodiment of the present application;
[0024] Figure 5 It is an example diagram of performing key point registration processing on the first target image and the second target image to obtain the registered first target image and the registered second target image provided by the embodiment of the present application;
[0025] Figure 6 It is an example diagram of the determination process of the primary connected component where the moving target is located in the optical flow binary image provided by the embodiment of the present application;
[0026] Figure 7 It is an example diagram of the determination process of the position area where the moving target is located provided by the embodiment of the present application;
[0027] Figure 8Schematic diagram of the determination process of the speed of a moving object provided by an embodiment of the present application. Detailed implementation manners
[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the forms such as "one or more", unless clearly indicated to the contrary in the context.
[0029] Reference to "one embodiment" or "some embodiments" etc. described in this specification means that a specific feature, structure, or characteristic described in combination with the embodiment is included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0030] The "multiple" involved in the embodiments of the present application means greater than or equal to two. It should be noted that in the description of the embodiments of the present application, the terms such as "first" and "second" are only used for the purpose of distinguishing descriptions and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying an order.
[0031] To enable those skilled in the art to better understand the solution of the present application, the application scenario of the technical solution of the present application will be described first.
[0032] See Figure 1 , which shows a schematic diagram of a scenario provided by an embodiment of the present application.
[0033] In this example scenario, a certain user holds a mobile phone to take pictures. It should be noted that various function buttons for improving the shooting effect will be deployed on the display interface of the mobile phone camera for the user to click to meet the personalized shooting needs of each user, such as adjusting the ratio of the captured image, whether to turn on the flash, etc. As Figure 1As shown, when a user holds a mobile phone to take pictures of a moving car, since the main object of focus is the constantly moving car, in order to improve the shooting clarity, the user can click the "Motion Focus" button on the display interface of the mobile phone camera to use multiple tracking frames (such as Figure 1 the 8 small boxes on the car in the upper right corner) to focus on the moving car in multiple consecutive frames of images taken.
[0034] However, currently when a user uses a terminal device to take pictures of a moving object, such as Figure 1 when the user uses a mobile phone to take pictures of a moving car, even if multiple tracking frames (such as Figure 1 the 8 small boxes on the car in the upper right corner) are used to focus on the moving car, it is still easy to have the problem of taking the moving object out of focus, such as the problem of taking the Figure 1 car in [a certain situation] "blurry", which reduces the user's shooting experience.
[0035] To overcome the above technical problems, the present application provides a moving object detection method, an electronic device and a storage medium. It can more accurately detect the position and moving speed of the moving object, so as to perform precise focusing according to the detected position and moving speed of the moving object, thus solving the problem of taking the moving object out of focus, and further improving the user's shooting experience.
[0036] The moving object detection method provided by the embodiments of the present application can be applied to electronic devices (i.e., terminal devices) such as mobile phones, tablet computers, personal digital assistants (PDAs), desktop, laptop, notebook computers, ultra-mobile personal computers (UMPCs), handheld computers, netbooks, and wearable devices.
[0037] To enable those skilled in the art to more clearly understand the frame sending display time anchoring method provided by the present application, the hardware architecture and software system architecture of the electronic device will be introduced in detail below.
[0038] Refer to Figure 2 , which shows a schematic diagram of the electronic device provided by the embodiments of the present application.
[0039] As Figure 2 shown, the electronic device 200 may include a processor 210, a mobile communication module 220, a wireless communication module 230, a display screen 240, an internal memory 241, a camera 242, an audio module 243, a speaker 243A, a receiver 243B, a microphone 243C, a headphone jack 243D, an antenna group 1, and an antenna group 2.
[0040] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than shown in the figures, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0041] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions. For example, it can determine a camera parameter acquisition instruction according to the received shooting request of the user, etc., to detect the position and speed of the captured moving target.
[0042] A memory may also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.
[0043] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0044] The internal memory 241 may be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 241 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound collection function, an image capture function, etc.). The data storage area may store data created during the use of the electronic device 200 (such as audio data, image data, etc.). In addition, the internal memory 241 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the electronic device 200 by running the instructions stored in the internal memory 241, and / or the instructions stored in the memory provided in the processor.
[0045] In some embodiments, the internal memory 241 stores instructions for executing a moving object detection method. The processor 210 can, by executing the instructions stored in the internal memory 241, implement continuously capturing two frames of images (defined as the first target image and the second target image) in response to a shooting operation triggered by the user on the terminal device, and then performing key point registration processing on the first target image and the second target image to obtain the registered first target image and the registered second target image. Next, according to the registered first target image and the registered second target image, a target optical flow map is calculated; and after performing binarization processing on the target optical flow map, the primary connected region where the moving object is located in the obtained optical flow binarized image is determined, and then frame difference and binarization processing are performed on the registered first target image and the registered second target image to obtain a difference binarized image; and the primary connected region is projected onto the difference binarized image to filter out the filtered connected region by using a preset threshold as the position region where the moving object is located. Furthermore, the filtered connected region can be projected onto the target optical flow map to calculate the speed of the moving object within the filtered connected region to obtain the detection result of the moving object (including position and speed).
[0046] The display screen 240 is used to display images, videos, etc., such as displaying the captured images or videos containing moving objects, for example Figure 1 In the display screen of the user's mobile phone, the captured images or videos containing a moving car can be displayed. The display screen 540 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 500 can include one or N display screens 240, where N is a positive integer greater than 1.
[0047] The camera 242 is used to capture static images or videos. For example, after the user holds the electronic device 200, the camera 242 of the electronic device 200 can be used to capture an image frame containing a moving object. Figure 1After the user holds the mobile phone, the camera of the mobile phone can be used to capture images or videos containing a moving car. In some embodiments, the electronic device 200 may include one or N cameras 242, where N is a positive integer greater than 1.
[0048] The electronic device 200 realizes the display function through the GPU, the display screen 240, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 240 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change the display information.
[0049] The electronic device 200 can realize the audio function through the audio module 243, the speaker 243A, the receiver 243B, the microphone 243C, the headphone jack 243D, and the application processor, etc. For example, the input and output of voices such as music playback and recording.
[0050] The audio module 243 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 243 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 243 may be disposed in the processor 210, or some functional modules of the audio module 243 may be disposed in the processor 210.
[0051] The speaker 243A, also called the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 200 can listen to music or hands-free calls through the speaker 243A.
[0052] The receiver 243B, also called the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 200 answers a call or a voice message, the receiver 243B can be used to listen to the voice by bringing it close to the human ear.
[0053] The microphone 243C, also called the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak close to the microphone 243C with the mouth to input the sound signal into the microphone 243C. The electronic device 200 may be provided with at least one microphone 243C. In some other embodiments, the electronic device 200 may be provided with two microphones 243C, which can not only collect sound signals but also achieve a noise reduction function. In some other embodiments, the electronic device 200 may also be provided with three, four or more microphones 243C to achieve sound signal collection, noise reduction, and also identify the sound source to achieve functions such as directional recording.
[0054] The headphone jack 243D is used to connect a wired headphone without restricting the standard attributes of the interface.
[0055] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is only for illustrative purposes and does not constitute a limitation on the structure of the electronic device 200.
[0056] The wireless communication function of the electronic device 200 can be implemented by the antenna 1, the antenna 2, the mobile communication module 220, the wireless communication module 230, the modulation and demodulation processor, and the baseband processor, etc.
[0057] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 200 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: the antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0058] The mobile communication module 220 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device 200. The mobile communication module 220 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 220 can receive electromagnetic waves by the antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 220 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 220 can be disposed in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 220 and at least some modules of the processor 210 can be disposed in the same device.
[0059] The wireless communication module 230 may provide solutions for wireless communications applied to the electronic device 200, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 230 may be one or more devices integrating at least one communication processing module. The wireless communication module 230 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 210. The wireless communication module 230 may also receive the signals to be sent from the processor 210, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation.
[0060] In addition, on top of the above components, the electronic device 200 runs an operating system. For example, iOS operating system, Android operating system, Windows operating system, etc. Application programs can be installed and run on the operating system.
[0061] See Figure 3 , which shows a schematic diagram of the software structure of the electronic device provided by the embodiments of the present application.
[0062] The software system of the electronic device 200 may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present application, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device 200.
[0063] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom, namely the application layer (APK), the application framework layer (Framework), the hardware abstraction layer (HAL), the driver layer, and the hardware layer.
[0064] The application layer may include a series of application packages (APPs). As Figure 3 shown, the application packages may include applications such as cameras, calls, navigation, WLAN, Bluetooth, and galleries. When the user holds the electronic device 200 to target a moving object (such as Figure 1When taking pictures of the moving car shown, the camera application can communicate with the camera-related devices in the camera access interface of the framework layer to request camera functions and obtain image data, etc.
[0065] The application framework layer (framework layer) provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. Such as Figure 3 As shown, the application framework layer may include a window manager, a notification manager, a resource manager, a camera access interface (including but not limited to camera management, camera devices, etc.). The application framework layer can be used to implement the interaction between the camera service and the camera API. That is, it provides a unified interface, enabling different camera hardware to interact with different camera applications. The framework layer is also used to handle many general aspects of camera functions, such as autofocus, motion focus, exposure control, etc.
[0066] Among them, the window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0067] The notification manager enables the application to display notification information in the status bar. It can be used to convey notification-type messages, which can disappear automatically after a short stay without user interaction. For example, the notification manager is used to inform that the download is complete, message reminders, etc.
[0068] The resource manager provides various resources for the application, such as localized strings, icons, pictures, layout files, video files, and so on.
[0069] In this embodiment, the hardware abstraction layer (HAL) provides a set of standard interfaces, enabling the framework layer to communicate with camera hardware from various different manufacturers without having to understand the underlying hardware details. The hardware abstraction layer (HAL) stores the hardware abstraction layer and the camera algorithm library, etc. It should be noted that the moving target detection algorithm provided in this application is stored in the camera algorithm library of the HAL layer, such as Figure 3 As shown.
[0070] Among them, the moving target detection algorithm is used to process two consecutive frames of images containing moving targets obtained by the camera (such as Figure 1 the image of the moving car taken by the user using the mobile phone camera) to more accurately detect the moving target (such as Figure 1The position and moving speed of a car in motion. Specifically, after acquiring two consecutive frames of images (represented by a first target image and a second target image) in response to a user's shooting operation on the camera trigger, key point registration processing can be first performed on the first target image and the second target image to obtain the registered first target image and the registered second target image. Then, based on the registered first target image and the registered second target image, a target optical flow map is calculated; and after performing binarization processing on the target optical flow map, a primary connected region where the moving target is located in the obtained optical flow binarized image is determined. Then, inter-frame difference and binarization processing are performed on the registered first target image and the registered second target image to obtain a difference binarized image; and the primary connected region is projected onto the difference binarized image to filter out the filtered connected region using a preset threshold as the position region where the moving target is located. Furthermore, the filtered connected region can be projected onto the target optical flow map to calculate the speed of the moving target within the filtered connected region, and the detection result of the moving target (including the position and moving speed) is obtained.
[0071] The hardware layer may include the hardware components of the aforementioned electronic device. Exemplarily, Figure 3 sensors, an image signal processor, a digital signal processor, and a graphics processor are shown.
[0072] Among them, the sensor is used for image exposure processing, etc.
[0073] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 200 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0074] In this way, through the application layer, framework layer, HAL layer, driver layer, and hardware layer, the interoperability of different camera applications and hardware devices on the Android platform can be achieved, and the purpose of more accurately detecting the position and moving speed of the moving target can be achieved.
[0075] The technical solutions involved in the following embodiments can all be implemented in an electronic device having the above-mentioned hardware architecture and software architecture.
[0076] Next, the specific implementation process of the moving target detection method provided in this application will be introduced in detail:
[0077] As Figure 4 shown, the specific implementation process of this moving target detection method may include the following steps S401 - S405:
[0078] S401: In response to a shooting operation triggered by a user on a terminal device, obtain a first target image and a second target image; the first target image and the second target image are two consecutive frames of images captured by the terminal device.
[0079] In this embodiment, when a user uses a terminal device such as a mobile phone to shoot a moving target (such as Figure 1 the shown moving car), in response to the shooting operation triggered by the user on the terminal device, such as in response to the user's click operation on the "Camera" icon on the mobile phone display screen, the terminal device will immediately start the camera APP, and will also start the camera (such as the camera pre-loaded in the terminal device) in the background to capture an image containing the moving target, and take any two consecutive frames of images captured by the terminal device as the target images to be detected, and define them as the first target image and the second target image respectively for performing the subsequent step S402.
[0080] It should be noted that this embodiment does not limit the types of the first target image and the second target image. For example, the target image can be a color image composed of the three primary colors of red (R), green (G), and blue (B), or a grayscale image, etc. It should be noted that the subsequent embodiments will introduce the first target image and the second target image as RGB images. In addition, this application does not limit the acquisition method of the first target image and the second target image. For example, two consecutive frames of images can be obtained from the camera preview stream in the tripod or hand-held state as the first target image and the second target image; or, any two consecutive frames of images in the pre-recorded video stream can be used as the first target image and the second target image, etc.
[0081] S402: Perform key point registration processing on the first target image and the second target image to obtain the registered first target image and the registered second target image.
[0082] After the terminal device obtains the first target image and the second target image to be detected by using the camera through step S401, it can further use existing or future key point detection algorithms to detect the image key points of the first target image and the second target image, and then use the obtained image key points to perform registration processing on the first target image and the second target image to eliminate the influence of camera jitter, ensure that the key points of the two frames of target images are aligned one by one, and obtain the registered first target image and the registered second target image for performing the subsequent step S403.
[0083] It should be noted that, in order to improve the accuracy of key point registration for the first target image and the second target image, an optional implementation method is that the implementation process of this step S402 can include the following steps S4021 - S4024:
[0084] S4021: Denoise the first target image and the second target image using the mean filtering algorithm to obtain the denoised first target image and the denoised second target image.
[0085] It should be noted that to improve the accuracy of image processing and eliminate the influence of noise in this application, after obtaining the first target image and the second target image, the mean filtering algorithm available now or in the future can be used to denoise the first target image and the second target image, making the images smoother, and obtaining the denoised first target image and the denoised second target image for subsequent step S4022.
[0086] For example: as Figure 5 shown, assume that the first target image and the second target image are respectively two images as shown in Figure 5 (a). If the moving target to be detected is a hand, then use the mean filtering algorithm to denoise them respectively, and two images as shown in Figure 5 (b) can be obtained.
[0087] S4022: Detect key points of the denoised first target image and the denoised second target image using a preset key point detection algorithm to obtain the key points contained in the denoised first target image and the denoised second target image respectively.
[0088] After the terminal device obtains the denoised first target image and the denoised second target image through step S4021, it can further use a preset key point detection algorithm to detect key points of the first target image and the second target image, and obtain the key points contained in the denoised first target image and the denoised second target image respectively (referring to the special points representing the key information of the denoised first target image and the denoised second target image) for subsequent step S4023.
[0089] Among them, the specific content of the preset key point detection algorithm is not limited and can be set according to the actual situation and empirical values. For example, an optional implementation method is to set the preset key point detection algorithm as the ORB key point detection algorithm, that is, the ORB key point detection algorithm can be used to detect key points of the denoised first target image and the denoised second target image to obtain the image key points contained in them respectively.
[0090] For example: as Figure 5 shown, based on the above example, still assume that the first target image and the second target image are respectively two images as shown in Figure 5 (a). If the moving target to be detected is a hand, and the mean filtering algorithm is used to denoise them respectively, two images as shown in Figure 5After the two denoised images shown in (b), the ORB key point detection algorithm or the like can be used to perform key point detection on the two images respectively, and obtain two images containing the key points of the images as shown in Figure 5 (c) in
[0091] S4023: Use the key points contained in the denoised first target image and the denoised second target image respectively to perform registration processing on the denoised first target image and the denoised second target image, and calculate the circumscribed rectangle according to the key points as the common area.
[0092] After the terminal device obtains the denoised first target image and the denoised second target image through step S4022, further, the key points contained in the denoised first target image and the denoised second target image can be used to perform registration processing on the denoised first target image and the denoised second target image to ensure that the key points of the two can be aligned one by one, and the positions of the moving targets can be overlapped as much as possible. And ensure that the moving target to be detected can be located in the middle position of the image as much as possible, and then perform the subsequent step S4024.
[0093] Among them, the specific content of the registration algorithm is not limited and can be set according to the actual situation and empirical values.
[0094] For example: as Figure 5 shown, based on the above example, it is still assumed that the first target image and the second target image are respectively two images as shown in Figure 5 (a) in Figure 5 where the moving target to be detected is a hand, and the mean filter algorithm is used to perform denoising processing on the two respectively to obtain two denoised images as shown in Figure 5 (b) in Figure 5 and the ORB key point detection algorithm or the like is used to perform key point detection on the two respectively to obtain two images containing the key points of the images as shown in Figure 5 (c) in
[0095] S4024: Perform cropping processing on the denoised first target image and the denoised second target image according to the common area to obtain the cropped and registered first target image and the cropped and registered second target image.
[0096] After the terminal device obtains the common area of the denoised first target image and the denoised second target image through step S4023, it can use this common area to crop the denoised first target image and the denoised second target image, obtaining the cropped and registered first target image and the cropped and registered second target image, which are used as the final registered first target image and the registered second target image to execute the subsequent step S403, further removing the influence of noise and improving the efficiency and accuracy of subsequent image processing.
[0097] For example: As Figure 5 shown, based on the above example, it is still assumed that the first target image and the second target image are respectively two images as shown in Figure 5 (a). Among them, the moving target to be detected is a hand, and the mean filter algorithm is used to perform denoising processing on both of them respectively, obtaining two denoised images as shown in Figure 5 (b), and the ORB key point detection algorithm and other algorithms are used to perform key point detection on both of them respectively, obtaining two images containing image key points as shown in Figure 5 (c), and using the key points of the two images in Figure 5 (c) to perform registration processing on these two images, and calculating the circumscribed rectangle according to the key points as the common area, obtaining two images as shown in Figure 5 (d) and their common area represented by a square box. After that, this common area can be used for cropping, obtaining two images as shown in Figure 5 (e), which are used as the registered first target image and the registered second target image.
[0098] S403: According to the registered first target image and the registered second target image, calculate the target optical flow map; and after performing binarization processing on the target optical flow map, determine the primary connected domain where the moving target is located in the obtained optical flow binarized image.
[0099] After the terminal device obtains the registered first target image and the registered second target image through step S402, further, it can use existing or future optical flow algorithms (the specific content is not limited) to perform optical flow calculation on the registered first target image and the registered second target image, obtaining the corresponding color optical flow map of the two, and defining it as the target optical flow map. And this target optical flow map is a color RGB image, and the brightness in the map represents the speed magnitude of the moving target. Then, after performing binarization processing on this target optical flow map, it can determine the primary connected domain where the moving target is located in the obtained optical flow binarized image, so as to execute the subsequent step S404.
[0100] Specifically, in an optional implementation, in order to improve the accuracy of the calculated target optical flow map, the FarneBack algorithm can be used to calculate the dense optical flow of the registered first target image and the registered second target image, thereby obtaining the target optical flow map.
[0101] Among them, the FarneBack algorithm refers to a motion estimation algorithm based on all pixel points in two consecutive frames of images. It realizes optical flow tracking through the displacement vectors of all pixel points in two consecutive frames of images, and has a good tracking effect on moving targets. Specifically, the main implementation process of the Farneback algorithm is as follows: The coordinate position of each pixel point is expanded into a polynomial through the neighborhood information of each pixel point (the weight is determined by the pixel value size and position of the neighborhood pixel points), obtaining a polynomial with the original coordinates (x0, y0) as the independent variable and the new coordinates (x, y) as the dependent variable, and substituting the coordinate data to calculate the movement amounts (dx, dy) of this pixel point in the x and y directions. In this way, the displacement vectors of each pixel point in two consecutive frames of images are obtained, including amplitude and phase. If the amplitude and phase information of each pixel point displacement vector are converted into H, S, and V three-channel information, the motion of the moving object can be visually observed in consecutive frame images, thus realizing the dense optical flow tracking of the moving object.
[0102] Furthermore, after obtaining the target optical flow map, first, the target optical flow map can be grayscale processed to obtain the optical flow grayscale map corresponding to the target optical flow map. The specific calculation formula is as follows:
[0103]
[0104]
[0105]
[0106] Among them, g represents the gray level of the optical flow grayscale map; ω represents the current angular velocity; V y represents the velocity of the current point in the y-axis direction; V x represents the velocity of the current point in the x-axis direction; V max and V min respectively represent the maximum velocity and the minimum velocity of the current frame.
[0107] Then, a preset image binarization algorithm can be used to binarize the optical flow grayscale image to obtain an optical flow binarized image corresponding to the optical flow grayscale image. The specific content of the preset image binarization algorithm is not limited and can be set according to the actual situation and empirical values. For example, an optional implementation method is to set the preset image binarization algorithm to the Otsu method (OTSU algorithm), that is, the OTSU algorithm can be used to binarize the optical flow grayscale image to obtain an optical flow binarized image corresponding to the optical flow grayscale image.
[0108] For example: As Figure 6 shown, based on the above example, assume Figure 6 that the two images shown in (a) are Figure 5 the registered first target image and the registered second target image shown in (e). Then, the Farneback algorithm can be used to calculate the corresponding target optical flow map for the two, such as Figure 6 the image shown in (b). Further, the target optical flow map can be grayscaled to obtain Figure 6 the optical flow grayscale image shown in (c). Then, the OTSU algorithm can be used to binarize the optical flow grayscale image to obtain Figure 6 the optical flow binarized image shown in (d).
[0109] On this basis, an optional implementation method is that after binarizing the target optical flow map to obtain an optical flow binarized image, in order to improve the accuracy of the image processing result, further morphological operations of dilation and / or erosion (the specific operation process and operation method are not limited) can be performed on the obtained optical flow binarized image to obtain an optical flow binarized image after morphological operations, and the primary connected region where the moving target is located can be determined according to the coordinates of the pixel points in the optical flow binarized image after morphological operations. For example, the set of coordinates of the white pixel points in the optical flow binarized image after morphological operations can be used to form the primary connected region where the moving target is located.
[0110] For example: As Figure 6 shown, based on the above example, still assume Figure 6 that the two images shown in (a) are Figure 5 the registered first target image and the registered second target image shown in (e), and the Farneback algorithm is used to calculate the corresponding target optical flow map for the two, such as Figure 6 the image shown in (b), and the target optical flow map is grayscaled to obtain Figure 6 the optical flow grayscale image shown in (c). Then, the OTSU algorithm is used to binarize the optical flow grayscale image to obtain Figure 6 the optical flow binarized image shown in (d). After that, morphological operations of dilation and / or erosion can be performed on the optical flow binarized image to obtainFigure 6 The optical flow binary image after the morphological operation shown in (e) of the figure. Then, the rectangle where the coordinate set of white pixel points in the optical flow binary image after the morphological operation is located is used to frame the position of the primary connected region where the moving target is located, as shown in Figure 6 (f) of the figure.
[0111] S404: Perform inter-frame difference and binarization processing on the registered first target image and the registered second target image to obtain a difference binary image; and project the primary connected region onto the difference binary image to filter out the filtered connected region using a preset threshold as the position area where the moving target is located.
[0112] After the terminal device obtains the registered first target image and the registered second target image through step S402, it not only needs to perform processing such as optical flow, grayscale, and binarization on them to determine the primary connected region where the moving target is located, but also needs to perform inter-frame difference (the specific algorithm content and calculation method are not limited) and binarization processing on the registered first target image and the registered second target image to obtain a difference binary image; and project the primary connected region onto the difference binary image to filter out the filtered connected region using a preset threshold as the position area where the moving target is located for subsequent step S405. Among them, the specific content of the preset dynamic threshold is not limited and can be set according to the actual situation and empirical values.
[0113] Specifically, an optional implementation method is that in order to improve the recognition accuracy of the position area where the moving target is located, the implementation process of "performing inter-frame difference and binarization processing on the registered first target image and the registered second target image to obtain a difference binary image" in the above step S404 may include: First, calculate the inter-frame difference between the registered first target image and the registered second target image, and statistically obtain the grayscale histogram of the inter-frame difference.
[0114] Among them, the inter-frame difference refers to detecting moving target objects in an image by comparing the pixel value differences between adjacent frame images. Specifically, the main implementation process of the inter-frame difference is: First, obtain two adjacent frame images (such as the registered first target image and the registered second target image), and convert them into grayscale images. Then, perform a difference operation on the two grayscale images to obtain a grayscale histogram of the inter-frame difference. In this difference image, the larger the pixel value, the greater the change of the pixel between the two frame images, which may be caused by the appearance or movement of a moving object.
[0115] Then, use a preset dynamic threshold to perform binarization processing on the grayscale histogram to obtain a binarized inter-frame difference image.
[0116] Among them, it should be noted that in order to detect moving target objects, it is necessary to perform threshold processing on the grayscale histogram of frame difference. Threshold processing is to compare the pixel values in the difference image with a preset threshold. If the pixel value is greater than the threshold, it is considered that the pixel belongs to the moving target object. Thus, the moving object can be separated from the background to achieve the detection of the moving target object. Specifically, after obtaining the grayscale histogram of frame difference through statistics, it can be sorted according to the frequency, and the top two are taken as the preset dynamic threshold to binarize the frame difference image (i.e., the grayscale histogram) to obtain the binarized frame difference image.
[0117] Next, morphological operations of dilation and / or erosion can be performed on the binarized frame difference image to filter out noise and obtain the frame difference image after morphological operations as the differential binarized image.
[0118] Furthermore, after obtaining the differential binarized image, first, the primary connected region where the moving target obtained through step S403 is located can be projected onto this differential binarized image, and the differential point density within the primary connected region is calculated. The specific calculation formula is as follows:
[0119]
[0120] Among them, ρ represents the differential point density within the primary connected region; d i represents the value of the differential binarized image within the connected region; S represents the area of the primary connected region.
[0121] Then, the primary connected regions with a density ρ less than the preset threshold can be filtered out, and the filtered connected regions are selected as the position region where the moving target is located. The specific content of the preset threshold is not limited and can be set according to the actual situation and empirical values. For example, the preset threshold can be set to 0.5, that is, the primary connected regions with a density ρ less than 0.5 will be filtered out, and the primary connected regions with a density ρ greater than 0.5 will be retained, so as to select the filtered connected regions as the position region where the moving target is located.
[0122] For example: As Figure 7 shown, based on the above example, assume Figure 7 that the two images shown in (a) are Figure 5 the registered first target image and the registered second target image shown in (e), then the frame difference between the two can be calculated, and the grayscale histogram of the frame difference is obtained through statistics, as Figure 7 shown in (b), and then the grayscale histogram is binarized using the preset dynamic threshold to obtain the binarized frame difference image, as Figure 7As shown in (c), morphological operations of dilation and / or erosion are further performed on the binarized inter-frame difference image to obtain an inter-frame difference image after morphological operations, which is used as a differential binarized image, as shown in Figure 7 (d). Further, the initially selected connected regions can be projected onto the differential binarized image to filter out the filtered connected regions using a preset threshold as the position area where the moving object (i.e., the hand) is located, as shown in Figure 7 (e).
[0123] S405: Project the filtered connected regions onto the target optical flow map and calculate the speed of the moving object within the filtered connected regions.
[0124] After the terminal device filters out the filtered connected regions as the position area where the moving object is located through step S404, the filtered connected regions can be further projected onto the target optical flow map calculated through step S403, and the average optical flow value of all points falling within the filtered connected regions is calculated as the moving speed of the moving object within the filtered connected regions, so that the detection result of the moving object (i.e., including the position and moving speed of the moving object) can be formed in combination with the determined position of the moving object.
[0125] It should be noted that in a possible implementation, there may be multiple moving objects in the first target image and the second target image at the same time. For example, Figure 1 there may be multiple moving cars in the figure, or there may be pedestrians walking, etc. Since the moving object that the same terminal device focuses on at the same time is often the moving object with the fastest speed in the captured image, after calculating the moving speeds of each moving object in the first target image and the second target image through the above step S405, as shown in Figure 8 , the moving speeds of each moving object can be sorted in descending order, and the moving object with the maximum moving speed is selected from the sorting result as the focusing subject, so as to use the autofocus algorithm to accurately focus on the moving object according to the detected position and speed of the moving object, and then solve the problem of shooting the moving object out of focus.
[0126] In this way, when the user holds a terminal device such as a mobile phone to shoot a moving object (such as Figure 1 the moving car in the figure), or when the terminal device such as a mobile phone is placed under a tripod to shoot a moving object, the terminal device can more accurately detect the position and moving speed of the moving object (such as Figure 1 the moving car in the figure) by executing the above steps S401 - S405, so as to realize the accurate focusing on the moving object according to the detected position and moving speed of the moving object (such as Figure 1 the moving car in the figure), and then solve the problem of shooting the moving object out of focus. Figure 1Precise focus tracking of a moving car in the scene), optimize the focus tracking result, thus solving the problem of shooting a moving target (such as a moving car in the scene) out of focus, and then improving the user shooting experience. Figure 1 The problem of shooting a moving car in the scene out of focus is solved, thereby improving the user shooting experience.
[0127] In addition, the embodiment of the present application also provides an electronic device (i.e., a terminal device). For the hardware structure and software framework of the electronic device, reference can be made to Figure 2 and Figure 3 the corresponding descriptions. The electronic device includes a memory and a processor. The memory stores a computer program, and the processor is configured to call and execute the computer program to implement the moving target detection method provided in the above description.
[0128] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by the processor of the terminal device, it is used to implement the moving target detection method provided in the above description.
[0129] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting moving targets, characterized in that, Applied to a terminal device, the method includes: In response to a shooting operation triggered by a user on the terminal device, obtaining a first target image and a second target image; the first target image and the second target image are two consecutive frames of images captured by the terminal device; Performing key point registration processing on the first target image and the second target image to obtain a registered first target image and a registered second target image; Calculating a target optical flow map based on the registered first target image and the registered second target image; and after performing binarization processing on the target optical flow map, determining a primary connected domain where a moving target is located in the obtained optical flow binarized image; Performing inter-frame difference and binarization processing on the registered first target image and the registered second target image to obtain a difference binarized image; and projecting the primary connected domain onto the difference binarized image to filter out a filtered connected domain using a preset threshold as the position area where the moving target is located; Projecting the filtered connected domain onto the target optical flow map and calculating the speed of the moving target within the filtered connected domain.
2. The method according to claim 1, wherein The performing key point registration processing on the first target image and the second target image to obtain a registered first target image and a registered second target image includes: Using a mean filtering algorithm to perform denoising processing on the first target image and the second target image to obtain a denoised first target image and a denoised second target image; Using a preset key point detection algorithm to perform key point detection on the denoised first target image and the denoised second target image to obtain the key points included in the denoised first target image and the denoised second target image respectively; Using the key points included in the denoised first target image and the denoised second target image respectively to perform registration processing on the denoised first target image and the denoised second target image, and calculating a circumscribed rectangle based on the key points as a common area; Performing cropping processing on the denoised first target image and the denoised second target image according to the common area to obtain a cropped and registered first target image and a cropped and registered second target image.
3. The method according to claim 2, characterized in that, The preset key point detection algorithm is an ORB key point detection algorithm.
4. The method according to claim 1, wherein The calculating a target optical flow map based on the registered first target image and the registered second target image includes: Using the FarneBack algorithm to calculate the dense optical flow of the registered first target image and the registered second target image to obtain a target optical flow map.
5. The method according to claim 1, wherein After calculating the target optical flow map based on the registered first target image and the registered second target image, the method further includes: Performing grayscale processing on the target optical flow map to obtain an optical flow grayscale map corresponding to the target optical flow map; Using a preset image binarization algorithm to perform binarization processing on the optical flow grayscale map to obtain an optical flow binarized image corresponding to the optical flow grayscale map.
6. The method according to claim 5, wherein The preset image binarization algorithm is the OTSU algorithm of maximum inter-class variance method.
7. The method according to claim 1 or 5, characterized in that After performing binarization processing on the target optical flow map, determining a primary connected region where a moving target is located in the obtained optical flow binarized image, includes: Performing morphological operations of dilation and / or erosion on the optical flow binarized image to obtain an optical flow binarized image after morphological operations, and determining a primary connected region where a moving target is located according to the coordinates of pixel points in the optical flow binarized image after morphological operations.
8. The method according to claim 1, wherein Performing inter-frame difference and binarization processing on the registered first target image and the registered second target image to obtain a difference binarized image, includes: Calculating the inter-frame difference between the registered first target image and the registered second target image, and statistically obtaining a gray-level histogram of the inter-frame difference; Performing binarization processing on the gray-level histogram using a preset dynamic threshold to obtain a binarized inter-frame difference image; Performing morphological operations of dilation and / or erosion on the binarized inter-frame difference image to obtain an inter-frame difference image after morphological operations, as the difference binarized image.
9. The method according to claim 1 or 8, characterized in that, Projecting the primary connected region onto the difference binarized image to filter out a filtered connected region using a preset threshold as the position region where the moving target is located, includes: Projecting the primary connected region onto the difference binarized image, calculating the density of difference points falling within the primary connected region, and filtering out primary connected regions with a density less than the preset threshold to select a filtered connected region as the position region where the moving target is located.
10. The method according to claim 9, wherein The preset threshold is 0.
5.
11. The method according to claim 1, wherein Projecting the filtered connected region onto the target optical flow map and calculating the speed of the moving target within the filtered connected region, includes: Projecting the filtered connected region onto the target optical flow map, and calculating the average optical flow of all points falling within the filtered connected region as the speed of the moving target within the filtered connected region.
12. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and the processor is configured to call and execute the computer program to implement the method according to any one of claims 1-11.
13. A computer-readable storage medium, characterized in that, A computer program is stored on a computer-readable storage medium, and when the computer program is executed by an electronic device, the method according to any one of claims 1-11 is implemented.
Citation Information
Patent Citations
Video moving target detection method
CN107133972A
Movement target detecting method based on light streams and inter-frame matching
CN108154520A
A method for detecting and tracking a moving object
CN109102523A
Crowd abnormity detection method and related device
CN110781853A
Moving target detection method and device, terminal equipment and storage medium
CN112967321A