Moving object detection method and related equipment
By calling the optical flow estimation algorithm in the terminal device and performing multi-stage denoising processing and area screening, the error detection problem caused by optical flow pattern noise is solved, and the accuracy of motion object detection and image quality are improved.
Patent Information
- Application Number
- CN202410033473.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-01-09
AI Technical Summary
In the prior art, there is a large amount of noise in the optical flow map, which leads to a high error detection rate for moving objects detection and a low accuracy in calculation of motion speed, which in turn affects the quality of the captured image.
By calling the optical flow estimation calculation method of the chip in the terminal device, the initial optical flow map is obtained, and through multi-stage denoising processing, molecular area division, and connection area screening, the moving object area is accurately determined, and the accuracy of motion speed calculation is improved.
The noise in the optical flow diagram is effectively removed, and the area of moving objects is accurately determined, which improves the accuracy of motion speed calculation and improves the quality of the captured images.
Smart Images

Figure CN120339337A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminals, and in particular, to a method for detecting a moving object and related devices. Background Art
[0002] In the field of computer vision, the moving object detection technology can be used to identify and track objects in a moving state from a video sequence or an image. When applied to a real-time shooting scenario, through moving object detection, it is possible to sense whether there are moving objects in the captured scene and related information of the moving objects (such as the image area where the moving object is located, the moving speed of the moving object, etc.). Thus, the exposure parameters can be determined based on this related information to adjust the exposure strategy of the shooting device in real time and improve the image clarity of the captured image.
[0003] In the related art, usually, the image processing function of the chip platform of the terminal device itself is used to determine the optical flow map between two adjacent image frames captured by the shooting device, and then determine whether there are moving objects in the current scene and the related information of the moving objects according to the optical flow map. However, due to a large amount of noise in the optical flow map obtained in the related art, problems such as misdetecting the noise as a moving object and a low accuracy rate of the calculated speed of the moving object occur. Furthermore, the exposure parameters of the shooting device are inaccurate, and the quality of the captured image is low. Summary of the Invention
[0004] In view of the above, it is necessary to provide a method for detecting a moving object and related devices, which can solve the problems that a large amount of noise in the optical flow map easily leads to misdetection of moving objects, a low accuracy rate of calculating the moving speed, and a low quality of the captured image caused by the above problems.
[0005] In a first aspect, this application provides a method for detecting a moving object, which is applied to a terminal device. The method includes: in response to a shooting operation of a user, obtaining two adjacent image frames by using a shooting device of the terminal device; calculating an initial optical flow map between the two image frames by using a preset optical flow estimation algorithm; performing denoising processing on the initial optical flow map to obtain a target optical flow map; dividing the target optical flow map into multiple sub-regions, and determining a first speed corresponding to each sub-region according to the optical flow value of each pixel point in each sub-region; determining multiple target regions in the target optical flow map according to the first speed, and determining a second speed corresponding to each target region, where each target region includes at least one sub-region; determining the region where there is a moving object in the first image frame of the two image frames according to the target region, and determining the moving speed of the moving object according to the second speed.
[0006] Through the above technical solution, the optical flow estimation algorithm of the chip of the terminal device can be called when shooting an image to obtain an initial optical flow map between two adjacent image frames; by denoising the initial optical flow map, the noise and abnormal optical flow in the initial optical flow map are gradually removed; by dividing the target optical flow map into multiple sub-regions, the connected regions where moving objects may exist in the target optical flow map can be determined according to the optical flow values in the target optical flow map; by screening the connected regions, the regions where moving objects exist in the target optical flow map can be accurately determined, avoiding misdetecting the static noise regions as moving objects and improving the accuracy of calculating the moving speed of moving objects.
[0007] In a possible implementation manner, each pixel point in the initial optical flow map includes at least a first channel and a second channel, where the first channel includes a first displacement value in a first direction, and the second channel includes a second displacement value in a second direction; the optical flow value of each pixel point in the initial optical flow map includes the first displacement value, the second displacement value, and a third displacement value corresponding to the displacement vector determined according to the first displacement value and the second displacement value.
[0008] Through the above technical solution, the displacement information of each pixel point in two adjacent image frames can be determined according to the values of each channel in the initial optical flow map, where the displacement information may include displacement values in two directions, and the displacement value of the displacement vector determined according to the displacement values in two directions, which can facilitate subsequent processes to determine the moving regions in the optical flow map according to these displacement values.
[0009] In a possible implementation manner, the denoising process of the initial optical flow map to obtain the target optical flow map includes: based on the optical flow value of each pixel point in the initial optical flow map and a preset optical flow value range, performing denoising processing on the initial optical flow map to obtain an updated optical flow map; based on the first image frame and the second image frame in the two image frames, using an improved difference algorithm to obtain a denoised difference map; using the denoised difference map to perform denoising processing on the updated optical flow map to obtain the target optical flow map.
[0010] Through the above technical solution, the noise in the initial optical flow map can be initially removed according to the preset optical flow value range to obtain an updated optical flow map; a denoised difference map of two image frames is obtained according to the improved difference algorithm, and the denoised difference map is used to further denoise the updated optical flow map, which can focus on distinguishing the differences between moving and static pixel points in the updated optical flow map. By resetting the optical flow values in the updated optical flow map, the optical flow noise that interferes with the distinction between moving and static regions can be effectively removed.
[0011] In a possible implementation, the optical flow value of each pixel point in the initial optical flow map includes a third displacement value. The denoising process of the initial optical flow map based on the optical flow value of each pixel point in the initial optical flow map and a preset optical flow value range includes: if the third displacement value of any pixel point in the initial optical flow map is not within the optical flow value range, updating the optical flow value of the any pixel point to a preset first value; or, if the third displacement value of the any pixel point is within the optical flow value range, maintaining the optical flow value of the any pixel point.
[0012] Through the above technical solution, slight noise caused by the shaking of the terminal device when taking pictures and abnormally large noise caused by the optical flow estimation algorithm can be removed according to the preset optical flow value range.
[0013] In a possible implementation, obtaining the denoised difference map by using an improved difference algorithm based on the first image frame and the second image frame in the two image frames includes: determining the first gray value of each pixel point in the first image frame and the second gray value of the corresponding pixel point in the second image frame; based on the first gray value and the second gray value corresponding to each pixel point, performing difference calculation on the first image frame and the second image frame by using the improved difference algorithm to obtain an initial difference image between the first image frame and the second image frame; performing binarization processing on the initial difference image to obtain a binarized image; and performing a post-processing operation on the binarized image to obtain the denoised difference map.
[0014] Through the above technical solution, it is possible to perform difference calculation on two image frames by using an improved difference algorithm according to the gray values of each corresponding pixel point in the two image frames, so as to focus on distinguishing the differences between moving and stationary pixel points in the updated optical flow map and obtain an initial difference image of the two image frames; then through binarization processing, the moving area and the stationary area in the initial difference image can be distinguished; thereafter, through post-processing operations such as image erosion processing and image dilation processing, white noise in the binarized image can be further removed, the edges and contours of the moving area and the stationary area in the binarized image can be enhanced, and the edges of different areas can be made smoother and more continuous.
[0015] In a possible implementation, the formula used in the improved difference algorithm includes:
[0016]
[0017] where D n (x,y) represents the gray value of the pixel point with coordinates (x,y) in the initial difference image, and f n (x,y) represents the first image frame f nThe first grayscale value of the pixel point with coordinates (x, y) in f n-1 (x, y) represents the second image frame f n-1 The second grayscale value of the pixel point with coordinates (x, y) in it.
[0018] Through the above technical solution, compared with the method of subtracting the absolute value of the grayscale values of corresponding pixel points in two image frames in the traditional difference algorithm, the formula used in the above difference algorithm can focus on distinguishing moving and stationary pixel points. Therefore, compared with the difference map with a large amount of noise obtained by the traditional difference algorithm, the initial difference image obtained by using the above difference algorithm can more accurately distinguish moving objects from stationary backgrounds.
[0019] In a possible implementation manner, the binaryzation processing of the initial difference image includes: if the grayscale value of any pixel point in the initial difference image is less than a preset grayscale threshold, updating the grayscale value of the any pixel point to a preset first value; or, if the grayscale value of the any pixel point is greater than or equal to the grayscale threshold, updating the grayscale value of the any pixel point to a preset second value.
[0020] Through the above technical solution, the grayscale values of the pixel points in the initial difference image can be updated to preset values, so that moving pixel points and stationary pixel points can be accurately identified according to the obtained binary image.
[0021] In a possible implementation manner, the denoising processing of the updated optical flow map by using the denoised difference map includes: performing a per-pixel AND operation on the denoised difference map and the updated optical flow map, and updating the optical flow value of each pixel point in the updated optical flow map according to the result of the AND operation.
[0022] In a possible implementation manner, the updating of the optical flow value of each pixel point in the updated optical flow map according to the result of the AND operation includes: if the grayscale value of any pixel point in the denoised difference map is equal to a preset first value, updating the optical flow value of the pixel point corresponding to the any pixel point in the updated optical flow map to the first value; or, if the grayscale value of the any pixel point is equal to a preset second value, maintaining the optical flow value of the pixel point corresponding to the any pixel point in the updated optical flow map.
[0023] Through the above technical solution, the corresponding optical flow value in the updated optical flow map can be reset to 0 according to the stationary pixel points indicated by the denoised difference map, and the optical flow value of the corresponding pixel point in the updated optical flow map can be maintained unchanged according to the moving pixel points indicated by the denoised difference map, so as to further eliminate the noise caused by abnormal optical flow values in the updated optical flow map.
[0024] In a possible implementation, the optical flow value of each pixel point in each sub-region includes a first displacement value and a second displacement value. The method of dividing the target optical flow map into multiple sub-regions and determining the first velocity corresponding to each sub-region according to the optical flow value of each pixel point in each sub-region includes: determining a first average value of the first displacement values of all pixel points in each sub-region, and a second average value of the second displacement values of all pixel points; determining the first velocity based on the first average value and the second average value.
[0025] Through the above technical solution, the first velocity corresponding to each sub-region can be preliminarily determined according to the average value of the first displacement values in the first direction of all pixel points and the average value of the second displacement values in the second direction in each sub-region, so as to facilitate screening out the region where the moving object is located from the target optical flow map in the subsequent process.
[0026] In a possible implementation, the method of determining multiple target regions in the target optical flow map according to the first velocity includes: if the first velocity of any sub-region is greater than a preset velocity threshold, taking the any sub-region as a target sub-region; determining the connected region where each target sub-region is located based on the connected graph algorithm, and taking the connected region as the target region, where each connected region includes at least one target sub-region.
[0027] Through the above technical solution, the target sub-regions exceeding the preset velocity threshold can be determined according to the first velocity of each sub-region; by connecting the target sub-regions, the connected region where there is a moving object can be determined as the target region.
[0028] In a possible implementation, before determining the connected region to which each target sub-region belongs based on the connected graph algorithm, the method further includes noise filtering of the target optical flow map, including: updating the optical flow value of the isolated target sub-region at the non-boundary position in the target optical flow map to a preset first value, where the isolated target sub-region includes a target sub-region that is not adjacent to other target sub-regions.
[0029] Through the above technical solution, the isolated target sub-region at the non-boundary in the target optical flow map can be updated to a non-moving region, so as to further denoise the target optical flow map.
[0030] In a possible implementation, the method further includes: cleaning the outliers in the first velocities of all sub-regions in each target region, and determining the third velocity of each target region according to the average value of the first velocities after cleaning; based on the sliding window algorithm, determining multiple sliding window average values corresponding to the first velocities of all sub-regions in each target region, and determining the fourth velocity of each target region according to the maximum value among the multiple sliding window average values.
[0031] Through the above technical solution, it is possible to remove the smaller and larger values in the first speed based on the principle of normal distribution, and obtain the average value of the first speed after removing the larger and smaller values as the third speed of the target area; it is possible to obtain the average value corresponding to the larger value in the first speed as the fourth speed of the target area through the sliding window algorithm.
[0032] In a possible implementation manner, the method further includes: determining the shooting state of the shooting device according to the target optical flow map, including: determining the ratio of the number of all target sub-regions in the target optical flow map to the total number of all sub-regions in the target optical flow map, where the target sub-region represents a sub-region where the first speed is greater than a preset speed threshold; if the ratio is greater than a preset ratio threshold, determining that the shooting state of the shooting device is a moving state; or, if the ratio is less than or equal to the ratio threshold, determining that the shooting state of the shooting device is an approximately stationary state.
[0033] Through the above technical solution, it is possible to determine the shooting state of the shooting device for two image frames corresponding to the target optical flow map according to the proportion of the moving sub-regions in the target optical flow map. For example, if the proportion of the moving sub-regions in the target optical flow map is large, it can be considered that there are many moving sub-regions in the target optical flow map because the shooting device is moving, so it can be considered that the shooting device is in a moving state; or, if the proportion of the moving sub-regions in the target optical flow map is small, it can be considered that the few moving sub-regions in the target optical flow map are the regions corresponding to the moving objects in the shooting scene, and the shooting device is approximately stationary, so it can be considered that the shooting device is in an approximately stationary state.
[0034] In a possible implementation manner, the method further includes: determining the second speed based on the shooting state of the shooting device, including: if the shooting state is the approximately stationary state and the fourth speed of each target area is less than the product of a preset third value and the third speed of each target area, determining the second speed according to the fourth speed; or, if the shooting state is the approximately stationary state and the fourth speed of each target area is greater than or equal to the product of the third value and the third speed of each target area, determining the second speed according to the third speed; or, if the shooting state is the moving state and the fourth speed of each target area is less than the product of a preset fourth value and the third speed of each target area, determining the second speed according to the fourth speed; or, if the shooting state is the moving state and the fourth speed of each target area is greater than or equal to the product of the fourth value and the third speed of each target area, determining the second speed according to the third speed.
[0035] Through the above technical solution, the second speed of the target area can be determined according to the shooting state of the shooting device, the third speed and the fourth speed of the target area, so as to avoid misjudgment of the motion speed of the target area caused by the shooting state of the shooting device and improve the calculation accuracy of the motion speed of the target area.
[0036] In a possible implementation manner, the method further includes: performing denoising processing on the target area, including: sorting all target areas in descending order of the area of the target area, selecting a preset number of target areas at the front from the sorted sequence, and determining noise areas from the selected target areas.
[0037] In a possible implementation manner, determining the noise area in the target area includes: determining the corresponding area of each target area in the first image frame or the second image frame, obtaining a traditional difference map between the two image frames of the corresponding area, and each pixel point in the traditional difference map corresponds to a difference value; determining the average difference value of the difference values of all pixel points in the traditional difference map corresponding to each target area; determining a difference threshold corresponding to each target area according to the number of sub-areas in each target area; and determining the target area corresponding to the average difference value less than or equal to the corresponding difference threshold as the noise area.
[0038] In a possible implementation manner, the difference threshold corresponding to each target area is determined according to the following formula:
[0039]
[0040] where threshold represents the difference threshold, num represents the number of sub-areas in each target area, and α, λ, β, γ, k represent preset parameters.
[0041] Through the above technical solution, the corresponding difference threshold can be determined according to the number of sub-areas of each target area, so that the noise area in the target area can be determined according to the average difference value and the difference threshold of the target difference area corresponding to each target area in the traditional difference maps of the two image frames, and the recognition accuracy of the motion area is improved.
[0042] In a second aspect, the present application provides a terminal device, where the terminal device includes a memory and a processor: wherein, the memory is used to store program instructions; the processor is used to read and execute the program instructions stored in the memory, and when the program instructions are executed by the processor, the terminal device executes the above-mentioned moving object detection method.
[0043] In a third aspect, the present application provides a chip, which is coupled to a memory in a terminal device, and the chip is configured to control a processor in the terminal device to execute the above-mentioned moving object detection method.
[0044] In a fourth aspect, the present application provides a computer storage medium storing program instructions, which when running on a terminal device, cause the processor of the terminal device to execute the above-mentioned moving object detection method.
[0045] In addition, for the technical effects brought by the second aspect to the fourth aspect, reference may be made to the descriptions related to the methods in the above method section for each design, which will not be elaborated here. Description of the Drawings
[0046] Figure 1 It is a schematic diagram of a visualized image of optical flow provided by an embodiment of the present application.
[0047] Figure 2 It is a schematic diagram of an initial optical flow map provided by an embodiment of the present application.
[0048] Figure 3 It is a software architecture diagram of a terminal device provided by an embodiment of the present application.
[0049] Figure 4 It is a flowchart of a moving object detection method provided by an embodiment of the present application.
[0050] Figure 5 It is an example diagram of denoising processing of an optical flow map provided by an embodiment of the present application.
[0051] Figure 6 It is an example diagram of a traditional difference map provided by an embodiment of the present application.
[0052] Figure 7 It is an example diagram of a denoised difference map provided by an embodiment of the present application.
[0053] Figure 8 It is a flowchart of a method for denoising an initial optical flow map provided by an embodiment of the present application.
[0054] Figure 9 It is a flowchart of a refined process of S302 provided by an embodiment of the present application.
[0055] Figure 10 It is a schematic diagram of dividing a target optical flow map into multiple sub-regions provided by an embodiment of the present application.
[0056] Figure 11 It is a schematic diagram of a motion state matrix corresponding to a target optical flow map provided by an embodiment of the present application.
[0057] Figure 12It is a schematic diagram for noise filtering of a target optical flow map provided by an embodiment of the present application.
[0058] Figure 13 It is a schematic diagram of a connected region provided by an embodiment of the present application.
[0059] Figure 14 It is a flowchart of a method for determining the shooting state of a shooting device provided by an embodiment of the present application.
[0060] Figure 15 It is a flowchart of a method for determining the third speed and the fourth speed of a target region provided by an embodiment of the present application.
[0061] Figure 16 It is a flowchart of a method for determining a noise region in a target region provided by an embodiment of the present application.
[0062] Figure 17 It is an example diagram of pixel points in a target region provided by an embodiment of the present application.
[0063] Figure 18 It is a flowchart of a method for differential denoising of optical flow provided by an embodiment of the present application.
[0064] Figure 19 It is a flowchart of a motion detection method provided by an embodiment of the present application.
[0065] Figure 20 It is a flowchart of a method for noise verification of a connected graph provided by an embodiment of the present application.
[0066] Figure 21 It is a hardware architecture diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0067] In an embodiment of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, words such as "exemplary" or "for example" are used to mean as an example, illustration or explanation. Any embodiment or design solution described as "exemplary" or "for example" in an embodiment of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application pertains. The terms used in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application. It should be understood that, unless otherwise stated in this application, " / " means "or". For example, A / B may mean A or B. The "and / or" in this application is merely a description of the relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone, these three situations. "At least one" means one or more. "Multiple" means two or more than two. For example, at least one of a, b, or c may mean: a, b, c, a and b, a and c, b and c, a, b, and c, these seven situations. Without conflict, the following embodiments and the features in the embodiments may be combined with each other.
[0069] In the field of computer vision, motion object detection technology can be used to identify and track objects in a moving state from a video sequence or an image. When applied to a real-time shooting scenario, through motion object detection, it is possible to sense whether there are moving objects in the captured scene and the relevant information of the moving objects (such as the image area where the moving object is located, the moving speed of the moving object, etc.). Thus, the exposure reduction parameters can be determined based on this relevant information to adjust the exposure strategy of the shooting device in real time and improve the image clarity of the captured image.
[0070] In related technologies, an optical flow estimation algorithm that usually uses the image processing function of the chip platform of the terminal device itself is used to determine the optical flow between two adjacent (indicating consecutive in the time series) image frames captured by the shooting device. Thus, based on the optical flow, it is determined whether there are moving objects (such as vehicles, people, animals, etc.) in the current scene and the relevant information of the moving objects.
[0071] Specifically, reference can be made to Figure 1 the schematic diagram of the visualization image of the optical flow (hereinafter referred to as the optical flow map) shown. The optical flow represents the movement of target pixel points in the image due to the movement of moving objects in the image and / or the movement of the shooting device between two adjacent image frames. The optical flow can be used to capture the displacement information of corresponding pixel points between two adjacent image frames. For example, the moving speed and moving direction of each pixel point, etc. Among them, the time interval between two adjacent image frames is very small, such as 1 / 60 second. Therefore, the optical flow estimation algorithm can calculate the optical flow between two adjacent image frames based on the following physical assumptions: the pixel intensity of the shooting scene is basically unchanged between two adjacent image frames, and adjacent pixels have similar movements.
[0072] The optical flow maps of two adjacent image frames obtained by the chip platform of the terminal device itself (hereinafter referred to as the initial optical flow maps) can be images of size width×height×channel, where width represents the image width of the initial optical flow map (equal to the image width of each image frame), height represents the image height of the initial optical flow map (equal to the image height of each image frame), and channel represents the number of channels included in each pixel point in the initial optical flow map. Among them, the unit length of the size of the image can be the size of one pixel point.
[0073] For example, the number of channels can be at least 2, indicating that each pixel point in the initial optical flow map includes at least 2 channels, and one of the channels can be, for example, Figure 1 the displacement vector in the horizontal X-axis direction as in Another channel can be, for example, Figure 1 the displacement vector in the vertical Y-axis direction as in The displacement vector and the displacement vector can be used to determine the displacement vector Figure 1 in, for example, by the sum value. The displacement value can be equal to the square root of the sum of the squares of the displacement value and the displacement value The larger the displacement value, the greater the movement speed of the object at the corresponding pixel point. Among them, the unit length of the displacement value can be the size of one pixel point.
[0074] For ease of understanding, the displacement vector can be understood as: the vector obtained by the position change of a certain feature point from the position of pixel point A in image frame f n-1 to the position of pixel point B in image frame f n where image frame f n represents the nth image frame (hereinafter referred to as the first image frame), and image frame f n-1 represents the (n - 1)th image frame adjacent to the first image frame f n (hereinafter referred to as the second image frame). The time node corresponding to the second image frame is before the time node corresponding to the first image frame. For example, according to the time sequence, the second image frame is obtained first, and then the first image frame is obtained. Among them, n can represent an integer greater than or equal to 2, and the sizes of the second image frame and the first image frame are the same.
[0075] In practical applications, due to a large amount of noise in the optical flow maps obtained in the related art, problems such as misdetecting the noise area as a moving object and a low accuracy rate of the calculated speed of the moving object occur, which in turn leads to inaccurate exposure parameters of the shooting device and low image quality of the captured images (for example, there are obvious blurred areas).
[0076] Refer to Figure 2 As shown, it is a schematic diagram of an initial optical flow map provided by an embodiment of the present application. Among them, as Figure 2 In the initial optical flow map shown in, the white area represents the area where the optical flow value is 0, and the depth of the gray scale of the pixel points is used to indicate the magnitude of the optical flow value at the corresponding position. The darker the gray scale, the larger the optical flow value. Except for the darker area in the lower right corner in the initial optical flow map, the other areas have lighter gray, and it can be seen that there are more small optical flow value noises in the lighter gray area of the initial optical flow map; and there are also areas with different shades of color in the darker area in the lower right corner of the initial optical flow map, and the darkest black dot area among them represents the large optical flow value noise area.
[0077] To solve the above problems, an embodiment of the present application provides a moving object detection method, which can call the optical flow estimation algorithm of the chip of the terminal device when shooting an image to obtain the initial optical flow map between two adjacent image frames; by denoising the initial optical flow map, gradually removing the noise and abnormal optical flow in the initial optical flow map; by dividing the target optical flow map into multiple sub-regions, the connected regions where moving objects may exist in the target optical flow map can be determined according to the optical flow values in the target optical flow map; by screening the connected regions, the regions where moving objects exist in the target optical flow map can be accurately determined, avoiding misdetecting the noise regions in the stationary state as moving objects, and improving the accuracy of calculating the moving speed of moving objects. The moving object detection method is applied to various terminal devices, such as mobile phones, tablet computers, wearable devices, camera devices, computers and other devices. The terminal device includes an application processor, and the application processor is used to run the operating system. The following combines Figure 3 Exemplarily illustrate the software structure of the terminal device.
[0078] Refer to Figure 3 As shown, it is the software architecture diagram of the terminal device provided by the embodiment of the present application. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. For example, the Android system from top to bottom is the application layer 101, the framework layer 102, the Android runtime and system libraries 103, the hardware abstraction layer 104, the kernel layer 105, and the hardware layer 106.
[0079] The application layer 101 may include a series of application packages. For example, the application packages may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, device control service, etc.
[0080] The framework layer 102 provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. For example, the application framework layer may include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc.
[0081] Among them, the window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc. The content provider is used to store and obtain data, and make this data accessible to applications. The data may include videos, images, audio, dialed and answered calls, browsing history and bookmarks, phone books, etc. The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a short message notification icon may include a view for displaying text and a view for displaying pictures. The telephone manager is used to provide the communication functions of the terminal device. For example, the management of call states (including connection, hanging up, etc.). The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc. The notification manager enables applications to display notification information in the status bar, can be used to convey notification-type messages, and can disappear automatically after a short stay without user interaction. For example, the notification manager is used to inform that the download is completed, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or a scroll bar text, such as a notification of a background running application, and can also be a notification that appears on the screen in the form of a dialog window. For example, prompt text information in the status bar, emit a prompt sound, the terminal device vibrates, the indicator light flashes, etc.
[0082] Android Runtime includes a core library and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system. The core library contains two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of Android.
[0083] The application layer 101 and the framework layer 102 run in the virtual machine. The virtual machine executes the Java files of the application layer and the framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.
[0084] The system library 103 may include multiple functional modules. For example, a surface manager, Media Libraries, a 3D graphics processing library (e.g., OpenGL ES), a 2D graphics engine (e.g., SGL), etc.
[0085] Among them, the surface manager is used to manage the display subsystem and provide the fusion of 2D and 3D layers for multiple applications. The Media Libraries support the playback and recording of multiple common audio and video formats, as well as static image files, etc. The Media Libraries can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc. The 2D graphics engine is a drawing engine for 2D drawing.
[0086] The hardware abstraction layer 104 runs in the user space, encapsulates the kernel layer drivers, and provides call interfaces to the upper layer.
[0087] The kernel layer 105 is the layer between the hardware and the software. The kernel layer 105 at least includes a display driver, a touch driver, an audio driver, and a sensor driver.
[0088] The kernel layer 105 is the core of the operating system of the terminal device, is the first layer of software extension based on the hardware, provides the most basic functions of the operating system, is the basis for the operation of the operating system, and is responsible for managing the system's processes, memory, device drivers, files, and network systems, and determines the performance and stability of the system. For example, the kernel layer can determine the operation time of an application program for a certain part of the hardware.
[0089] The kernel layer 105 includes programs closely related to the hardware, such as interrupt handlers, device drivers, etc., and also includes basic, common, and frequently running modules, such as a clock management module, a process scheduling module, etc., and also includes key data structures. The kernel layer can be set in the processor or solidified in the internal memory.
[0090] The hardware layer 106 includes the hardware of the terminal device, such as a display screen, buttons, a camera, etc.
[0091] Refer to Figure 4 As shown, it is a flowchart of a moving object detection method provided by an embodiment of the present application. The moving object detection method is applied to a terminal device and includes the following processes.
[0092] S201, in response to a user's shooting operation, use the shooting device of the terminal device to obtain two adjacent image frames.
[0093] In an embodiment of the present application, the user's shooting operation may be an operation to open the camera application, or an operation to click the shooting control after the camera application is opened. The terminal device may obtain a series of image frames in response to the user's shooting operation, and select two adjacent image frames from the obtained image frames. The two adjacent image frames may be two adjacent preview image frames. For example, the user opens the camera application and aims the camera at a certain object to take a photo or record a video, and the terminal device responds to the user operation to obtain two adjacent image frames of the object. In another embodiment, the user's shooting operation may also be, when using an application (APP), to call the camera application through the photo-taking function provided by the application to obtain two adjacent image frames. For example, call the camera application through an instant messaging application to photograph any object.
[0094] In an embodiment of the present application, the two adjacent image frames may include: the nth image frame f n (hereinafter referred to as the first image frame), the first image frame f n the adjacent (n - 1)th image frame f n-1 (hereinafter referred to as the second image frame), where n represents an integer greater than or equal to 2.
[0095] The first time node corresponding to the first image frame is after the second time node corresponding to the second image frame. For example, the first time node is 1 / 60 second later than the second time node, and the second image frame has the same size as the first image frame. The first image frame and the second image frame may be multi-channel images. For example, each pixel point includes an RGB image with three color channels of red (R), green (G), and blue (B). The first image frame and the second image frame may also be grayscale images or images in other formats, and the present application does not make specific limitations on this.
[0096] S202, use a preset optical flow estimation algorithm to calculate the initial optical flow map between the two image frames.
[0097] In an embodiment of the present application, the optical flow estimation algorithm of the chip platform of the terminal device may be called to calculate the optical flow map between the two image frames as the initial optical flow map. For the introduction of the initial optical flow map, reference may be made to the above description of Figure 1 and Figure 2 .
[0098] In an embodiment of the present application, each pixel point in the initial optical flow map includes at least a first channel and a second channel. Among them, the first channel includes a first displacement value in a first direction (for example Figure 1 the displacement value in the horizontal X-axis direction in ), and the second channel includes a second displacement value in a second direction (for example Figure 1 the displacement value in the vertical Y-axis direction in )。The optical flow value of each pixel point in the initial optical flow map includes a first displacement value, a second displacement value, and a displacement vector determined according to the first displacement value and the second displacement value (for example Figure 1 the displacement vector in ) corresponding third displacement value (for example Figure 1 the displacement value in ). The optical flow value (such as the third displacement value) in the initial optical flow map represents a value greater than or equal to 0. When the optical flow value is 0, it indicates that the corresponding pixel point is in a stationary state, and the larger the optical flow value, the greater the movement speed of the corresponding pixel point.
[0099] S203. Denoise the initial optical flow map to obtain a target optical flow map.
[0100] In an embodiment of the present application, as Figure 2 shown, due to reasons such as the shaking of the shooting device, the limitations of the mathematical model and physical assumptions of the optical flow estimation algorithm (for example, the pixel intensity of the shooting scene is basically unchanged between two adjacent image frames, and adjacent pixels have similar motions), and the complex and changeable shooting scene, the initial optical flow map may contain a large amount of noise, and it is necessary to denoise the initial optical flow map to improve the accuracy of moving object detection in subsequent processes.
[0101] In an embodiment of the present application, for different noises in the initial optical flow map caused by different reasons, different denoising methods can be used to perform multiple denoising processes on the initial optical flow map in multiple stages, so as to specifically remove different noises in the initial optical flow map in each stage and improve the denoising effect on the initial optical flow map.
[0102] In an embodiment of the present application, the denoising process (hereinafter referred to as the first denoising process) for the abnormal small optical flow value noise (such as the small optical flow value noise caused by the small shaking of the shooting device) and the abnormal large optical flow value noise (such as the large optical flow value noise caused by the optical flow estimation algorithm) in the initial optical flow map may include: by setting an optical flow value range, reset the optical flow values of the pixel points in the initial optical flow map whose corresponding optical flow values do not belong to this optical flow value range, for example, reset them to 0, so as to remove the influence of the above small optical flow value noise and large optical flow value noise and obtain an updated optical flow map.
[0103] Refer to Figure 5 shown, which is an example diagram of the optical flow map denoising process provided by an embodiment of the present application. Figure 5 Each picture in shows an optical flow map. The depth of the gray scale of the pixel points in the optical flow map is used to indicate the size of the optical flow value at the corresponding position. The darker the gray scale, the larger the optical flow value, and white indicates that the optical flow value is 0. Figure 5 In the initial optical flow map in, there is more small optical flow value noise in the lighter gray area, and large optical flow value noise in the black dot area.Figure 5 The updated optical flow map represents the optical flow map obtained after the first denoising process on the initial optical flow map. It can be seen that after the first denoising process, the light gray small optical flow value noise regions in the initial optical flow map are processed into white, achieving the removal of small optical flow value noise; in addition, the black dot large optical flow value noise regions in the initial optical flow map are processed into white, achieving the removal of large optical flow value noise.
[0104] In an embodiment of the present application, the further denoising process (hereinafter referred to as the second denoising process) of the updated optical flow map can be achieved through the difference map between the first image frame and the second image frame. The principle includes: the optical flow vectors in the updated optical flow map represent the displacement situation of each pixel point between the two image frames, and the noise therein is usually random and has little impact on the pixel distribution of the entire image; while the difference image between the two image frames can highlight the dynamic changes of the pixel points between the two image frames, and the noise therein usually does not cause large changes. Therefore, the dynamic change information in the difference image between the two image frames can be used to filter the optical flow noise in the updated optical flow map, achieving the further denoising of the updated optical flow map, so as to more accurately distinguish the foreground moving objects and the background static regions in the optical flow map.
[0105] In an embodiment of the present application, in the related art, the traditional difference method is usually used to obtain the traditional difference map between the two image frames. Among them, the formula used in the traditional difference method includes: D n (x,y) = |f n (x,y) - f n-1 (x,y)|, where f n represents the nth image frame (for example, the first image frame), and f n (x,y) represents the first gray value of the pixel point with coordinates (x,y) in f n ; f n-1 represents the (n - 1)th image frame adjacent to the nth image frame (for example, the second image frame), and f n-1 (x,y) represents the second gray value of the pixel point with coordinates (x,y) in f n-1 . Among them, as Figure 1 shown, taking the upper left corner point of the image frame (for example, the first image frame) as the coordinate origin O to establish a rectangular coordinate system OXY, with the horizontal direction as the direction of the X axis and the vertical direction as the direction of the Y axis, so as to determine the coordinates of each pixel point in each image frame.
[0106] In an embodiment of the present application, a traditional difference map usually contains a large amount of noise, resulting in problems such as being unable to obtain the complete edges of objects in the image based on the traditional difference map and being unable to distinguish foreground moving objects from background static regions in the image. Even after performing binarization processing and conventional denoising on the traditional difference map, due to too much noise in the traditional difference map, there will still be a lot of noise in the denoised traditional difference map obtained (for example Figure 6 as shown), and it cannot be directly used for subsequent denoising of the updated optical flow map.
[0107] The embodiment of the present application provides an improved difference algorithm, which can use the following formula:
[0108]
[0109] where D n (x, y) represents the gray value of the pixel point with coordinates (x, y) in the initial difference image, and f n (x, y) represents the first gray value of the pixel point with coordinates (x, y) in the first image frame f n , and f n-1 (x, y) represents the second gray value of the pixel point with coordinates (x, y) in the second image frame f n-1 .
[0110] In the traditional difference algorithm, the absolute value of the difference between the gray values of corresponding pixel points in two image frames is taken. Compared with the traditional difference algorithm, the formula used in the improved difference algorithm provided by the embodiment of the present application can focus on distinguishing moving and static pixel points. Therefore, compared with the traditional difference algorithm that will obtain a difference map with a large amount of noise, the initial difference image obtained by using the difference algorithm provided by the embodiment of the present application has very little noise and can more accurately distinguish moving objects from static backgrounds (for example Figure 7 as shown).
[0111] In an embodiment of the present application, further post-processing operations can be performed on the initial difference map, and the post-processed initial difference map is used as the denoised difference map for the second denoising process of the updated optical flow map. For example, the post-processing operations can include, but are not limited to, a combination of one or more of the following methods: binarization processing, image erosion processing, and image dilation processing, etc., so as to achieve the following effects: further distinguish the moving region and the static region in the initial difference image, remove the white noise in the binarized image, enhance the edges and contours of the moving region and the static region in the binarized image, and make the edges of different regions smoother and more continuous.
[0112] In an embodiment of the present application, when denoising the updated optical flow map using the denoised difference map, a per-pixel AND operation can be performed on the denoised difference map and the updated optical flow map, and the optical flow value of each pixel in the updated optical flow map can be updated according to the result of the AND operation. Specifically, the optical flow value of the pixel corresponding to the pixel indicating the motion state in the denoised difference map in the updated optical flow map can be maintained unchanged, and the optical flow value of the pixel corresponding to the pixel indicating the stationary state in the denoised difference map in the updated optical flow map can be updated to 0, thereby further removing the noise in the updated optical flow map.
[0113] For example Figure 5 as shown Figure 5 the target optical flow map in shows the optical flow map obtained after performing a second denoising process on the updated optical flow map. It can be seen that after the second denoising process, the noise regions in the updated optical flow map are further removed, and the noise influence is basically removed in the obtained target optical flow map.
[0114] In an embodiment of the present application, the method for denoising the initial optical flow map can further refer to the description of the embodiment Figure 8 as shown.
[0115] S204, divide the target optical flow map into multiple sub-regions, and determine the first velocity corresponding to each sub-region according to the optical flow value of each pixel in each sub-region.
[0116] In an embodiment of the present application, the target optical flow map can be divided into M×N sub-regions, and each sub-region includes m×m pixels, where M, N, and m can be positive integers greater than 1 according to actual needs. For example, if the target optical flow map contains 640×480 pixel points, the value of M can be 128, the value of N can be 96, and the value of m can be 5. For example Figure 10 as shown, is a schematic diagram of dividing the target optical flow map into multiple sub-regions provided by an embodiment of the present application, where each grid in the right image represents a sub-region.
[0117] In an embodiment of the present application, the optical flow value of each pixel in each sub-region includes a first displacement value and a second displacement value. Based on the first displacement values of all pixels in each sub-region, a first average value can be determined; based on the second displacement values of all pixels, a second average value can be determined; based on the first average value and the second average value, a first velocity is determined. Exemplarily, the following formula can be used:
[0118]
[0119] where, V velicity represents the first velocity of any sub-region, V XThe first average value V of the first displacement values of all pixel points within the sub-region Y The second average value of the second displacement values of all pixel points within the sub-region.
[0120] Based on the above embodiments, the first speed corresponding to each sub-region can be preliminarily determined according to the average value of the first displacement values in the first direction and the average value of the second displacement values in the second direction of all pixel points in each sub-region, so as to facilitate screening out the regions where the moving objects are located from the target optical flow map in the subsequent process.
[0121] S205. Determine multiple target regions in the target optical flow map according to the first speed, and determine the second speed corresponding to each target region.
[0122] In an embodiment of the present application, each target region includes at least one sub-region. Determining multiple target regions in the target optical flow map according to the first speed may include the following processes (1)-(2):
[0123] (1) If the first speed of any sub-region is greater than a preset speed threshold, use the any sub-region as a target sub-region.
[0124] In an embodiment of the present application, the preset speed threshold can be set according to actual needs. For example, the speed threshold can be 1, and the present application does not make specific limitations thereto. In an embodiment of the present application, if the first speed of any sub-region is greater than the speed threshold, the sub-region can be used as a target sub-region in the motion state. Refer to Figure 11 As shown, the target optical flow map can be abstracted into a motion state matrix containing only two elements, "0" and "1", to determine whether there is a motion state within each sub-region. Among them, each element in the motion state matrix represents the motion state of a sub-region at the corresponding position. As Figure 11 shown, the grid corresponding to the element "1" in the motion state matrix represents the target sub-region in the motion state, and the grid corresponding to the element "0" in the motion state matrix represents the non-target sub-region, that is, the sub-region in the stationary state.
[0125] Through the above embodiments, it is possible to preliminarily screen the sub-regions in the motion state in the target optical flow map and determine the positions of the target sub-regions in the motion state in the target optical flow map. In addition, by abstracting each sub-region in the target optical flow map into an element in a motion state matrix, data dimensionality reduction can be achieved, improving the computational efficiency of the algorithm.
[0126] (2) Based on the connected graph algorithm, determine the connected region where each target sub-region is located, and use the connected region as the target region.
[0127] In an embodiment of the present application, before determining the connected regions to which each target sub-region belongs based on the connected graph algorithm, noise filtering can also be performed on the target optical flow map first, including: updating the optical flow values (such as the third optical flow value) of the isolated target sub-regions at non-boundary positions in the target optical flow map to a preset first value (such as 0). The isolated target sub-regions include target sub-regions that are not adjacent to other target sub-regions.
[0128] For example Figure 12 As shown, it is a schematic diagram of noise filtering of the target optical flow map provided by an embodiment of the present application. Among them, Figure 12 The elements around the element "1" within the rectangular frame in the left motion state matrix are all "0". It can be determined that the element "1" within the rectangular frame is isolated data. By updating the element "1" within the rectangular frame to the element "0", the filtering of the isolated data can be achieved. Correspondingly, the target sub-region corresponding to the element "1" within the rectangular frame is an isolated target sub-region at a non-boundary position. The third optical flow value of the non-boundary isolated target sub-region can be updated to 0, so as to update the non-boundary isolated target sub-region to a non-moving region and further denoise the target optical flow map.
[0129] In an embodiment of the present application, the connected graph algorithm is an algorithm for detecting and traversing connected regions, which can include but is not limited to depth-first search algorithm, breadth-first search algorithm, etc. Taking the depth-first search algorithm as an example, determining the connected regions where each target sub-region is located can include: regarding each sub-region in the target optical flow map as a node, where each target sub-region represents a connectable node, and each non-target sub-region represents a non-connectable node; starting from a starting node in the target optical flow map, recursively visiting its adjacent nodes. If its adjacent node is a connectable node, it is determined that the node is connected to the adjacent node, and the visited node is marked; by continuously traversing the unvisited adjacent nodes until reaching a non-connectable node and no longer being able to expand, a connected region can be obtained.
[0130] In an embodiment of the present application, each connected region includes at least one target sub-region, representing a larger moving region; each target optical flow map can include at least one connected region. For example Figure 13 As shown, it is a schematic diagram of the connected region provided by an embodiment of the present application. Among them, by Figure 13 connecting the data points where the elements in the left motion state matrix are "1" and adjacent, a connected region can be obtained. As shown in Figure 13 the right motion state matrix, a unique label (such as numbers, characters, English, etc.) is assigned to each connected region to distinguish different connected regions. For example Figure 13 the natural numbers 2 to 7 in the right motion state matrix in represent different connected regions respectively.
[0131] Through the above embodiments, by connecting the target sub-regions, the connected region where a moving object may exist can be determined as the target region, thereby realizing the division of the regions where different moving objects are located.
[0132] In an embodiment of the present application, since the first speed of at least one sub-region included in each target region is known, in the related art, the average value of the first speeds of all sub-regions included in each target region is usually used as the speed of the corresponding target region. However, the accuracy of this method is relatively low. The reason is that the related art does not consider the possible outliers in the first speed due to reasons such as noise, nor does it consider the influence of the state of the shooting device on the first speed.
[0133] Among them, the state of the shooting device may include an approximately stationary state (for example, the shooting device is held by the user or in the tripod mode) and a moving state (for example, the shooting device is in a moving state). The shooting state of the shooting device corresponding to the target optical flow map can be determined according to the proportion of the moving sub-regions in the target optical flow map. For example, if the proportion of the moving sub-regions in the target optical flow map is relatively large, it can be considered that the relatively large number of moving sub-regions in the target optical flow map is caused by the movement or shaking of the shooting device, so it can be considered that the shooting device is in a moving state; or, if the proportion of the moving sub-regions in the target optical flow map is relatively small, it can be considered that the relatively small number of moving sub-regions in the target optical flow map is the region corresponding to the moving object in the shooting scene, and the shooting device is approximately stationary, so it can be considered that the shooting device is in an approximately stationary state. The method for determining the shooting state of the shooting device can also refer to the description of the embodiments shown below. Figure 14 shown in the embodiments.
[0134] According to the above description, it can be known that the moving state of the shooting device will seriously affect the numerical distribution of the first speeds of all sub-regions included in each target region. For example, a shooting device in a moving state will result in very few non-zero values in the first speed. In addition, the moving speed of a shooting device in a moving state may cause many abnormally large values in the first speed. To solve the above problems, two different methods can be used to determine the speed of the corresponding target region according to the first speeds of all sub-regions included in each target region.
[0135] Among them, the first method may include: cleaning outliers that may exist in the first speeds of all sub-regions of each target region. For example, cleaning the smaller and larger values in the first speeds of all sub-regions of each target region, and then determining the average value of the first speeds of all sub-regions of each target region after cleaning as the speed of the corresponding target region (hereinafter referred to as the third speed). This method can regard the smaller and larger values in the first speed as values that may be abnormal and remove the outliers, and obtain the average value of the first speed after filtering out the possible outliers as the third speed of the target region, thereby avoiding the influence of the possible outliers in the first speed on the calculation accuracy of the speed of the target region.
[0136] The second method may include: sorting the first speeds of all sub-regions of each target region in descending order, selecting multiple first speeds at the front from the sorted sequence, and determining the average value of the selected multiple first speeds as the speed of the corresponding target region (hereinafter referred to as the fourth speed). The fourth speed of each target region is generally greater than or equal to the third speed. This method can determine the average value corresponding to the larger value in the first speed as the fourth speed of the target region for the shooting device in the motion state that may cause more abnormally large values in the first speed, and avoid the influence of the state of the shooting device on the calculation accuracy of the speed of the target region by judging whether the fourth speed is the average value of the abnormally large values caused by the shooting device in the subsequent process.
[0137] In an embodiment of the present application, the method for determining the third speed and the fourth speed of the target region may also refer to the description of the embodiment Figure 15 shown.
[0138] In an embodiment of the present application, determining the second speed based on the shooting state of the shooting device may include: if the shooting state is an approximate stationary state and the fourth speed of each target region is less than the product of a preset third value and the third speed of each target region, determining the second speed according to the fourth speed; or, if the shooting state is an approximate stationary state and the fourth speed of each target region is greater than or equal to the product of the third value and the third speed of each target region, determining the second speed according to the third speed. Among them, the third value can be set according to prior data. For example, the third value can be 5, and the present application does not make specific limitations on this.
[0139] Specifically, if the shooting state is an approximate stationary state, it can be determined that there may not be many abnormally large values in the first speed caused by the shooting device in the motion state, that is, the fourth speed may not be the average value of the abnormally large values. On this basis, if the fourth speed is less than the product of the third value and the third speed of each target area, it can be considered that there may be no abnormal values among the larger values removed in the above first method. Therefore, the accuracy of using the third speed obtained by the first method as the speed of the target area is relatively low, and the fourth speed can be used as the speed of the target area (such as the second speed); or, if the fourth speed is greater than or equal to the product of the third value and the third speed of each target area, it can be considered that there may indeed be abnormal values among the larger values removed in the above first method, and the accuracy of using the third speed as the speed of the target area is relatively high. Therefore, the third speed can be used as the speed of the target area (such as the second speed).
[0140] In an embodiment of the present application, determining the second speed based on the shooting state of the shooting device may further include: if the shooting state is a motion state and the fourth speed of each target area is less than the product of a preset fourth value and the third speed of each target area, determining the second speed according to the fourth speed; or, if the shooting state is a motion state and the fourth speed of each target area is greater than or equal to the product of the fourth value and the third speed of each target area, determining the second speed according to the third speed. Among them, the fourth value can be set according to prior data. For example, the fourth value can be 1.2, and the present application does not make specific limitations on this.
[0141] Specifically, if the shooting state is a motion state, it can be determined that there may be many abnormally large values in the first speed caused by the shooting device in the motion state, that is, the fourth speed may be the average value of the abnormally large values. On this basis, if the fourth speed is less than the product of the fourth value and the third speed of each target area, it can be considered that the fourth speed is not the average value of the abnormally large values, and there may be no abnormal values among the larger values removed in the above first method. Therefore, the accuracy of using the third speed obtained by the first method as the speed of the target area is relatively low, and the fourth speed can be used as the speed of the target area (such as the second speed); or, if the fourth speed is greater than or equal to the product of the fourth value and the third speed of each target area, it can be considered that the fourth speed is the average value of the abnormally large values, and it can also be considered that there may indeed be abnormal values among the larger values removed in the above first method, and the accuracy of using the third speed as the speed of the target area is relatively high. Therefore, the third speed can be used as the speed of the target area (such as the second speed).
[0142] Through the above embodiments, the second speed of the target area can be determined according to the shooting state of the shooting device, the third speed and the fourth speed of the target area, thereby avoiding misjudgment of the movement speed of the target area caused by the shooting state of the shooting device and improving the calculation accuracy of the movement speed of the target area.
[0143] S206. Determine the area where there is a moving object in the first image frame among the two image frames according to the target area, and determine the movement speed of the moving object according to the second speed.
[0144] In an embodiment of the present application, the accuracy of directly using the area corresponding to the target area in the first image frame as the area where there is a moving object may be relatively low, because there may be a situation where the entire target area is a noise area. Therefore, denoising processing can be performed on the target area to eliminate the noise area in the target area and improve the accuracy of moving area detection.
[0145] Since the target optical flow map in the above embodiment is obtained by performing multiple denoising processes on the initial optical flow map, and the target area is obtained by filtering isolated noises (such as isolated target sub-areas) in the target optical flow map, the noise area in the target area may be a noise area with a relatively large area. The noise area in the target area can be determined by performing noise inspection on a preset number of target areas with the largest areas, so as to achieve denoising processing on the target area.
[0146] In an embodiment, the denoising processing on the target area may include: sorting all target areas in descending order of the area of the target area, selecting a preset number of target areas at the front from the sorted sequence, and determining the noise area from the selected target areas. The preset number represents a positive integer less than or equal to the total number of target areas and can be set according to actual needs. For example, the preset number can be 7, and the present application does not make specific limitations on this.
[0147] In an embodiment of the present application, the area of any target area is proportional to the number of the target sub-areas. Therefore, when sorting all target areas in descending order of the area of the target area, all target areas can be sorted in descending order of the number of target sub-areas included in the target area.
[0148] In an embodiment of the present application, when determining a noise area from a selected target area, the corresponding area of the selected target area in the first image frame or the second image frame can be determined, and a traditional difference map between the two image frames of the corresponding area can be obtained; for the corresponding area of any selected target area, the average difference value of the difference values in the traditional difference map of the area can be calculated; a difference threshold can be set for any sub-area according to the number of sub-areas in any selected target area; if the average difference value of any selected target area is less than or equal to the difference threshold corresponding to any sub-area, it can be considered that any sub-area is a noise area.
[0149] Through the above embodiments, the corresponding difference threshold can be determined according to the number of sub-areas of each target area, so as to determine the noise area in the target area according to the area size of each target area, thereby improving the recognition accuracy of the moving area.
[0150] In an embodiment of the present application, the method for determining the noise area in the target area may also refer to the description of the embodiment shown below for Figure 16 shown embodiment.
[0151] In an embodiment of the present application, after denoising the target area, the area corresponding to each remaining target area in the first image frame can be used as an area where a moving object exists, and the second speed of each remaining target area can be used as the moving speed of the moving object in the area where the corresponding moving object exists.
[0152] Through the above-mentioned multiple embodiments, the optical flow estimation algorithm of the chip of the terminal device can be called when taking pictures to obtain an initial optical flow map between two adjacent image frames; by denoising the initial optical flow map, the noise and abnormal optical flow in the initial optical flow map can be gradually removed; by dividing the target optical flow map into multiple sub-areas, the connected areas where moving objects may exist in the target optical flow map can be determined according to the optical flow values in the target optical flow map; by screening the connected areas, the areas where moving objects exist in the target optical flow map can be accurately determined, avoiding misdetecting the noise area in the static state as a moving object, and improving the accuracy of calculating the moving speed of the moving object. Further, the exposure parameters of the shooting device can be adjusted according to the area where the moving object exists and the moving speed of the moving object, thereby improving the image quality (such as clarity) of the captured image.
[0153] Refer to as Figure 8 shown, which is a flowchart of a method for denoising an initial optical flow map provided by an embodiment of the present application. The method is applied to a terminal device, and the method for denoising the initial optical flow map includes:
[0154] S301. Denoise the initial optical flow map based on the optical flow values of each pixel in the initial optical flow map and a preset optical flow value range to obtain an updated optical flow map.
[0155] In an embodiment of the present application, if the displacement value of any pixel in the initial optical flow map (hereinafter simply referred to as the "third displacement value" to distinguish it from other displacement values in the above text) is not within the optical flow value range, update the optical flow value of any pixel to a preset first value; or, if the third displacement value of any pixel is within the optical flow value range, maintain the optical flow value of any pixel. Among them, the optical flow value range can be set according to actual needs. For example, the optical flow value range can be [a, b] = [1, 60], and the present application does not make specific limitations on this.
[0156] The above denoising processing method can also be expressed as: if g(x, y) ≤ a or g(x, y) ≥ b, then let G(x, y) = 0; or, if a < g(x, y) < b, then let G(x, y) = g(x, y), where g(x, y) represents the optical flow value (such as the third displacement value) of the pixel with coordinates (x, y) in the initial optical flow map, and G(x, y) represents the optical flow value (such as the third displacement value) of the pixel with coordinates (x, y) in the updated optical flow map.
[0157] Through the above embodiments, denoising of abnormal small optical flow value noise (such as small optical flow value noise caused by small shaking of the shooting device) and abnormal large optical flow value noise (such as large optical flow value noise caused by the optical flow estimation algorithm) in the initial optical flow map can be achieved.
[0158] S302. Based on the first image frame and the second image frame in two image frames, use an improved difference algorithm to obtain a denoised difference map.
[0159] In an embodiment of the present application, refer to as Figure 9 shown, which is a flowchart of the detailed process of S302 provided by an embodiment of the present application. The detailed process of S302 includes:
[0160] S3021. Determine the first grayscale value of each pixel in the first image frame and the second grayscale value of the corresponding pixel in the second image frame.
[0161] S3022. Based on the first grayscale value and the second grayscale value corresponding to each pixel, use an improved difference algorithm to perform difference calculation on the first image frame and the second image frame to obtain an initial difference image between the first image frame and the second image frame.
[0162] In an embodiment of the present application, the formula used in the improved difference algorithm can refer to the description in S203 in the process shown above for Figure 4 shown.
[0163] S3023, perform binarization processing on the initial difference image to obtain a binarized image.
[0164] In an embodiment of the present application, performing binarization processing on the initial difference image includes: if the gray value of any pixel point in the initial difference image is less than a preset gray threshold, then update the gray value of any pixel point to a preset first value; or, if the gray value of any pixel point is greater than or equal to the gray threshold, update the gray value of any pixel point to a preset second value. Wherein, when the gray value of a pixel point is the first value (for example, 0), it represents a pixel point in a black background area (or called a stationary area), and when the gray value of a pixel point is the second value (for example, 1 or 255), it represents a pixel point in a white foreground area (or called a moving area). The specific values of the first value and the second value can be set according to actual needs, and the present application does not make specific limitations on this.
[0165] In an embodiment of the present application, the formula used for binarization processing may include:
[0166]
[0167] where D n (x, y) represents the gray value of the pixel point with coordinates (x, y) in the initial difference image D n , R n (x, y) represents the gray value of the pixel point with coordinates (x, y) in the binarized image R n . 0 represents the first value, 1 represents the second value, and T represents the gray threshold (for example, 3).
[0168] Through the above embodiments, the gray value of the pixel points in the initial difference image can be updated to preset values, so that the moving pixel points and stationary pixel points can be accurately identified according to the obtained binarized image.
[0169] S3024, perform post-processing operations on the binarized image to obtain a denoised difference image.
[0170] In an embodiment of the present application, the post-processing operations may include, but are not limited to, a combination of one or more of the following methods: image erosion processing, image dilation processing.
[0171] In an embodiment of the present application, the image erosion processing may include: defining a structural element (such as a rectangle, circle, cross, etc.); placing the center point of the structural element at each pixel point of the image (such as the binarized image), determining the minimum value of the gray values of all pixel points within the coverage area of the structural element, and updating the gray value of the pixel point at the center point to this minimum value.
[0172] By performing image erosion processing on the binary image, the protruding points on the periphery of the foreground region of the binary image can be eroded. By shrinking the boundary of the foreground region, noise and small discontinuous regions can be removed, and the boundaries of different moving objects that may exist in the foreground region can be distinguished, and the moving objects and the static background can be distinguished more accurately.
[0173] In an embodiment of the present application, the image dilation processing may include: defining a structural element (such as a rectangular, circular, cross-shaped, etc. element); placing the center point of the structural element at each pixel point of the image (such as a binary image), determining the maximum value of the gray values of all pixel points within the coverage area of the structural element, and updating the gray value of the pixel point at the center point to this maximum value.
[0174] By performing image dilation processing on the binary image, pixel points can be expanded on the boundary of the foreground region in the image (such as a binary image) to achieve the effect of widening the boundary of the foreground region, making the edges of different foreground regions smoother and more continuous.
[0175] S303, denoise the updated optical flow map using the denoised difference map to obtain the target optical flow map.
[0176] In an embodiment of the present application, an AND operation can be performed on a per-pixel basis between the denoised difference map and the updated optical flow map, and the optical flow value of each pixel point in the updated optical flow map can be updated according to the result of the AND operation. For example, if the gray value of any pixel point in the denoised difference map is equal to a preset first value, the optical flow value of the pixel point corresponding to any pixel point in the updated optical flow map is updated to the first value (such as 0); or, if the gray value of any pixel point is equal to a preset second value (such as 1), the optical flow value of the pixel point corresponding to any pixel point in the updated optical flow map is maintained.
[0177] The above method of denoising the updated optical flow map using the denoised difference map can also be expressed as: if Dq(x,y) = 0, then let G ′ (x,y) = 0; or, if Dq(x,y) = 1, then let G ′ (x,y) = G(x,y), where Dq(x,y) represents the gray value of the pixel point with coordinates (x,y) in the denoised difference map, G(x,y) represents the optical flow value of the pixel point with coordinates (x,y) in the updated optical flow map (such as the third displacement value), and G ′ (x,y) represents the optical flow value of the pixel point with coordinates (x,y) in the target optical flow map (such as the third displacement value).
[0178] Through the above embodiments, the optical flow values corresponding in the updated optical flow map can be reset to 0 according to the stationary pixel points indicated by the denoised difference map (the pixel points with a black background gray value of 0), and the optical flow values of the corresponding pixel points in the updated optical flow map can be maintained unchanged according to the moving pixel points indicated by the denoised difference map (the pixel points with a white foreground gray value of 1), so as to further remove the noise caused by the abnormal optical flow values in the updated optical flow map.
[0179] Refer to as Figure 14 shown, which is a flowchart of a method for determining the shooting state of a shooting device provided by an embodiment of the present application. The method is applied to a terminal device, and the method for determining the shooting state of the shooting device includes:
[0180] S401, determine the ratio of the number of all target sub-regions in the target optical flow map to the total number of all sub-regions in the target optical flow map.
[0181] In an embodiment of the present application, a target sub-region represents a sub-region where the first speed is greater than a preset speed threshold. Referring to the description in S205, the number I of elements "1" in the motion state matrix can be determined as the number of all target sub-regions in the target optical flow map. Referring to the description in S204, the total number of all sub-regions in the target optical flow map can be M×N. The ratio of the number I of all target sub-regions in the target optical flow map to the total number M×N of all sub-regions in the target optical flow map can be exemplarily expressed as:
[0182]
[0183] S402, determine whether the above ratio is greater than a preset ratio threshold. If the above ratio is greater than the ratio threshold, execute S403; or, if the above ratio is less than or equal to the ratio threshold, execute S404.
[0184] In an embodiment of the present application, the preset ratio threshold can be set according to actual needs. The value range of this ratio threshold is (0,1). For example, this ratio threshold can be 0.75, and the present application does not make specific limitations on this.
[0185] If the above ratio is greater than the ratio threshold, in S403, determine that the shooting state of the shooting device is a motion state.
[0186] In an embodiment of the present application, when the above ratio is greater than the ratio threshold, it can be considered that most of the optical flow in the target optical flow map is in motion. This situation can be considered to be due to the movement of the shooting device (for example, moving and swaying), resulting in a full-frame change of the picture. Therefore, it can be determined that the shooting state of the shooting device is a motion state.
[0187] If the above ratio is less than or equal to the ratio threshold, at S404, determine that the shooting state of the shooting device is an approximately stationary state.
[0188] In an embodiment of the present application, when the above ratio is less than or equal to the ratio threshold, it can be considered that the moving optical flow in the target optical flow map is caused by local moving objects in the scene. Therefore, it can be determined that the shooting state of the shooting device is an approximately stationary state. For example, the state of the camera is the user's hand-held mode or the tripod mode.
[0189] Through the above embodiments, it is possible to determine the shooting state of the shooting device corresponding to two image frames of the target optical flow map according to the proportion of the moving sub-regions in the target optical flow map. For example, if the proportion of the moving sub-regions in the target optical flow map is relatively large, it can be considered that there are more moving sub-regions in the target optical flow map due to the movement of the shooting device. Therefore, it can be considered that the shooting device is in a moving state; or, if the proportion of the moving sub-regions in the target optical flow map is relatively small, it can be considered that the relatively few moving sub-regions in the target optical flow map are the regions corresponding to the moving objects in the shooting scene, and the shooting device is approximately stationary. Therefore, it can be considered that the shooting device is in an approximately stationary state.
[0190] Refer to as Figure 15 shown, which is a flowchart of a method for determining the third speed and the fourth speed of a target area provided by an embodiment of the present application. The method is applied to a terminal device. The method for determining the third speed and the fourth speed of the target area includes:
[0191] S501, clean the outliers in the first speeds of all sub-regions in each target area, and determine the third speed of each target area according to the average value of the cleaned first speeds.
[0192] In an embodiment of the present application, the outliers in the first speeds can be cleaned according to the principle of normal distribution. The principle includes: according to the standard normal distribution, about 99.7% of the data will fall within the range of the average value μ plus or minus three standard deviations σ. Therefore, cleaning the outliers in the first speeds of all sub-regions in each target area can include: calculating the average value μ and the standard deviation σ of the first speeds of all sub-regions in each target area; determining the range of normal values, which can be expressed as [μ - 3σ, μ + 3σ]; taking the first speeds of all sub-regions in each target area that are not within the range of normal values as outliers, and cleaning the outliers or resetting them to 0.
[0193] Through the above embodiments, it is possible to remove the smaller values and the larger values in the first speeds based on the principle of normal distribution, and obtain the average value of the first speeds after filtering out the possible outliers as the third speed of the target area.
[0194] S502. Based on the sliding window algorithm, determine the multiple sliding window averages corresponding to the first speeds of all sub-regions of each target region, and determine the fourth speed of each target region according to the maximum value among the multiple sliding window averages.
[0195] In an embodiment of the present application, based on the sliding window algorithm, determining the multiple sliding window averages corresponding to the first speeds of all sub-regions of each target region may include: sorting the first speeds of all sub-regions of each target region in descending order to obtain a sequence composed of the first speeds of all sub-regions of each target region; determining the size of the window, that is, the number of first speeds that the window can accommodate. The size of the window may be less than the total number of first speeds in the above sequence. For example, if the total number of first speeds in the above sequence is 5, the window size may be 3; initialize the window position, make the first data in the window be the first first speed in the above sequence, calculate the average value l1 of the first speeds in the window to obtain the first sliding window average value l1; move the window position one data position backward, make the first data in the window be the second first speed in the above sequence, calculate the average value l2 of the first speeds in the window to obtain the second sliding window average value l2; and so on, continuously move the sliding window backward until the last data in the window is the last first speed in the above sequence, or until the sliding window average value l p , take the position where the window is located at this time as the last placement position of the window, where p represents a preset positive integer. For example, p = 3; determine the sliding window average value l of the first speeds in the window at the i-th position in the above process i , where, l i+1 ≤l i , and i represents a positive integer.
[0196] In an embodiment of the present application, when determining the fourth speed of each target region according to the maximum value among the multiple sliding window averages, the multiple sliding window averages may be arranged in descending order, and select the first p ′ sliding window averages: l1, l2,..., l p , where p ′ represents a preset positive integer less than or equal to p. For example, when p ′ = 3, the first p ′ sliding window averages are l1, l2, l3; if the difference between every two adjacent sliding window averages among the first p ′ sliding window averages is less than or equal to the preset dynamic parameter θ, then select the maximum value among the multiple sliding window averages as the fourth speed of the target region. For example, if l1 - l2 ≤ θ and l2 - l3 ≤ θ, then determine the fourth speed = l1.
[0197] In an embodiment of the present application, the dynamic parameter θ can be determined according to the maximum speed of all grids in each connected region. For example, if the maximum speed is greater than the preset threshold of 10, θ can be set to 1; or, if the maximum speed is less than or equal to the preset threshold of 10, θ can be set to the maximum speed × 0.1, and the value range of θ can be [0, 1].
[0198] Through the above embodiment, the average value corresponding to the larger value in the first speed can be obtained as the fourth speed of the target region by the sliding window algorithm.
[0199] Refer to Figure 16 As shown in the figure, it is a flowchart of a method for determining a noise region in a target region provided by an embodiment of the present application. The method is applied to a terminal device, and the method for determining a noise region in a target region includes:
[0200] S601, determine the corresponding region of each target region in the first image frame or the second image frame, and obtain the traditional difference map between the two image frames of the corresponding region.
[0201] In an embodiment of the present application, the position of each pixel point in each target region (for example, each target region in the target region selected in S206) in the entire target optical flow map can be determined, so as to determine the corresponding region of each target region in the first image frame or the second image frame according to the set of positions of all pixel points in each target region.
[0202] For example Figure 17 As shown in the figure, it is an example diagram of pixel points of a target region provided by an embodiment of the present application. Among them, Figure 17 The left image shows that the target optical flow map is divided into 64×48 sub-regions, and each target sub-region is represented by a black small block, and each non-target region is represented by a gray small block. The region formed by the black small blocks within the rectangular frame in the left image is a target region, where each black small block represents a target sub-region, and each target sub-region includes 5×5 pixel points as shown in the Figure 17 right image. The positions of the pixel points of each target sub-region are known. Therefore, the corresponding region of each target region in the first image frame or the second image frame can be determined by the positions of each pixel point of each target region in the entire target optical flow map.
[0203] In an embodiment of the present application, each pixel point in the traditional difference map corresponds to a difference value. The method for obtaining the traditional difference map can refer to the description in S203. By calculating the traditional difference map of the corresponding region of each target region, the difference calculation result can be made more accurate.
[0204] S602. Determine the average difference value of the difference values of all pixel points in the traditional difference map corresponding to each target area.
[0205] In an embodiment of the present application, the average difference value may be equal to the ratio of the sum of the difference values of all pixel points in the traditional difference map corresponding to each target area to the number of all pixel points in each target area.
[0206] S603. Determine the difference threshold corresponding to each target area according to the number of sub-areas within each target area.
[0207] In an embodiment of the present application, the difference threshold corresponding to each target area may be determined according to the following formula:
[0208]
[0209] where threshold represents the difference threshold, num represents the number of sub-areas (such as target sub-areas) within each target area, and α, λ, β, γ, k represent preset parameters. For example, α = -0.0125, λ = 18.125, β = -0.02, γ = 10.0, k = 250. The above parameters can be determined according to prior data, and the present application does not make specific limitations thereto. Through the above embodiments, a larger difference threshold can be set for a target sub-area with a larger area.
[0210] S604. Determine the target area corresponding to the average difference value less than or equal to the corresponding difference threshold as the noise area.
[0211] In an embodiment of the present application, if the average difference value is less than or equal to the difference threshold, it can be considered that the corresponding target area is a noise area; if the average difference value is greater than the difference threshold, it can be considered that the corresponding target area is a non-noise area with a moving object.
[0212] Through the above embodiments, the corresponding difference threshold can be determined according to the number of sub-areas of each target area, so that the noise area in the target area can be determined according to the difference average value and the difference threshold of the traditional difference map corresponding to each target area, improving the recognition accuracy of the moving area.
[0213] Refer to Figure 18As shown in the figure, it is a flowchart of a method for differential denoising of optical flow provided by an embodiment of the present application. The method for differential denoising of optical flow may include: performing threshold denoising on the input original optical flow (such as an initial differential map), where the threshold range may be [a, b], resetting the optical flow values outside the threshold range to 0 to obtain a first-stage threshold-denoised optical flow; using an improved differential algorithm to perform differential calculation on two input image frames (such as a first image frame and a second image frame) to obtain a differential map; performing binarization processing on the differential map to obtain a binarized differential map; performing erosion processing on the binarized differential map to obtain an eroded differential map; performing dilation processing on the eroded differential map to obtain a denoised differential map; performing a pixel-by-pixel AND operation on the first-stage threshold-denoised optical flow and the denoised differential map to obtain a second-stage differential-denoised optical flow.
[0214] Through the above embodiments, the noise in the initial optical flow can be initially removed according to the preset optical flow value range to obtain a first-stage threshold-denoised optical flow; a denoised differential map of two image frames can be obtained according to the improved differential algorithm, and the denoised differential map is used to further denoise the first-stage threshold-denoised optical flow, which can focus on distinguishing the differences between moving and stationary pixel points in the updated optical flow. By resetting the optical flow values in the first-stage threshold-denoised optical flow, the optical flow noise that interferes with the distinction between moving and stationary regions can be effectively removed.
[0215] Refer to Figure 19 As shown in the figure, it is a flowchart of a motion detection method provided by an embodiment of the present application. The motion detection method may include: dividing the denoised optical flow map (such as a target optical flow map) into multiple grids (such as multiple sub-regions); calculating the average displacement value of the optical flow (such as a first speed) for each grid; selecting motion grids (such as target sub-regions) where motion objects may exist according to the average displacement value; filtering the noise grids in the motion grids, such as filtering non-boundary isolated motion grids (such as isolated target sub-regions); calculating the connected graph (such as a target region) formed by the filtered motion grids; in a preset order (such as, in accordance with Figure 13In the order of numbers 2 to 7, the following process is performed for each connected graph: Calculate two possible motion speeds of the current connected graph (such as the third speed and the fourth speed); Determine the motion state of the camera (such as the camera being in a motion state, or an approximately stationary state where the user holds it by hand or places it on a tripod), and verify the two possible motion speeds according to the state of the camera to determine the final speed of the current connected graph (such as the second speed); Determine whether the current connected graph is a noise area. If the current connected graph is not a noise area, output the speed of the current connected graph (such as the second speed) as the motion speed of the moving object in the corresponding area of the image frame (such as the first image frame). If there are other connected graphs that have not been processed with the above process, perform the above process on the next connected area in the above order. In another embodiment, the input of the motion detection method can also be the optical flow without noise removal, because a certain denoising effect can also be achieved during the motion detection process.
[0216] Through the above embodiments, it is possible to accurately determine the motion speed of each connected area according to the state of the shooting device, and remove the noise areas from the connected areas, improving the calculation accuracy of the motion speed of the motion area and the moving object.
[0217] Refer to Figure 20 As shown, it is a flowchart of a method for noise verification of a connected graph provided by an embodiment of the present application. The method for noise verification of a connected graph may include: performing an AND operation of full-image difference on the areas corresponding to the connected graphs in two frames of images (such as the first image frame and the second image frame) to obtain a connected sub-graph area difference image of the area corresponding to the connected graph; determining the average difference value of each connected sub-graph area difference image; determining a corresponding difference threshold according to the area size of each connected graph; and determining whether the connected graph is a noise area by comparing the average difference value of each connected graph with the difference threshold.
[0218] Through the above embodiments, it is possible to determine the noise areas in the target area and improve the recognition accuracy of the motion area.
[0219] An embodiment of the present application also provides a terminal device 100. Refer to Figure 21As shown, the terminal device 100 may be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an Ultra-mobile Personal Computer (UMPC), a netbook, a cellular phone, a Personal Digital Assistant (PDA), an Augmented Reality (AR) device, a Virtual Reality (VR) device, an Artificial Intelligence (AI) device, a wearable device, a vehicle-mounted device, a smart home device, and / or a smart city device. The embodiments of the present application do not impose special restrictions on the specific type of the terminal device 100.
[0220] The terminal device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a Universal Serial Bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a Subscriber Identification Module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0221] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0222] The processor 110 may include one or more processing units. For example, the processor 110 may include an Application Processor (AP), a modem processor, a Graphics Processing Unit (GPU), an Image Signal Processor (ISP), a controller, a video codec, a Digital Signal Processor (DSP), a baseband processor, and / or a Neural-network Processing Unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0223] The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0224] A memory may also be provided in the processor 110 for storing instructions and data. In an embodiment of the present application, the memory in the processor 110 is a cache memory. The memory may save the instructions or data just used or recycled by the processor 110. If the processor 110 needs to use the instructions or data again, it can directly call them from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0225] In an embodiment of the present application, the processor 110 may include one or more interfaces. The interfaces may include an Inter-integrated Circuit (I2C) interface, an Inter-integrated Circuit Sound (I2S) interface, a Pulse Code Modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a Mobile Industry Processor Interface (MIPI), a General-Purpose Input / Output (GPIO) interface, a Subscriber Identity Module (SIM) interface, and / or a Universal Serial Bus (USB) interface, etc.
[0226] The I2C interface is a bidirectional synchronous serial bus, including a Serial Data Line (SDA) and a Serial Clock Line (SCL). In an embodiment of the present application, the processor 110 may include multiple groups of I2C buses. The processor 110 can be respectively coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces. For example, the processor 110 can be coupled to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface to implement the touch function of the terminal device 100.
[0227] The I2S interface can be used for audio communication. In an embodiment of the present application, the processor 110 may include multiple groups of I2S buses. The processor 110 can be coupled to the audio module 170 through the I2S bus to achieve communication between the processor 110 and the audio module 170. In an embodiment of the present application, the audio module 170 can transmit audio signals to the wireless communication module 160 through the I2S interface to implement the function of answering a call through a Bluetooth headset.
[0228] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In an embodiment of the present application, the audio module 170 and the wireless communication module 160 can be coupled through the PCM bus interface. In an embodiment of the present application, the audio module 170 can also transmit audio signals to the wireless communication module 160 through the PCM interface to implement the function of answering a call through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0229] The UART interface is a general-purpose serial data bus for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In an embodiment of the present application, the UART interface is usually used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface to implement the Bluetooth function. In an embodiment of the present application, the audio module 170 can transmit audio signals to the wireless communication module 160 through the UART interface to implement the function of playing music through a Bluetooth headset.
[0230] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes the Camera Serial Interface (CSI), the Display Serial Interface (DSI), etc. In an embodiment of the present application, the processor 110 and the camera 193 communicate through the CSI interface to implement the shooting function of the terminal device 100. The processor 110 and the display screen 194 communicate through the DSI interface to implement the display function of the terminal device 100.
[0231] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In an embodiment of the present application, the GPIO interface can be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0232] The USB interface 130 is an interface that conforms to the USB standard specification, and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the terminal device 100, and can also be used to transfer data between the terminal device 100 and peripheral devices. It can also be used to connect headphones to play audio. The interface can also be used to connect other terminal devices 100, such as AR devices, etc.
[0233] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present invention is only for illustrative purposes and does not constitute a structural limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 can also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0234] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 can receive the charging input from a wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 can receive the wireless charging input through the wireless charging coil of the terminal device 100. While charging the battery 142, the charging management module 140 can also supply power to the terminal device 100 through the power management module 141.
[0235] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives the inputs from the battery 142 and / or the charging management module 140, and powers the processor 110, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as the battery capacity, the number of battery cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 can also be disposed in the same device.
[0236] The wireless communication function of the terminal device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.
[0237] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the terminal device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: the antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0238] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the terminal device 100. The mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by the antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In an embodiment of the present application, at least some functional modules of the mobile communication module 150 can be disposed in the processor 110. In an embodiment of the present application, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 can be disposed in the same device.
[0239] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.), or displays an image or video through the display screen 194. In an embodiment of the present application, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 110 and be disposed in the same device as the mobile communication module 150 or other functional modules.
[0240] The wireless communication module 160 may provide solutions for wireless communications applied to the terminal device 100, including Wireless Local Area Networks (WLAN) (such as Wireless Fidelity (Wi-Fi) networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), Infrared (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signal, and transmits the processed signal to the processor 110. The wireless communication module 160 may also receive the signal to be transmitted from the processor 110, perform frequency modulation and amplification on it, and convert it into electromagnetic waves through the antenna 2 and radiate it out.
[0241] In an embodiment of the present application, the antenna 1 of the terminal device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, enabling the terminal device 100 to communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include Global System For Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), Beidou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).
[0242] The terminal device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0243] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In an embodiment of the present application, the terminal device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.
[0244] The terminal device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0245] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin color of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In an embodiment of the present application, the ISP can be set in the camera 193.
[0246] The camera 193 is used to capture static images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In an embodiment of the present application, the terminal device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.
[0247] The digital signal processor is used to process digital signals. In addition to being able to process digital image signals, it can also process other digital signals. For example, when the terminal device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0248] The video codec is used to compress or decompress digital videos. The terminal device 100 can support one or more video codecs. In this way, the terminal device 100 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0249] The NPU is a Neural-Network (NN) computing processor. By learning from the biological neural network structure, such as learning from the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the terminal device 100 can be realized, such as: image recognition, face recognition, moving object detection, speech recognition, text understanding, etc.
[0250] The internal memory 121 may include one or more Random Access Memories (RAM) and one or more Non-Volatile Memories (NVM).
[0251] The random access memory may include Static Random-Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM, for example, the fifth-generation DDR SDRAM is generally called DDR5 SDRAM), etc.;
[0252] The non-volatile memory may include disk storage devices, flash memory.
[0253] Flash memory can be classified into NOR Flash, NAND Flash, 3D NAND Flash, etc. according to the operating principle, single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), etc. according to the number of potential levels of storage cells, and universal flash storage (UFS), embedded multi-media card (eMMC), etc. according to the storage specification.
[0254] The random access memory can be directly read and written by the processor 110, and can be used to store the operating system or executable programs (such as machine instructions) of other running programs, and can also be used to store data of users and application programs, etc.
[0255] The non-volatile memory can also store executable programs and data of users and application programs, etc., and can be pre-loaded into the random access memory for direct reading and writing by the processor 110.
[0256] The external memory interface 120 can be used to connect to an external non-volatile memory to expand the storage capacity of the terminal device 100. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external non-volatile memory.
[0257] The internal memory 121 or the external memory interface 120 is used to store one or more computer programs. One or more computer programs are configured to be executed by the processor 110. One or more computer programs include multiple instructions. When the multiple instructions are executed by the processor 110, the screen display detection method executed on the terminal device 100 in the above embodiments can be implemented to realize the screen display detection function of the terminal device 100.
[0258] The terminal device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, and the application processor, etc. For example, music playback, recording, etc.
[0259] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In an embodiment of the present application, the audio module 170 can be disposed in the processor 110, or a partial functional module of the audio module 170 can be disposed in the processor 110.
[0260] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The terminal device 100 can listen to music or hands-free calls through the speaker 170A.
[0261] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the terminal device 100 answers a call or a voice message, the voice can be listened to by bringing the receiver 170B close to the human ear.
[0262] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The terminal device 100 can be provided with at least one microphone 170C. In some other embodiments, the terminal device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the terminal device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0263] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5 mm Open Mobile Terminal Platform (OMTP) standard interface, or a Cellular Telecommunications Industry Association of the USA (CTIA) standard interface.
[0264] The keys 190 include a power-on key, volume keys, etc. The keys 190 can be mechanical keys or touch keys. The terminal device 100 can receive key inputs to generate key signal inputs related to the user settings and function controls of the terminal device 100.
[0265] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and also for touch vibration feedback. For example, touch operations applied to different applications (such as taking pictures, audio playing, etc.) can correspond to different vibration feedback effects. For touch operations applied to different areas of the display screen 194, the motor 191 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminder, receiving messages, alarm clock, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0266] The indicator 192 can be an indicator light and can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc.
[0267] The SIM card interface 195 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation from the terminal device 100. The terminal device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The terminal device 100 interacts with the network through the SIM card to achieve functions such as calls and data communication. In an embodiment of the present application, the terminal device 100 uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the terminal device 100 and cannot be separated from the terminal device 100. The embodiment of the present application also provides a computer storage medium, and computer instructions are stored in the computer storage medium. When the computer instructions run on the terminal device 100, the terminal device 100 is caused to execute the above-related method steps to implement the moving object detection method in the above embodiment.
[0268] The embodiment of the present application also provides a computer program product. When the computer program product runs on a computer, the computer is caused to execute the above-related steps to implement the moving object detection method in the above embodiment.
[0269] In addition, the embodiment of the present application also provides a device, which can specifically be a chip, component or module. The device can include a processor and a memory connected to each other; wherein, the memory is used to store computer execution instructions. When the device runs, the processor can execute the computer execution instructions stored in the memory to cause the chip to execute the moving object detection method in each of the above method embodiments.
[0270] Among them, the terminal device, computer storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0271] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0272] In several embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0273] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0274] In addition, each functional unit in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0275] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0276] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. A method for detecting a moving object, applied to a terminal device, characterized in that, The method includes: In response to a user's shooting operation, obtaining two adjacent image frames by using the shooting device of the terminal device; Calculating an initial optical flow map between the two image frames by using a preset optical flow estimation algorithm; Performing denoising processing on the initial optical flow map to obtain a target optical flow map; Dividing the target optical flow map into multiple sub-regions, and determining a first speed corresponding to each sub-region according to the optical flow values of each pixel point in each sub-region; Determining multiple target regions in the target optical flow map according to the first speed, and determining a second speed corresponding to each target region, where each target region includes at least one sub-region; Determining the region where a moving object exists in the first image frame of the two image frames according to the target region, and determining the moving speed of the moving object according to the second speed.
2. The motion object detection method according to claim 1, wherein Each pixel point in the initial optical flow map includes at least a first channel and a second channel, where the first channel includes a first displacement value in a first direction, and the second channel includes a second displacement value in a second direction; The optical flow value of each pixel point in the initial optical flow map includes the first displacement value, the second displacement value, and a third displacement value corresponding to a displacement vector determined according to the first displacement value and the second displacement value.
3. The method for detecting a moving object according to claim 1 or 2, wherein The performing denoising processing on the initial optical flow map to obtain a target optical flow map includes: Performing denoising processing on the initial optical flow map based on the optical flow value of each pixel point in the initial optical flow map and a preset optical flow value range to obtain an updated optical flow map; Based on the first image frame and the second image frame in the two image frames, obtaining a denoising difference map by using an improved difference algorithm; Performing denoising processing on the updated optical flow map by using the denoising difference map to obtain the target optical flow map.
4. The motion object detection method according to claim 3, wherein The optical flow value of each pixel point in the initial optical flow map includes a third displacement value, and the performing denoising processing on the initial optical flow map based on the optical flow value of each pixel point in the initial optical flow map and a preset optical flow value range includes: If the third displacement value of any pixel point in the initial optical flow map is not within the optical flow value range, updating the optical flow value of the any pixel point to a preset first value; or, If the third displacement value of the any pixel point is within the optical flow value range, maintaining the optical flow value of the any pixel point.
5. The motion object detection method according to claim 3, wherein The obtaining a denoising difference map by using an improved difference algorithm based on the first image frame and the second image frame in the two image frames includes: Determining a first gray value of each pixel point in the first image frame, and a second gray value of the corresponding pixel point in the second image frame; Based on the first gray value and the second gray value corresponding to each pixel point, performing difference calculation on the first image frame and the second image frame by using the improved difference algorithm to obtain an initial difference image between the first image frame and the second image frame; Performing binarization processing on the initial difference image to obtain a binarized image; Performing a post-processing operation on the binarized image to obtain the denoising difference map.
6. The method for detecting a moving object according to claim 5, wherein, The formula used in the improved difference algorithm includes: Among them, D n (x, y) represents the gray value of the pixel point with coordinates (x, y) in the initial difference image, f n (x, y) represents the first gray value of the pixel point with coordinates (x, y) in the first image frame f n in, and f n-1 (x, y) represents the second gray value of the pixel point with coordinates (x, y) in the second image frame f n-1 in.
7. The method for detecting a moving object according to claim 5, wherein The binary processing of the initial difference image includes: If the gray value of any pixel point in the initial difference image is less than a preset gray threshold, update the gray value of the any pixel point to a preset first value; or, If the gray value of the any pixel point is greater than or equal to the gray threshold, update the gray value of the any pixel point to a preset second value.
8. The method for detecting a moving object according to claim 3, wherein The denoising process of the updated optical flow map using the denoised difference map includes: Perform a pixel-by-pixel AND operation on the denoised difference map and the updated optical flow map, and update the optical flow value of each pixel point in the updated optical flow map according to the result of the AND operation.
9. The motion object detection method according to claim 8, wherein The updating of the optical flow value of each pixel point in the updated optical flow map according to the result of the AND operation includes: If the gray value of any pixel point in the denoised difference map is equal to a preset first value, update the optical flow value of the pixel point corresponding to the any pixel point in the updated optical flow map to the first value; or, If the gray value of the any pixel point is equal to a preset second value, maintain the optical flow value of the pixel point corresponding to the any pixel point in the updated optical flow map.
10. The motion object detection method according to claim 1, wherein The optical flow value of each pixel point in each sub-region includes a first displacement value and a second displacement value. The method of dividing the target optical flow map into multiple sub-regions and determining the first speed corresponding to each sub-region according to the optical flow value of each pixel point in each sub-region includes: Determine a first average value of the first displacement values of all pixel points in each sub-region, and a second average value of the second displacement values of all pixel points. Based on the first average value and the second average value, determine the first speed.
11. The motion object detection method according to claim 1, characterized in that, The determining of multiple target regions in the target optical flow map according to the first speed includes: If the first speed of any sub-region is greater than a preset speed threshold, use the any sub-region as a target sub-region; Based on the connected graph algorithm, determine the connected region where each target sub-region is located, and use the connected region as the target region, where each connected region includes at least one target sub-region.
12. The method for detecting a moving object according to claim 11, wherein, Before determining the connected region to which each target sub-region belongs based on the connected graph algorithm, the method further includes noise filtering of the target optical flow map, including: Update the optical flow value of the isolated target sub-region at the non-boundary position in the target optical flow map to a preset first value, where the isolated target sub-region includes a target sub-region that is not adjacent to other target sub-regions.
13. The motion object detection method according to claim 1, wherein The method further includes: Clean the outliers in the first speeds of all sub-regions in each target region, and determine the third speed of each target region according to the average value of the cleaned first speeds; Based on the sliding window algorithm, determine multiple sliding window average values corresponding to the first speeds of all sub-regions of each target region, and determine the fourth speed of each target region according to the maximum value among the multiple sliding window average values.
14. The method for detecting a moving object according to any one of claims 1 to 13, characterized in that, The method further includes: determining the shooting state of the shooting device according to the target optical flow map, including: Determine the ratio of the number of all target sub-regions in the target optical flow map to the total number of all sub-regions in the target optical flow map, where the target sub-region represents a sub-region with a first speed greater than a preset speed threshold; If the ratio is greater than a preset ratio threshold, determine that the shooting state of the shooting device is a motion state; or, If the ratio is less than or equal to the ratio threshold, determine that the shooting state of the shooting device is an approximately stationary state.
15. The motion object detection method according to claim 14, wherein The method further includes: determining the second speed based on the shooting state of the shooting device, including: If the shooting state is the approximately stationary state and the fourth speed of each target region is less than the product of a preset third value and the third speed of each target region, determine the second speed according to the fourth speed; or, if the shooting state is the approximately stationary state and the fourth speed of each target region is greater than or equal to the product of the third value and the third speed of each target region, determine the second speed according to the third speed; or, If the shooting state is the motion state and the fourth speed of each target region is less than the product of a preset fourth value and the third speed of each target region, determine the second speed according to the fourth speed; or, if the shooting state is the motion state and the fourth speed of each target region is greater than or equal to the product of the fourth value and the third speed of each target region, determine the second speed according to the third speed.
16. The motion object detection method according to claim 1, wherein The method further includes: performing denoising processing on the target region, including: Sort all target regions in descending order of the area of the target region, and select a preset number of target regions at the front from the sorted sequence, and determine the noise regions from the selected target regions.
17. The method for detecting a moving object according to claim 1 or 16, wherein Determining the noise regions in the target region includes: Determine the corresponding region of each target region in the first image frame or the second image frame, obtain the traditional difference map between the two image frames of the corresponding region, and each pixel point in the traditional difference map corresponds to a difference value; Determine the average difference value of the difference values of all pixel points in the traditional difference map corresponding to each target region; Determine the difference threshold corresponding to each target region according to the number of sub-regions in each target region; Determine the target regions corresponding to the average difference values less than or equal to the corresponding difference threshold as noise regions.
18. The method for detecting a moving object according to claim 17, wherein Determine the difference threshold corresponding to each target region according to the following formula: where threshold represents the difference threshold, num represents the number of sub-regions in each target region, and α, λ, β, γ, k represent preset parameters.
19. A terminal device, characterized in that, The terminal device includes a memory and a processor: The memory is used to store program instructions; The processor is used to read and execute the program instructions stored in the memory. When the program instructions are executed by the processor, the terminal device executes the moving object detection method according to any one of claims 1 to 18.
20. A computer storage medium, characterized in that, The computer storage medium stores program instructions, which, when running on a terminal device, cause the processor of the terminal device to execute the moving object detection method according to any one of claims 1 to 18.
Citation Information
Patent Citations
Method and device for detecting displacement of motion image as well as optical mouse
CN102243537A
Vehicle moving target detection method, device and system
CN111351474A
Moving target detection method, device and equipment and medium
CN111882583A
Speed information acquisition method and device, equipment and medium
CN113450579A
Video denoising method and device based on optical flow motion detection, and computer storage medium
CN115967777A