Motion detection method based on low-power-consumption embedded device

Through the intermittent working mode and dynamic sound pressure threshold combined with image frame sequence analysis, the balance problem of power consumption and accuracy of low-power embedded devices in motion detection is solved, and low-power consumption and high-reliability motion object detection is achieved, suitable for intelligent security and IoT perception scenarios.

CN120431129APending Publication Date: 2025-08-05翟勇慧

Patent Information

Application Number
CN202510540743.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The motion detection methods of existing low-power embedded devices are difficult to balance between power consumption control and detection accuracy, and the prior art often adopts high-performance processors or fixed threshold detection and are prone to false alarms, and stable motion detection cannot be achieved under low power consumption conditions.

Method used

The intermittent working mode is used to collect sound wave signals, combine dynamic sound pressure threshold and image frame sequence analysis, and the intermittent acquisition of the acoustic module and accurate awakening of the visual unit, combined with multimodal feature fusion and resource hierarchical scheduling, dynamic environment perception and efficient motion object detection are achieved.

Benefits of technology

High-precision motion object detection is achieved under low power consumption conditions, reducing false alarm rate, improving the battery life of the equipment and adaptability to complex scenarios, forming a closed-loop technical link, and providing low-power and highly reliable motion detection solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431129A_ABST
    Figure CN120431129A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of target detection, and discloses a motion detection method based on low-power-consumption embedded equipment. Comprises: acquiring an original sound wave signal and calculating an environment sound pressure level; when the environment sound pressure level is larger than or equal to a sound pressure threshold value for N times continuously, image acquisition is triggered, and an image frame sequence is generated; performing motion target analysis on the image frame sequence to obtain motion parameters of the object; judging whether an effective moving target exists or not according to the motion parameters, and sending the motion parameters to a detection terminal; and efficient balance among detection precision, power consumption control and scene adaptability is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and more particularly to a motion detection method based on a low-power embedded device. Background Art

[0002] With the rapid development of the Internet of Things, smart homes, smart wearables and other fields, the application of low-power embedded devices is becoming more and more extensive. Such devices usually need to operate for a long time with limited battery capacity, so the power consumption control requirements are extremely high. For example, devices such as door and window sensors and smart watches in smart homes all need to achieve stable motion detection functions under low power conditions to meet users' requirements for device endurance and functionality.

[0003] Patent application publication number CN103002195A discloses a motion detection method and a motion detection device. The motion detection method includes: capturing a current image; generating a current brightness image based on the current image; performing a differential operation on the current brightness image and the background brightness image to generate a foreground image; generating a foreground binary image based on the foreground image and sensitivity; and updating the background brightness image based on the update frequency and sensitivity.

[0004] However, the above-mentioned motion detection method has high performance requirements for the device and requires a high-performance processor to process large amounts of image data. It cannot be implemented on embedded devices with low processor performance. In order to improve the detection accuracy of motion detection, the existing technology generally adopts a "continuous wake-up" mode (such as the camera working all the time), resulting in a large average power consumption of the device. Simply reducing power consumption (such as reducing the sampling rate) will significantly sacrifice performance, and periodic sleep is prone to missing short movements. Fixed threshold detection (such as PIR sensors) is prone to false alarms due to temperature changes, environmental vibrations or the device's own movement. In addition, the existing technology mostly adopts independent processing of acoustic and visual signals, lacking a joint decision-making mechanism for the two.

[0005] In view of this, the present invention proposes a motion detection method based on a low-power embedded device to solve the above problem. Summary of the Invention

[0006] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solution: a motion detection method based on a low-power embedded device, comprising:

[0007] Step S1: collecting original sound wave signals in a preset intermittent working mode, calculating the sound pressure of the original sound wave signals, and obtaining the ambient sound pressure level;

[0008] Step S2: Collect the ambient sound pressure level within a fixed historical time period, record it as the historical ambient sound pressure level, construct a sound pressure threshold function based on the current ambient sound pressure level and the historical ambient sound pressure level, and thus obtain the sound pressure threshold; when the ambient sound pressure level is greater than or equal to the sound pressure threshold for N consecutive times, obtain an ambient image to generate an image frame sequence;

[0009] Step S3: performing motion target analysis on the image frame sequence to generate a foreground mask image; performing motion parameter calculation on the image frame sequence based on the foreground mask image to obtain the motion parameters of the object;

[0010] Step S4: Determine whether there is a valid moving target based on the motion parameters of the object. If it is determined that there is a valid moving target, send the motion parameters to the motion detection terminal.

[0011] Furthermore, the low-power embedded device includes an audio acquisition module, an image acquisition module, an image display module and a main control processor;

[0012] When the ambient sound pressure level is not greater than or equal to the sound pressure threshold for N consecutive times, the image acquisition module and the image display module are in a dormant state, and the main control processor is in a preset low power consumption mode;

[0013] When the ambient sound pressure level is greater than or equal to the sound pressure threshold for N consecutive times, the main control processor switches to a preset high power consumption mode and wakes up the image acquisition module and the image display module.

[0014] Furthermore, the working cycle of the intermittent working mode is composed of an acquisition window and a sleep window alternatingly, and each working cycle includes i acquisition windows and sleep windows;

[0015] In the last sleep window of the working cycle, the main control processor normalizes the original sound wave signal collected in the working cycle to generate a standardized sound pressure sampling sequence, and performs sound pressure calculation on the standardized sound pressure sampling sequence to obtain the ambient sound pressure level.

[0016] Furthermore, the moving target analysis includes:

[0017] Performing color space conversion on the image frame sequence and extracting brightness channel data of the color space, and generating a reference image based on the brightness channel data;

[0018] Downsampling the reference image to obtain a luminance image, dividing the luminance image into luminance windows of a preset size, and calculating the mean of the luminance windows; generating a luminance template matrix based on the mean of the luminance windows;

[0019] The reference image is divided into macroblocks of a preset size, the texture feature vectors of the macroblocks are extracted using an improved local binary pattern operator, and a three-dimensional texture coding template matrix is generated based on the texture feature vectors of the macroblocks;

[0020] Perform differential operation on the image frame based on the brightness template matrix and the three-dimensional texture coding template matrix to generate an initial foreground difference map;

[0021] The initial foreground difference map is binarized and segmented to generate a binary mask image. The noise points in the binary mask image are removed using the connected component labeling algorithm to generate a foreground mask image.

[0022] Furthermore, the method of extracting the texture feature vector of the macroblock by using the improved local binary pattern operator includes:

[0023] Construct a comparison window and use the brightness value of the central pixel of the comparison window as the comparison threshold. Compare the brightness values of all pixels in the comparison window with the comparison threshold. Pixels with brightness values greater than or equal to the comparison threshold are marked as 1, and pixels with brightness values less than the comparison threshold are marked as 0, thereby obtaining a binary sequence.

[0024] Performing a cyclic shift operation on the binary sequence to generate multiple shift variants, and selecting the minimum value among the decimal values corresponding to the shift variants as the encoding value of the macroblock;

[0025] The coding values of all pixels in the macroblock are statistically analyzed according to a predefined interval to obtain a histogram statistical result, and a texture feature vector is generated based on the statistical result.

[0026] Furthermore, the process of generating a binary mask image includes:

[0027] The initial foreground difference map is divided into statistical windows of a preset size, and the brightness mean and brightness standard deviation of the statistical windows are calculated based on the brightness values of the pixels in the statistical windows;

[0028] A dynamic threshold function is constructed based on the brightness mean and brightness standard deviation of the statistical window, and the dynamic threshold corresponding to the statistical window is obtained using the dynamic threshold function;

[0029] The initial foreground difference map is segmented based on a dynamic threshold to generate a binary mask image.

[0030] Furthermore, the method of obtaining the motion parameters of the moving object includes:

[0031] Construct a three-level Gaussian pyramid and use it to downsample the image frame sequence layer by layer to obtain the corresponding LO level image, L1 level image and L2 level image, and mark the effective motion area in each level image with the foreground mask image;

[0032] Based on the effective motion area between adjacent image frames in the L2 layer image, the displacement of the effective motion area of the adjacent image frames is calculated by the gradient operator and the optical flow algorithm to obtain the initial displacement vector;

[0033] Based on the effective motion area between adjacent image frames in the L1 layer image, the initial displacement vector is iteratively optimized by the Newton iteration method to obtain the optimized displacement vector;

[0034] Based on the effective motion area between adjacent image frames in the L0 level image, the optimized displacement vector is bilinearly interpolated by bilinear interpolation to obtain the displacement vectors Δx and Δy;

[0035] The direction angle and speed of the object's movement are calculated based on the displacement vectors Δx and Δy; the object's movement direction is determined based on the rate of change of the effective motion area in the foreground mask image; the movement direction, direction angle and speed together constitute the motion parameters.

[0036] The technical effects and advantages of the present invention are as follows:

[0037] This embodiment achieves an efficient balance between detection accuracy, power consumption control and scene adaptability through dynamic environment perception, multimodal feature fusion and resource hierarchical scheduling mechanism; through intermittent acquisition of acoustic modules and dynamic threshold analysis technology, it tracks the background noise baseline in real time and builds adaptive trigger logic, suppressing instantaneous interference while accurately identifying the acoustic characteristics of continuously moving targets, waking up the visual unit only when the conditions are met, and significantly reducing the energy consumption of static monitoring; the visual processing layer integrates the dual detection model of brightness difference analysis and texture feature matching, combines local adaptive threshold segmentation and morphological optimization technology, eliminates illumination artifacts and image noise interference, extracts the complete moving target outline, and constructs a multi-scale pyramid parsing framework through the improved optical flow method to achieve hierarchical iterative calculation of sub-pixel displacement vectors. The system can calculate the motion direction, speed and other parameters with both accuracy and efficiency. The embedded resource management adopts a modular decoupling design, compresses the computing load through downsampling processing, parallel window statistics and binary texture coding, and implements differentiated feature management for static background and dynamic foreground in combination with the regional dynamic template update strategy, thereby improving the robustness of the algorithm to changes in scene structure. The system forms a closed-loop technical chain from environmental perception to motion analysis through three core mechanisms: precise control of acoustic triggering, optimization of visual analysis levels and coordinated scheduling of software and hardware. It effectively solves the technical bottlenecks of high false alarm rate and weak endurance of traditional solutions, provides low-power and highly reliable motion detection solutions for scenarios such as smart security and Internet of Things perception, and promotes the development of embedded intelligent devices towards high efficiency and practicality. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 A schematic diagram of a motion detection method based on a low-power embedded device according to the present invention;

[0039] Figure 2 Schematic diagram of the foreground mask image of the present invention. DETAILED DESCRIPTION

[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0041] Example 1

[0042] See also Figure 1 As shown, the motion detection method based on a low-power embedded device described in this embodiment includes:

[0043] Step S1: collecting original sound wave signals in a preset intermittent working mode, calculating the sound pressure of the original sound wave signals, and obtaining the ambient sound pressure level;

[0044] Step S2: Collect the ambient sound pressure level within a fixed historical time period, record it as the historical ambient sound pressure level, construct a sound pressure threshold function based on the current ambient sound pressure level and the historical ambient sound pressure level, and thus obtain the sound pressure threshold; when the ambient sound pressure level is greater than or equal to the sound pressure threshold for N consecutive times, obtain an ambient image to generate an image frame sequence;

[0045] Step S3: performing motion target analysis on the image frame sequence to generate a foreground mask image; performing motion parameter calculation on the image frame sequence based on the foreground mask image to obtain the motion parameters of the object;

[0046] Step S4: Determine whether there is a valid moving target based on the motion parameters of the object. If it is determined that there is a valid moving target, send the motion parameters to the motion detection terminal.

[0047] The low-power embedded device includes an audio acquisition module, an image acquisition module, an image display module and a main control processor; when the ambient sound pressure level is not greater than or equal to a sound pressure threshold for N consecutive times, the image acquisition module and the image display module of the embedded device are in a dormant state to reduce the overall power consumption of the device and reduce unnecessary resource consumption; the main control processor is in a preset low-power mode to calculate the ambient sound pressure level; when the ambient sound pressure is greater than or equal to the sound pressure threshold for N consecutive times, the image acquisition module and the image display module are awakened, and the main control processor switches to a preset high-power mode to process image frames.

[0048] Specifically, the detailed implementation steps of step S1 include:

[0049] Based on a preset intermittent working mode, the working cycle of the audio acquisition module is divided into an alternating collection window and a dormant window, and one working cycle includes i collection windows and dormant windows;

[0050] In the acquisition window, the original sound wave signal in the environment is collected through the audio acquisition module;

[0051] In the last sleep window of a working cycle, the main control processor normalizes the amplitude of all original sound wave signals in the working cycle to generate a standardized sound pressure sampling sequence; calculates the mean sound pressure of the working cycle, and performs logarithmic conversion on the mean sound pressure to obtain the ambient sound pressure level of the current working cycle;

[0052] The working cycle of the preset intermittent working mode consists of an acquisition window and a sleep window; the duration of a working cycle is 200ms, of which the duration of the acquisition window accounts for 5% to 10% of the entire working cycle; within the acquisition window of the audio acquisition module, the audio acquisition module samples the ambient sound through the microphone to obtain the original sound wave signal of the environment; within the sleep window of the audio acquisition module, the main control processor normalizes the amplitude of the original sound wave signal in a low-power state to obtain a standardized sound pressure sampling sequence; calculates the root mean square value of the standardized sound pressure sampling sequence to obtain the sound pressure mean, unit: Pa; based on the preset sound pressure parameter (20μPa), the obtained sound pressure mean is logarithmically converted to obtain the ambient sound pressure level of the working cycle.

[0053] The above steps significantly reduce the operating power consumption of embedded devices and improve the reliability of environmental sound pressure monitoring through the coordinated optimization of intermittent working mode and efficient data processing strategy. On the one hand, this process greatly reduces the continuous energy consumption of audio acquisition and processing units through periodic module sleep. On the other hand, it uses signal amplitude normalization to eliminate hardware fluctuation interference and combines root mean square calculation to suppress transient noise, ensuring that the sound pressure value stably represents the environmental background sound energy level and providing accurate input for the dynamic threshold trigger mechanism. At the same time, the main control processor reuses idle computing power in the sleep window to complete lightweight data processing, avoiding resource conflicts and frequent wake-up of high-power computing units. This deep integration of hardware scheduling and algorithm optimization not only ensures extremely low power operation of the device in a static environment, but also reduces the probability of false wake-up of the image module through anti-interference processing, laying a high-efficiency and high-reliability acoustic perception foundation for subsequent moving target detection.

[0054] Specifically, the method of obtaining the dynamic threshold in step S2 includes:

[0055] An adaptive threshold function is constructed based on the current ambient sound pressure level and the historical ambient sound pressure level, and the sound pressure threshold is obtained using the adaptive threshold function. The expression of the adaptive threshold function is: T(t) = u(t) + k·σ(t); where T(t) is the sound pressure threshold (unit: decibel) calculated for the current working cycle, which is used to determine whether to trigger sound wake-up; k is the sensitivity coefficient, which is used to control the offset multiple of the threshold relative to the baseline noise; t is the time index, which represents the timing of the working cycle; μ(t) is the exponential moving mean of the background sound pressure (unit: decibel), which represents the baseline level of the current ambient noise; the function expression of μ(t) is: μ(t) = γ·μ(t-1) + (1-γ)·E(t); where E(t) is the ambient sound pressure level of the current working cycle; γ is a smoothing factor used to control the weight of the historical data mean, with a value range of 0.9-0.99. The larger the γ value, the stronger the ability to resist transient interference. Here, it is 0.96; μ(t-1) is the exponential moving mean of the background sound pressure of the previous working cycle of the current working cycle, which is used to retain historical noise information and smooth the sudden change of the current sound pressure level; σ(t) represents the exponential moving standard deviation of the background sound pressure (unit: decibel), which is used to represent the fluctuation intensity of the ambient noise; the functional expression of σ(t) is Where γ is the smoothing factor, which is 0.96; σ 2 (t-1) is the exponential moving variance of the previous working cycle of the current working cycle (unit: decibel 2 ); μ(t) is the exponential moving average of the background sound pressure in the previous working cycle; E(t) is the ambient sound pressure level in the current working cycle; (μ(t)-E(t)) 2 It represents the square deviation between the ambient sound pressure level and the exponential moving mean of the background sound pressure during the current working cycle, reflecting the instantaneous fluctuation intensity of the current sound pressure level.

[0056] If the ambient sound pressure level is not greater than or equal to the sound pressure threshold for N consecutive working cycles, it means that there is no moving object at this time, and the image acquisition module and image display module of the embedded device continue to remain in the dormant state;

[0057] When the ambient sound pressure level exceeds the calculated sound pressure threshold for N consecutive working cycles, the main control processor sends a wake-up signal to the image display module and the image acquisition module, causing them to exit sleep mode. Subsequently, the image acquisition module captures the ambient image based on a preset scanning mode (1600×1200@15fps) to obtain an image frame sequence.

[0058] The above process dynamically adjusts the sound pressure threshold through an adaptive threshold function, achieving dynamic perception of environmental background sounds and precise wake-up control of the image acquisition module; the adaptive threshold function uses exponential moving mean weighted calculation and exponential moving standard deviation to analyze the background noise baseline level in real time, and constructs a sound pressure threshold calculation mechanism through statistical analysis of the noise fluctuation intensity; this mechanism can not only effectively suppress short-term interference such as wind noise and electrical noise, but also accurately capture the sounds produced by continuously moving targets such as footsteps, forming a sound pressure threshold that can be adaptively adjusted with the environmental noise, so that the system can adapt to complex and changing acoustic environments without human intervention; this design not only improves the reliability of detection, but also optimizes the battery life of the equipment, achieving a good balance between performance and energy consumption.

[0059] In addition, by adopting a multi-cycle continuous trigger verification mechanism, the probability of false alarm is greatly reduced while ensuring the sensitivity of moving target detection; in terms of algorithm design, the system adopts a lightweight strategy, enabling the main control processor to complete sound pressure threshold calculation and status evaluation in a low-power state; combined with a modular wake-up strategy, the system implements hierarchical power consumption management of the image acquisition unit, ensuring extremely low energy consumption operation in a static environment.

[0060] Specifically, the detailed implementation steps of step S3 include:

[0061] The main control processor converts the image frame into the HSV color space and extracts the brightness value (0-255) of the brightness channel to generate a reference image, while recording the time parameters of the image frame;

[0062] Downsample the reference image to generate a luminance image with a resolution of 800×600, ensuring geometric center alignment;

[0063] The luminance image is divided into luminance windows of size 32×24 pixels. The horizontal step size of adjacent luminance windows is 16 pixels (50% overlap) and the vertical step size is 12 pixels (50% overlap). To ensure that adjacent luminance windows overlap by 50% and avoid edge information loss, 2401 luminance windows are obtained (49 horizontally and 49 vertically).

[0064] Calculate the mean and standard deviation of the brightness window, and round the mean to the nearest integer. The mean of all brightness windows constitutes the brightness template matrix M L ;

[0065] Based on the preset standard deviation threshold, the brightness window is classified into stable area (standard deviation ≤ 15) and dynamic area (standard deviation > 15). A low-frequency update strategy (10 seconds / time) is adopted for the stable area, and a real-time update strategy is adopted for the dynamic area. The brightness mean update formula is μ new =(1-α)·μ old +α·μ current; where μ old is the historical statistical value of the brightness mean in the window of the previous frame; μ current is the real-time statistical value of the brightness mean in the brightness window of the current frame; α is the update factor, and its value range is (0.02, 0.4). When the brightness mean of the stable area is updated, α is 0.05, and when the brightness mean of the dynamic area is updated, α is 0.3;

[0066] For windows with brightness mutations exceeding 20% for three consecutive frames, the mean of the brightness window is recalculated and the brightness template matrix M is updated. L The main control processor calculates the statistics of all brightness windows in parallel and quantizes the mean into 8-bit integer storage to reduce memory usage.

[0067] The reference image is divided into 64×48 pixel units. If the image size is not an integer multiple, the boundary is expanded by using the mirror filling method.

[0068] A 3×3 neighborhood scan is performed on each macroblock (with symmetric padding at the boundaries) to obtain an 8-bit binary string. The binary string is then cyclically left-shifted to generate eight shift combinations (including the original sequence). The binary string is converted to decimal values and the minimum value is taken as the final LBP code. For example, if the original binary code is 11001001, after the cyclic shift operation, eight different binary codes are obtained, of which the smallest value is 00011100. Converting this to decimal is 28, which is the final LBP code value for the neighborhood.

[0069] The neighborhood scanning rule is as follows: taking the neighborhood center pixel as the benchmark, the brightness values of its eight neighboring pixels (located at the top, bottom, left, right, upper left, upper right, lower left, and lower right positions) are compared one by one with the brightness value of the center pixel; if the brightness value of the neighboring pixel is greater than or equal to the brightness value of the center pixel, the neighboring pixel is marked as 1; if the brightness value of the neighboring pixel is less than the brightness value of the center pixel, the neighboring pixel is marked as 0;

[0070] Perform 16-bin histogram statistics on the LBP coding values in each macroblock to generate a texture feature vector containing 16 bins; arrange the 16-dimensional feature vectors of each macroblock into an m×n×16 three-dimensional texture coding template matrix in raster scanning order (m and n are the number of vertical and horizontal macroblocks respectively);

[0071] It should be noted that the 16-bin histogram means that the LBP code value (0-255) is evenly divided into 16 intervals, and each interval spans 16. The 16-bin histogram is obtained by counting the number of LBP code values of the macroblock in each bin interval; among them, bin1: 0-15; bin2: 16-31; and so on to obtain the interval range of each bin;

[0072] For example, the LBP coding values of the macroblock are: 5, 12, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250; according to the above interval division, the number of LBP coding values in each interval is counted, and the following results are obtained: bin1 (0-15): 2 (5, 12); bin2 (16-31): 2 (20, 30); bin3 (32-47): 1 (40); bin4 (48-63): 1 (50); bin5 (64-79): 1 (60); bin6 (80-95): 1 (80); bin7 (96-111): 1 (1 00); bin8 (112-127): 1 (120); bin9 (128-143): 1 (130); bin10 (144-159): 1 (140); bin11 (160-175): 1 (160); bin12 (176-191): 1 (170); bin13 (192-207): 1 (19 0); bin14 (208-223): 1 (210); bin15 (224-239): 1 (230); bin16 (240-255): 2 (240, 250); then, the texture feature vector composed of the statistical results of these 16 bins is: [2, 2, 1, 1, 1, 1, 1, 1, 1, 1, 1, 2].

[0073] The similarity of adjacent macroblocks is calculated by the similarity function, and the similarity function expression is: The value range of D is [0, +∞), and the smaller the value of D is, the higher the similarity of the textures of the two macroblocks is; f m and f n Represents the LBP histogram feature vectors of two adjacent macroblocks, each vector contains 16 bins (i.e., i = 1, 2, ..., 16); f m (i) represents the frequency statistics of macroblock m in the i-th bin (e.g. bin 5 corresponds to the number of pixels in the LBP coding range 32-39); f n (i) represents the frequency statistics of macroblock n in the i-th bin; i represents the histogram bin index; for example, f m (5) = 120 means that the LBP codes of 120 pixels in macroblock m fall within bin5;

[0074] Traverse all adjacent macroblocks. If their similarity D is less than 3.0, merge them into the same texture region. Update the boundary coordinates of the merged region and recalculate the merged texture feature vector (take the mean of the feature vectors of each macroblock).

[0075] When the similarity D of a macroblock is greater than 5.0 (i.e., the similarity with the adjacent macroblock is less than 60%), it is marked as a "mutation block" and the histogram data of the "mutation block" is cleared. Its texture feature vector is recalculated based on the LBP coding of the last 5 frames; the similarity D of the mutation block and the adjacent area is recalculated; if the similarity D of the macroblock is less than 3.0, the two macroblocks are re-merged; otherwise, they are split into independent areas; when the mutation block is split or merged, the area boundaries in the 3D texture coding template matrix need to be updated, and the feature vectors of other areas remain unchanged;

[0076] Every 30 seconds, all macroblocks are updated to regenerate the 3D texture coding template matrix to adapt to structural changes in the scene (such as furniture displacement, door and window opening and closing).

[0077] Perform HSV color space conversion on the current image frame, extract the brightness value of the brightness channel, and form a brightness matrix based on the arrangement of image pixels.

[0078] The brightness difference intensity of the current frame is calculated based on the brightness matrix and the brightness template matrix. The calculation formula is: D L =|I L -M L |; Among them, I L Represents the brightness matrix of the current frame, M L Represents the brightness template matrix;

[0079] The current image frame is encoded using the improved LBP operator to obtain the encoding matrix, and the Hamming distance D between the encoding matrix and the three-dimensional texture encoding template matrix is calculated. T ;

[0080] The initial foreground difference map F0 is generated based on the brightness difference intensity and Hamming distance calculated above; the function expression is: F0 = k1·D L +k2·D T ; Where k1 and k2 are weight coefficients, and their sum is 1; the brightness value range of the initial foreground difference map is 0-255;

[0081] The initial foreground difference map F0 is divided into a statistical window of size 16×12 pixels, and the mean μ of the pixel values in the statistical window is calculated. local and standard deviation σ local ;

[0082] Mean μ based on statistical window local and standard deviation σ localConstruct a dynamic threshold function, the function expression is T(x,y)=μ local +λ·σ local ; where λ is a constant, when σ local ≥10, the value is 1.8. local When <10, the value is 1.2;

[0083] The dynamic threshold function is used to obtain the corresponding dynamic threshold, and the initial foreground difference map F0 is segmented based on the dynamic threshold to generate a binary mask image M b ; The segmentation rules are:

[0084] When the value of a point (x, y) in the initial foreground difference map F0(x, y) is less than or equal to T(x, y), the binary mask image M b The value of the point in (x, y) is set to 0;

[0085] On the contrary, when the value of a point (x, y) in the initial foreground difference map F0(x, y) is greater than T(x, y), the binary mask image M b The value of the point in (x,y) is set to 255.

[0086] Binarized mask image M b The specific process of morphological optimization and connected domain filtering is as follows:

[0087] A morphological opening operation is performed on the binary mask image, first eroding and then dilating the binary mask image using a 3×3 circular structuring element. During the erosion phase, isolated noise points (such as sensor thermal noise or small dynamic interference) with an area smaller than 3×3 pixels in the binary mask image are eliminated. During the dilation phase, the edges of moving targets lost due to erosion in the binary mask image are restored to ensure the integrity of the moving target contour.

[0088] See also Figure 2 As shown in the figure, the binary mask image after morphological opening operation is traversed by the connected domain labeling algorithm, and the area of the circumscribed rectangle of each connected domain is counted; the fragmented areas with an area less than 50 pixels (such as flying insects, dust disturbances, etc.) are eliminated, and the connected domains with an area greater than 50 pixels and compact shape are retained and marked as valid motion areas, thereby obtaining the foreground mask image M;

[0089] Construct a three-level Gaussian pyramid; the Gaussian pyramid levels are as follows:

[0090] L0 level: original resolution (e.g. 1600×1200), used for final accuracy optimization;

[0091] L1 layer: 1 / 2 resolution (800×600), used for intermediate iterations;

[0092] L2 level: 1 / 4 resolution (400×300), used for initial coarse estimation.

[0093] The standard deviation of Gaussian filtering in each layer is 1.0, 1.5, and 2.0 respectively;

[0094] Use the three-level Gaussian pyramid to analyze the current frame I t and the next frame I of the current frame t+1 Sampling is performed separately to obtain the image I of the corresponding resolution level t L0 , I t L1 and I t L2 , and

[0095] The foreground mask images of the current frame and the next frame of the current frame are downsampled to each pyramid level at the same ratio to generate the corresponding M L0 、M L1 and M L2 Three levels of images.

[0096] Based on the foreground mask image M L2 The effective motion area marked in the L2 level image I t L2 and Mark the motion area in the image; calculate I by improving the optical flow equation t L2 and The displacement of the moving area between the moving objects is obtained by the initial displacement vector Δx of the moving object at the L2 level. L2 and Δy L2 ; The function expression of the improved optical flow equation is: Where; W is a 3×3 weight kernel, used to normalize the matrix; For (x,y) in I t L2 The gradient along the x and y directions in the image is calculated by the Sobel operator;

[0097] Through the foreground mask image M L1 The effective motion area marked in the L1 level image I t L1 and Mark the effective motion area in the image;

[0098] The initial displacement vector of the L2 layer is in the L1 layer image I t L1 and The displacement vector is iteratively optimized within the effective motion area in , thereby obtaining the optimized displacement vector Δx at the L1 level L1 and Δy L1 ; Δx L1 and Δy L1 The function expression is: Δx L1 =Δx L2 ×2+δx L1 ;Δy L1 =Δx L2 ×2+δy L1 ; Among them, δx L1 and δy L1 Represents the L1 layer image I t L1 and The residual displacement increment obtained by local iterative solution; the residual calculation formula is: in, For (x,y) in I t L1 The gradient along the x and y directions in the image is calculated by the Sobel operator; δx L1 and δy L1 Represents the residual displacement increment to be solved; the constraint condition is: the iteration stops when the maximum number of iterations reaches 4 or the residual is less than 0.1;

[0099] Through the foreground mask image M L0 The effective motion area marked in I t L0 and Mark the movement area;

[0100] The optimized displacement vector of the L1 level is scaled up to the L0 level according to the resolution ratio; t L0 and In the motion area, the displacement vector is bilinearly interpolated by bilinear interpolation to obtain displacement vectors Δx and Δy;

[0101] The direction angle θ and velocity v of the object are calculated based on the displacement vectors Δx and Δy. The calculation formula is: Wherein, Δt is the time interval between two adjacent frames;

[0102] The direction of motion is determined based on the rate of change of the effective motion region in the foreground mask image: a positive rate of change indicates that the object is approaching the embedded device; a negative rate of change indicates that the object is moving away from the embedded device.

[0103] The object's movement direction, direction angle and movement speed together constitute the movement parameters.

[0104] The above method achieves high-precision moving target detection in complex scenes through multimodal feature fusion and hierarchical motion analysis strategy; this method combines brightness dynamic perception and texture feature analysis to construct a dual difference detection model, uses local adaptive threshold segmentation to effectively distinguish real moving targets from light interference, eliminates noise and extracts the complete moving area contour through morphological optimization; the Gaussian pyramid multi-scale displacement estimation technology based on the improved optical flow method maintains the accuracy of motion vector calculation while significantly reducing the computational complexity, and realizes real-time motion parameter analysis under the resource constraints of embedded devices; the regional dynamic update mechanism enhances the algorithm's adaptability to scene structure changes by classifying and processing the feature templates of static background and dynamic foreground, ensuring the stability of parameter output such as motion direction and speed; the multi-level verification mechanism effectively integrates the complementary advantages of acoustic triggering and visual detection, forming a reliable moving target discrimination capability under low power conditions.

[0105] The detailed implementation steps of step S4 include:

[0106] When the motion direction data in the motion parameters is not 0, it indicates that there is a valid moving target at this time; when either the direction angle or the motion speed is 0 and the motion direction data is 0, it indicates that there is no valid moving target at this time;

[0107] When there is no valid moving target, the delay counter is started; when the delay counter reaches the preset time, the embedded device is controlled to return to the sleep state;

[0108] When a valid moving target is detected, the motion parameters are sent to the motion detection terminal; when the valid moving target disappears or is in a stationary state and the ambient sound pressure level is lower than the dynamic threshold, the delay counter is started; when the delay counter reaches the preset time, the embedded device is controlled to return to the sleep state.

[0109] This embodiment achieves an efficient balance between detection accuracy, power consumption control and scene adaptability through dynamic environment perception, multimodal feature fusion and resource hierarchical scheduling mechanism; through intermittent acquisition of acoustic modules and dynamic threshold analysis technology, it tracks the background noise baseline in real time and builds adaptive trigger logic, suppressing instantaneous interference while accurately identifying the acoustic characteristics of continuously moving targets, waking up the visual unit only when the conditions are met, and significantly reducing the energy consumption of static monitoring; the visual processing layer integrates the dual detection model of brightness difference analysis and texture feature matching, combines local adaptive threshold segmentation and morphological optimization technology, eliminates illumination artifacts and image noise interference, extracts the complete moving target outline, and constructs a multi-scale pyramid parsing framework through the improved optical flow method to achieve hierarchical iterative calculation of sub-pixel displacement vectors. The system can calculate the motion direction, speed and other parameters with both accuracy and efficiency. The embedded resource management adopts a modular decoupling design, compresses the computing load through downsampling processing, parallel window statistics and binary texture coding, and implements differentiated feature management for static background and dynamic foreground in combination with the regional dynamic template update strategy, thereby improving the robustness of the algorithm to changes in scene structure. The system forms a closed-loop technical chain from environmental perception to motion analysis through three core mechanisms: precise control of acoustic triggering, optimization of visual analysis levels and coordinated scheduling of software and hardware. It effectively solves the technical bottlenecks of high false alarm rate and weak endurance of traditional solutions, provides low-power and highly reliable motion detection solutions for scenarios such as smart security and Internet of Things perception, and promotes the development of embedded intelligent devices towards high efficiency and practicality.

[0110] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art will be able to modify the technical solutions described in the foregoing embodiments or to substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

[0111] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0112] In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0113] In the description of the present invention, unless otherwise specified, "plurality" means two or more.

[0114] In the description of the present invention, “several” means one or more, and “a large number” means two or more.

[0115] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0116] The formulas in this manual are all dimensionless and calculated using numerical values. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field based on actual conditions.

[0117] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

Claims

1. A motion detection method based on low-power embedded devices, characterized in that: include: Step S1: collecting original sound wave signals in a preset intermittent working mode, calculating the sound pressure of the original sound wave signals, and obtaining the ambient sound pressure level; Step S2: Collect the ambient sound pressure level within a fixed historical time period, record it as the historical ambient sound pressure level, construct a sound pressure threshold function based on the current ambient sound pressure level and the historical ambient sound pressure level, and thus obtain the sound pressure threshold; when the ambient sound pressure level is greater than or equal to the sound pressure threshold for N consecutive times, obtain an ambient image to generate an image frame sequence; Step S3: performing motion target analysis on the image frame sequence to generate a foreground mask image; Calculate the motion parameters of the image frame sequence based on the foreground mask image to obtain the motion parameters of the object; Step S4: Determine whether there is a valid moving target based on the motion parameters of the object. If it is determined that there is a valid moving target, send the motion parameters to the motion detection terminal.

2. The motion detection method based on low-power embedded devices according to claim 1, characterized in that: The low-power embedded device includes an audio acquisition module, an image acquisition module, an image display module and a main control processor; When the ambient sound pressure level is not greater than or equal to the sound pressure threshold for N consecutive times, the image acquisition module and the image display module are in a dormant state, and the main control processor is in a preset low power consumption mode; When the ambient sound pressure level is greater than or equal to the sound pressure threshold for N consecutive times, the main control processor switches to a preset high power consumption mode and wakes up the image acquisition module and the image display module.

3. The motion detection method based on low-power embedded devices according to claim 2, characterized in that: The working cycle of the intermittent working mode is composed of acquisition windows and sleep windows alternating, and each working cycle contains i acquisition windows and sleep windows; In the last sleep window of the working cycle, the main control processor normalizes the original sound wave signal collected in the working cycle to generate a standardized sound pressure sampling sequence, and performs sound pressure calculation on the standardized sound pressure sampling sequence to obtain the ambient sound pressure level.

4. The motion detection method based on low-power embedded devices according to claim 3, characterized in that: The moving target analysis includes: Performing color space conversion on the image frame sequence and extracting brightness channel data of the color space, and generating a reference image based on the brightness channel data; Downsampling the reference image to obtain a luminance image, dividing the luminance image into luminance windows of a preset size, and calculating the mean of the luminance windows; Generate a brightness template matrix based on the mean of the brightness window; The reference image is divided into macroblocks of a preset size, the texture feature vectors of the macroblocks are extracted using an improved local binary pattern operator, and a three-dimensional texture coding template matrix is generated based on the texture feature vectors of the macroblocks; Perform differential operation on the image frame based on the brightness template matrix and the three-dimensional texture coding template matrix to generate an initial foreground difference map; The initial foreground difference map is binarized and segmented to generate a binary mask image. The noise points in the binary mask image are removed using the connected component labeling algorithm to generate a foreground mask image.

5. The motion detection method based on low-power embedded devices according to claim 4, characterized in that: The method of extracting the texture feature vector of the macroblock by using the improved local binary pattern operator includes: Construct a comparison window and use the brightness value of the central pixel of the comparison window as the comparison threshold. Compare the brightness values of all pixels in the comparison window with the comparison threshold. Pixels with brightness values greater than or equal to the comparison threshold are marked as 1, and pixels with brightness values less than the comparison threshold are marked as 0, thereby obtaining a binary sequence. Performing a cyclic shift operation on the binary sequence to generate multiple shift variants, and selecting the minimum value among the decimal values corresponding to the shift variants as the encoding value of the macroblock; The coding values of all pixels in the macroblock are statistically analyzed according to a predefined interval to obtain a histogram statistical result, and a texture feature vector is generated based on the statistical result.

6. The motion detection method based on low-power embedded devices according to claim 5, characterized in that: The process of generating a binary mask image includes: The initial foreground difference map is divided into statistical windows of a preset size, and the brightness mean and brightness standard deviation of the statistical windows are calculated based on the brightness values of the pixels in the statistical windows; A dynamic threshold function is constructed based on the brightness mean and brightness standard deviation of the statistical window, and the dynamic threshold corresponding to the statistical window is obtained using the dynamic threshold function; The initial foreground difference map is segmented based on a dynamic threshold to generate a binary mask image.

7. The motion detection method based on low-power embedded devices according to claim 6, characterized in that: The method of obtaining the motion parameters of the moving object includes: Construct a three-level Gaussian pyramid and use it to downsample the image frame sequence layer by layer to obtain the corresponding LO level image, L1 level image and L2 level image, and mark the effective motion area in each level image with the foreground mask image; Based on the effective motion area between adjacent image frames in the L2 layer image, the displacement of the effective motion area of adjacent image frames is calculated by the gradient operator and the optical flow algorithm to obtain the initial displacement vector; In the effective motion region between adjacent image frames in the L1 level image, the initial displacement vector is iteratively optimized by the Newton iteration method to obtain the optimized displacement vector; In the effective motion area between adjacent image frames in the L0 level image, the optimized displacement vector is bilinearly interpolated by bilinear interpolation to obtain displacement vectors Δx and Δy; The direction angle and speed of the object's movement are calculated based on the displacement vectors Δx and Δy; the object's movement direction is determined based on the rate of change of the effective motion area in the foreground mask image; the movement direction, direction angle and speed together constitute the motion parameters.

Citation Information

Patent Citations

  • Motion detecting method and motion detecting device

    CN103002195A

Cited By

  • Dynamic texture atlas recombination and compression method and system based on airspace sparsity

    CN121810824A