Low-power monitoring wake-up method, device and equipment based on humanoid recognition

By downsampling and graying processing of low-power monitoring equipment, combined with background differential and multi-scale feature extraction, an adaptive monitoring strategy is generated, which solves the problems of low energy utilization efficiency and insufficient recognition accuracy of traditional low-power monitoring equipment in humanoid target recognition, and achieves efficient and accurate monitoring effects.

CN118982847BActive Publication Date: 2025-07-18SHENZHEN ANKED SHITONG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411220085.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-02
Publication Date
2025-07-18
Estimated Expiration
2044-09-02

AI Technical Summary

Technical Problem

Traditional low-power monitoring equipment is difficult to dynamically adjust the working status according to actual scenario requirements, resulting in low energy utilization efficiency and insufficient accuracy and real-time performance in humanoid target recognition, which is prone to false alarms and missed alarms.

Method used

By downsampling and grayscale processing of the original image sequence, combined with background differential technology, humanoid features are extracted and initial humanoid recognition model is trained, multi-scale feature extraction and classification are generated, multi-mode adaptive monitoring strategies are dynamically adjusted, and the model is optimized through incremental learning.

Benefits of technology

It improves the energy utilization efficiency of monitoring equipment, extends the battery life, improves the accuracy and real-timeness of humanoid target recognition, and provides a more reliable and intelligent security solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118982847B_ABST
    Figure CN118982847B_ABST
Patent Text Reader

Abstract

The present invention relates to a low-power monitoring wake-up method, device and equipment based on human form recognition. The method includes: performing downsampling and background difference on the original image sequence collected by the low-power monitoring device to obtain a preliminary target mask; performing human form feature extraction to obtain a pseudo label of the human form area; training to obtain an initial human form recognition model and performing multi-scale feature extraction to obtain a target feature vector; performing classification to obtain a preliminary human form recognition result and optimizing it to obtain a human form target tracking sequence; calculating a scene activity index, generating a multi-mode adaptive monitoring strategy, and adjusting the low-power working scheme; performing incremental learning to obtain a target human form recognition model and generating a multi-level wake-up control strategy. The implementation of the present invention improves the energy utilization efficiency of the monitoring device, extends the battery service life, and can also improve the accuracy and real-time performance of monitoring, providing a more reliable and intelligent security solution for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human form recognition, and particularly to a low-power monitoring wake-up method, device, and equipment based on human form recognition. Background Art

[0002] With the rapid development of Internet of Things and smart home technologies, low-power monitoring devices are widely used in fields such as security and home monitoring. However, traditional low-power monitoring devices often adopt a fixed working mode and are difficult to dynamically adjust the working state according to actual scenario requirements, resulting in low energy utilization efficiency. At the same time, the accuracy and real-time performance of these devices in human target recognition are also insufficient, and false alarms and missed detections often occur, affecting the monitoring effect.

[0003] In addition, most of the existing monitoring wake-up methods rely on simple motion detection or fixed threshold triggering, and are unable to effectively distinguish human targets from other moving objects, and are easily awakened frequently by environmental interference, increasing unnecessary energy consumption. In complex and changeable actual monitoring scenarios, how to achieve low-power operation while ensuring recognition accuracy has become an urgent technical problem to be solved. Summary of the Invention

[0004] The objective of the present invention is to provide a low-power monitoring wake-up method, device, and equipment based on human form recognition, so as to improve the energy utilization efficiency of the monitoring device, extend the battery life, and also be able to improve the accuracy and real-time performance of monitoring, providing a more reliable and intelligent security solution for users.

[0005] To achieve the above objective, the present invention provides a low-power monitoring wake-up method based on human form recognition, including the following steps:

[0006] Downsample and grayscale the original image sequence collected by the low-power monitoring device to obtain a preprocessed image and perform background difference to obtain a preliminary target mask;

[0007] Extract human form features from the preliminary target mask to obtain candidate human form regions and calculate shape descriptors to obtain human form region pseudo-labels;

[0008] Train an initial human form recognition model based on the human form region pseudo-labels and perform multi-scale feature extraction to obtain a local spatial depth feature map and a target feature vector;

[0009] Classify and perform weighted voting on the target feature vector to obtain a preliminary human form recognition result, and optimize the recognition results of consecutive frames to obtain a human target tracking sequence;

[0010] Calculate the scene activity index based on the humanoid target tracking sequence, generate a multi-mode adaptive monitoring strategy, and dynamically adjust the low-power working scheme of the low-power monitoring device based on the multi-mode adaptive monitoring strategy;

[0011] Perform incremental learning and model update according to the humanoid target tracking sequence to obtain a target humanoid recognition model, and generate a multi-level wake-up control strategy based on the target humanoid recognition model and the low-power working scheme.

[0012] The present invention also provides a low-power monitoring wake-up device based on humanoid recognition, including:

[0013] A preprocessing module, configured to downsample and grayscale the original image sequence collected by the low-power monitoring device, obtain a preprocessed image and perform background difference to obtain a preliminary target mask;

[0014] A calculation module, configured to extract humanoid features from the preliminary target mask, obtain a candidate humanoid region and perform shape descriptor calculation to obtain a humanoid region pseudo-label;

[0015] A training module, configured to train an initial humanoid recognition model based on the humanoid region pseudo-label and perform multi-scale feature extraction to obtain a local spatial depth feature map and a target feature vector;

[0016] A classification module, configured to classify and perform weighted voting on the target feature vector to obtain a preliminary humanoid recognition result, and optimize the recognition results of consecutive frames to obtain a humanoid target tracking sequence;

[0017] An adjustment module, configured to calculate the scene activity index based on the humanoid target tracking sequence, generate a multi-mode adaptive monitoring strategy, and dynamically adjust the low-power working scheme of the low-power monitoring device based on the multi-mode adaptive monitoring strategy;

[0018] A generation module, configured to perform incremental learning and model update according to the humanoid target tracking sequence to obtain a target humanoid recognition model, and generate a multi-level wake-up control strategy based on the target humanoid recognition model and the low-power working scheme.

[0019] The present invention also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0020] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0021] In summary, the technical solution provided by the present invention preprocesses the original image sequence through downsampling and grayscale conversion, and combines background difference technology to effectively reduce the amount of data processing and computational complexity, providing a high-quality preliminary target mask for subsequent humanoid recognition. By adopting multi-scale feature extraction and spatial pyramid pooling modules, combined with discrete wavelet transform, it can capture multi-level feature information of humanoid targets, improving the robustness and generalization ability of the recognition model. The introduction of a weighted voting mechanism and spatio-temporal consistency analysis effectively reduces the error of single-frame recognition and improves the accuracy and stability of humanoid target tracking. Based on multi-dimensional information such as scene activity index, lighting conditions, and device power, an adaptive monitoring strategy is generated to dynamically optimize the working mode of the monitoring device, maximizing the energy-saving potential while ensuring the monitoring effect. Through incremental learning and model update mechanisms, the humanoid recognition model can continuously adapt to new scenarios and target features, improving the long-term operating performance of the system. The design of a multi-level wake-up control strategy realizes a progressive wake-up process from low energy consumption to high performance, avoiding unnecessary full-power operation while ensuring fast response, and further optimizing the energy utilization efficiency. The present invention can achieve efficient and accurate humanoid recognition and low-power operation in complex and variable actual application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 FIG. is a schematic diagram of the steps of a low-power monitoring wake-up method based on humanoid recognition according to an embodiment of the present invention;

[0023] Figure 2 FIG. is a block diagram of the structure of a low-power monitoring wake-up device based on humanoid recognition according to an embodiment of the present invention;

[0024] Figure 3 FIG. is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.

[0025] The implementation, functional features, and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0027] Referring to Figure 1 , this embodiment provides a low-power monitoring wake-up method based on humanoid recognition, including the following steps:

[0028] S1. Downsample and grayscale the original image sequence collected by the low-power monitoring device, obtain a preprocessed image, and perform background difference to obtain a preliminary target mask;

[0029] Specifically, the resolution of the original image sequence collected by the low-power monitoring device is reduced. The downsampled image is obtained through downsampling, and the downsampled image is subjected to grayscale conversion to generate a grayscale image. The grayscale image is subjected to image smoothing processing to eliminate high-frequency noise and fine textures in the image, enhance the smoothness of the image, and obtain a preprocessed image. The preprocessed image is subjected to edge enhancement processing. By highlighting the edge information in the image, the contour clarity of the objects in the image is improved, and a preprocessed image sequence containing more obvious edge features is generated. Multiple frames of images are selected from the generated preprocessed image sequence, and the pixels of these images are averaged to generate a background reference image. The background reference image represents the static background part in the scene and is used to compare with the newly input preprocessed image. By performing a difference operation on the newly input preprocessed image and the background reference image, a difference image is obtained, revealing possible moving objects or changing areas in the image. The difference image is subjected to threshold processing. A first threshold is set, and the pixels in the difference image greater than the threshold are marked as the foreground area to generate a binary difference image containing possible foreground objects. The binary difference image is subjected to noise removal processing to eliminate isolated points and small areas in the image, thereby optimizing the moving area image. Edge detection is performed on the optimized moving area image. By identifying the contours of the objects in the image, a target contour image is obtained, and the target contour line segments in the image are extracted. The target contour line segments are compared with a preset humanoid contour template, the similarity score between the contour line segments and the template is calculated, and a humanoid feature description is generated. The humanoid feature description is compared with a second threshold. Only when the similarity score exceeds the second threshold will the corresponding area be recognized as a target area that may contain a humanoid, and finally a preliminary target mask is obtained through screening.

[0030] S2. Extract humanoid features from the preliminary target mask, obtain candidate humanoid regions, and calculate shape descriptors to obtain humanoid region pseudo-labels;

[0031] Specifically, the preliminary target mask is framed by a rectangle to define candidate rectangular regions that may contain human figures. According to a preset human body proportion range, the candidate rectangular regions are screened to exclude those regions that do not conform to the human body proportion, and a preliminary human figure region is obtained. Feature extraction is performed on the preliminary human figure region to extract its gradient features and generate a histogram of oriented gradients. The histogram of oriented gradients reflects the distribution of pixel gradient directions within the region. By analyzing the distribution features, the feature similarity is calculated to obtain a likelihood score that measures the possibility of the region being a human figure. This score is used to preliminarily judge the probability of a human figure appearing in the candidate region. Based on the human figure likelihood score, the preliminary human figure region is screened again to obtain a more refined and more likely human figure region. The overlapping regions of the refined human figure regions are merged to obtain a unified candidate human figure region. Shape description data is extracted from the candidate human figure region, and the obtained shape description data is used to represent the geometric shape features of the region. The shape description data is compressed in features to generate a compact shape descriptor. The compact shape descriptors are matched. The descriptors are matched with a preset standard shape template to obtain a shape matching result. According to the matching result, the candidate regions are sorted by similarity to identify the best-matching shape. The best-matching shape usually refers to the shape descriptor that is closest to the standard human figure template and represents the most likely part of the candidate region to be a human figure. Combining the best-matching shape with the previously calculated human figure likelihood score, a confidence score is comprehensively calculated to generate a confidence map of the entire human figure region. The confidence map reflects the likelihood of each region in the image being considered a human figure. Region expansion is performed on the confidence map of the human figure region to cover and capture the complete human figure region through the expansion operation, ensuring that all parts that may contain a human figure are taken into account. Feature description vectors are extracted from the complete human figure region. This vector contains all the important features related to the human figure within the region. By screening and analyzing the feature description vectors, it is determined which features meet the preset standards, and finally, the corresponding pseudo-labels of the human figure region are generated.

[0032] S3. An initial human figure recognition model is trained based on the pseudo-labels of the human figure region and multi-scale feature extraction is performed to obtain a local spatial depth feature map and a target feature vector;

[0033] Specifically, data augmentation is performed on the pseudo-labels of the humanoid region to expand the diversity of training data and obtain a more representative augmented training set. The augmented training set is normalized to ensure the consistency and stability of the input data, resulting in normalized training data. In terms of model construction, a lightweight convolutional neural network backbone consisting of 5 convolutional blocks is designed. Each convolutional block is composed of a convolutional layer, a batch normalization layer, and a ReLU activation function layer. This structure can effectively extract the local features of the image while keeping the model lightweight, suitable for the real-time processing requirements of low-power devices. Through this convolutional neural network backbone, feature extraction is performed on the input preprocessed image to obtain a preliminary feature map. A discrete wavelet transform layer is inserted between the 3rd and 4th convolutional blocks of the convolutional neural network backbone. By performing discrete wavelet transform on the feature map passing through the 3rd convolutional block, the feature map is decomposed into low-frequency approximation coefficients and high-frequency detail coefficients. The low-frequency approximation coefficients mainly contain the overall information of the image, while the high-frequency detail coefficients retain the edge and texture information of the image. By adaptively weighted fusion of these two types of coefficients, a multi-scale feature map is obtained. The multi-scale feature map is input into the 4th convolutional block to obtain an enhanced feature representation. A spatial pyramid pooling module is added after the last convolutional block of the convolutional neural network backbone. This module generates feature vectors of different scales by performing multi-scale pooling on the output feature map of the last convolutional block. By concatenating the feature vectors of different scales, a complete multi-scale feature representation is generated, and dimensionality reduction processing is performed on this multi-scale feature representation through 1x1 convolution operation to obtain a compressed feature vector. A fully connected layer network consisting of two fully connected layers and a Dropout layer is constructed, and the compressed feature vector is input into this fully connected layer network to generate the final prediction result. By using the cross-entropy loss function to calculate the error between the prediction result and the true label, a loss value is obtained. The parameters of the network are optimized through the backpropagation algorithm, and the network parameters are updated until the loss value converges or reaches the preset number of iterations, obtaining an initial humanoid recognition model. After completing the training of the initial humanoid recognition model, the new preprocessed image is input into the model, and after being processed by all convolutional blocks, discrete wavelet transform layers, and spatial pyramid pooling modules, a local spatial depth feature map is generated. Global average pooling operation is performed on the local spatial depth feature map to extract the target feature vector.

[0034] S4. Classify and perform weighted voting on the target feature vector to obtain a preliminary humanoid recognition result, and optimize the recognition results of consecutive frames to obtain a humanoid target tracking sequence;

[0035] Specifically, the target feature vector is normalized to adjust the value range of the feature vector to a standard range, eliminating the magnitude differences between different features. The normalized standardized feature vector is input into multiple pre-configured kernel extreme learning machines, and each kernel extreme learning machine generates a set of classification results. The classification results of different kernel extreme learning machines are weighted and averaged to obtain a comprehensive classification score. According to the preset classification threshold, the comprehensive classification score is binarized to obtain a preliminary humanoid recognition result, which marks the preliminary determination of the humanoid target in the current frame image by the monitoring system. Time series analysis is performed on the preliminary humanoid recognition results of multiple consecutive frames to evaluate the temporal consistency of the humanoid target in multiple consecutive frames, and a temporal consistency score is calculated. Based on the temporal consistency score, the single-frame recognition result is corrected to obtain a temporally smoothed recognition result, reducing errors caused by instantaneous changes or noise. Spatial connectivity analysis is performed on the temporally smoothed recognition result to evaluate the spatial consistency of the recognition result, and a spatial consistency score is calculated. By combining the spatial consistency score and the temporally smoothed recognition result, a spatio-temporally consistent humanoid region is obtained, representing the region in the surveillance video that is stably recognized as a humanoid by the system. The boundary box of the spatio-temporally consistent humanoid region is extracted to determine the specific position and size information of the humanoid target. This information is used to construct a target state vector, which includes the position, size, and other relevant features of the humanoid target in the current frame, thereby obtaining an initial tracking state. To predict the position of the humanoid target in the next frame, a Kalman filter is used to predict the initial tracking state. The Kalman filter is a classic state estimation tool that estimates the next frame's predicted position based on the current state estimate. The predicted position is correlated with the actual recognition result of the next frame to obtain an updated tracking state, thus achieving continuous tracking of the humanoid target in the video sequence. Motion consistency constraints are imposed on the updated tracking state, and abnormal tracking results that do not conform to the regular motion pattern are eliminated according to a reasonable motion pattern. In this way, abnormal motion trajectories caused by noise or recognition errors are eliminated, and an optimized tracking sequence is obtained. The optimized tracking sequence is smoothed to generate a humanoid target tracking sequence.

[0036] S5. Calculate the scene activity index according to the humanoid target tracking sequence, generate a multi-mode adaptive monitoring strategy, and dynamically adjust the low-power working scheme of the low-power monitoring device based on the multi-mode adaptive monitoring strategy;

[0037] Specifically, count the number of targets in the humanoid target tracking sequence to obtain the frequency of target appearance per unit time. Calculate the scene crowd density index based on the target appearance frequency, and obtain the crowd density score according to this index, which reflects the intensity of crowd activities in the monitoring scene. At the same time, analyze the movement speed of the targets in the humanoid target tracking sequence to obtain the average movement speed and the speed change rate. These data are used to calculate the dynamic index of the scene and obtain the scene dynamic score, which reflects the movement activity of the targets in the monitoring scene. Combine the crowd density score and the scene dynamic score, and calculate a comprehensive scene activity index through weighted summation. The comprehensive index can comprehensively reflect the overall activity state of the current monitoring scene. Collect the light intensity data of the current environment to obtain the light intensity value. According to the light intensity value, select appropriate image enhancement parameters and formulate a light adaptation processing strategy, which can optimize the image quality under the condition of light change and ensure the stability of the monitoring effect. At the same time, read the current battery information of the device to obtain the battery percentage. According to the battery percentage, set the corresponding power consumption control level and generate a battery self-adaptive control strategy to ensure that the device can reasonably adjust the power consumption and extend the running time in the case of low battery. Based on the scene activity index, the light adaptation processing strategy and the battery self-adaptive control strategy, generate a multi-mode self-adaptive monitoring strategy. This strategy intelligently adjusts the working mode of the monitoring device according to different scene requirements and device states. Set the image acquisition frequency according to the multi-mode self-adaptive monitoring strategy to form a dynamic sampling scheme, and adjust the working parameters of the image sensor according to this scheme to obtain an optimized image acquisition strategy. This strategy can reduce the power consumption during the image acquisition process while ensuring the monitoring effect. According to the multi-mode self-adaptive monitoring strategy, select an appropriate image processing complexity and generate an algorithm complexity control scheme. Combine the algorithm complexity control scheme and the optimized image acquisition strategy to formulate a low-power working scheme for the low-power monitoring device. This working scheme flexibly controls the device power consumption in different scenes by dynamically adjusting the image acquisition frequency and processing complexity of the device, and achieves a balance between efficient monitoring and energy consumption management.

[0038] S6. Perform incremental learning and model update according to the humanoid target tracking sequence to obtain a target humanoid recognition model, and generate a multi-level wake-up control strategy based on the target humanoid recognition model and the low-power working scheme.

[0039] Specifically, the target images in the humanoid target tracking sequence are cropped and aligned to obtain standardized target samples to ensure consistency in subsequent processing. Feature extraction is performed on the standardized target samples to generate a new sample feature set. The new sample feature set is compared with the historical sample feature library, and the novelty of the samples is evaluated through the comparison process. Based on the evaluation results of the sample novelty, a sample set suitable for incremental learning is selected. The incremental learning sample set represents the new or changed target features in the current monitoring scenario, and these features are used to optimize the existing recognition model. Based on the selected incremental learning sample set, the last layer of the initial humanoid recognition model is fine-tuned to obtain updated model parameters. To ensure the smoothness of model updates, the updated model parameters are weighted and fused with the original model parameters to generate a temporarily updated model. Performance evaluation is performed on the temporarily updated model, and the recognition accuracy and computational complexity are calculated to obtain model update evaluation metrics. These evaluation metrics are used to determine whether the model update is effective, and based on the evaluation results, it is decided whether to accept this update. If the update results meet the expected standards, a final target humanoid recognition model is generated. Based on the target humanoid recognition model, multiple confidence thresholds are set to generate a multi-level humanoid recognition judgment criterion. These judgment criteria allow the system to quickly screen real-time images according to different confidence levels and generate a preliminary wake-up trigger signal. Combining with the scene activity index in the low-power working scheme, the weight of the preliminary wake-up trigger signal is adjusted to obtain a scene-adaptive wake-up score. Multiple wake-up thresholds are set based on the scene-adaptive wake-up score to achieve flexible response to different scenes. According to the multiple wake-up thresholds, the wake-up process of the device is divided into multiple energy consumption levels to form a stepped wake-up strategy. This strategy minimizes energy consumption to the greatest extent while ensuring the effectiveness of monitoring through hierarchical wake-up. Energy consumption optimization is performed on the stepped wake-up strategy to ensure balanced distribution of energy consumption at different energy consumption levels, and a multi-level wake-up scheme with balanced energy consumption is obtained. The multi-level wake-up scheme with balanced energy consumption is combined with the low-power working scheme to generate a comprehensive multi-level wake-up control strategy. This strategy can intelligently adjust the wake-up and working modes of the device according to the real-time changes in the monitoring scenario and the power consumption status of the device to achieve efficient monitoring under low power consumption.

[0040] In one example, the original image sequence collected by the low-power monitoring device is downsampled and grayscaled to obtain a preprocessed image and perform background difference to obtain a preliminary target mask, including:

[0041] The resolution of the original image sequence collected by the low-power monitoring device is reduced to obtain a downsampled image, and the downsampled image is grayscaled to obtain a grayscale image;

[0042] The grayscale image is smoothed to obtain a preprocessed image, and the preprocessed image is edge-enhanced to obtain a preprocessed image sequence;

[0043] Select multiple frames of images from the preprocessed image sequence for pixel-level averaging to obtain a background reference image, and perform a difference operation between the newly input preprocessed image and the background reference image to obtain a difference image;

[0044] Perform threshold processing on the difference image, mark the pixels greater than the first threshold as foreground to obtain a binary difference image, and perform noise removal processing on the binary difference image to obtain an optimized moving region image;

[0045] Perform edge detection on the optimized moving region image to obtain a target contour image, and extract the target contour line segments of the target contour image;

[0046] Compare the target contour line segments with a preset humanoid contour template, calculate the similarity score to obtain a humanoid feature description, and compare the humanoid feature description with a second threshold to screen and obtain a preliminary target mask.

[0047] In this example, perform a resolution reduction operation on the original image sequence collected by the low-power monitoring device to obtain a downsampled image, reduce the number of pixels in the image, thereby reducing the computational complexity and the consumption of device resources. The resolution reduction can be achieved by downsampling the original image, that is, by averaging the color values of adjacent pixels or directly selecting one of the pixel values to represent a larger area, reducing the overall resolution of the image. For example, assume the resolution of the original image is , and through the downsampling factor , the resolution of the obtained downsampled image will become . Convert the downsampled image to grayscale to generate a grayscale image. Simplify the color image to a single-channel grayscale image to reduce the amount of calculation. The usual grayscale conversion method is to use the weighted average method. For example, the grayscale value can be calculated by the following formula:

[0048]

[0049] where is the grayscale value, are the pixel values of the red, green, and blue channels in the image respectively. Through this formula, each color pixel is converted into the corresponding grayscale value to generate a single-channel grayscale image. Perform image smoothing processing on the grayscale image. The commonly used smoothing processing method is Gaussian blur. Gaussian blur makes the value of each pixel become the weighted average of the values of its surrounding pixels through a convolution operation, thereby achieving a smoothing effect. The convolution kernel of Gaussian blur can be expressed as:

[0050]

[0051] where is the weight of the convolution kernel, represents the standard deviation of the Gaussian function, which determines the degree of blurring. After Gaussian blurring, a preprocessed image is obtained. Edge enhancement is performed on the preprocessed image. Edge enhancement is usually achieved through gradient operators such as the Sobel operator and the Laplace operator. By edge enhancement, the edge information of the objects in the image is highlighted, and a clearer object contour is obtained. The set of processed multi-frame images is used as a preprocessed image sequence, and these images are used for subsequent background difference analysis. Multiple frames of images are selected from the preprocessed image sequence for pixel-level averaging to obtain a background reference image. The background reference image represents the static background part of the scene and is used to compare with the newly input preprocessed image. The background reference image can be calculated by the following formula:

[0052]

[0053] where, represents the pixel value of the -th frame image at position , represents the number of selected frames. The obtained background reference image is subjected to a difference operation with the newly input preprocessed image to generate a difference image. The purpose of the difference operation is to compare the current image with the background and identify the moving objects that appear in the scene. The difference image can be obtained by the following formula:

[0054]

[0055] where, is the preprocessed image of the current frame, is the background reference image,

[0056] is the difference image, representing the changing part of the scene. Threshold processing is performed on the difference image, and the pixels greater than the preset threshold are marked as the foreground to obtain a binary difference image. This processing step separates the moving objects in the image from the background. The binary processing is performed in the following way:

[0057]

[0058] where, is the binary image, is the threshold. When the pixel value in the difference image is greater than the threshold, the pixel is marked as foreground (value 1), otherwise it is marked as background (value 0). Noise removal is performed on the binary difference image. Isolated small noise points are removed through morphological operations such as opening or closing operations to obtain a more coherent moving region image. The opening operation eliminates small noise points by eroding first and then dilating, while the closing operation fills small holes by dilating first and then eroding. After the noise removal process, an optimized moving region image is obtained. Edge detection is performed on the optimized moving region image to extract the contour of the object in the image. Common edge detection methods include the Canny edge detection algorithm. Through this algorithm, a target contour image is obtained, and the target contour line segments in the image are extracted. The target contour line segments are compared with a preset humanoid contour template. The similarity score of the contour is calculated. The higher the similarity score, the closer the target contour is to the humanoid contour template. The similarity can be calculated through shape matching metrics such as the Hausdorff distance. The similarity score is used to generate the final humanoid feature description. The calculated humanoid feature description is compared with a second threshold. Only when the similarity score exceeds this threshold will the corresponding region be recognized as a target region that may contain a humanoid, and finally a preliminary target mask is obtained through screening. The target mask is used to mark the regions in the image that are recognized as humanoids.

[0059] In one example, humanoid feature extraction is performed on the preliminary target mask to obtain candidate humanoid regions and calculate shape descriptors, resulting in humanoid region pseudo-labels, including:

[0060] The preliminary target mask is framed with a rectangle to obtain candidate rectangular regions, and according to a preset human body proportion range, the candidate rectangular regions are screened to obtain preliminary humanoid regions;

[0061] Gradient features are extracted from the preliminary humanoid regions to obtain a histogram of gradient directions, and the feature similarity is calculated based on the histogram of gradient directions to obtain a humanoid likelihood score;

[0062] Based on the humanoid likelihood score, the preliminary humanoid regions are screened again to obtain refined humanoid regions, and the overlapping regions of the refined humanoid regions are merged to obtain candidate humanoid regions;

[0063] Shape description data is extracted from the candidate humanoid regions, and the shape description data is feature-compressed to obtain a compact shape descriptor;

[0064] The compact shape descriptors are matched to obtain shape matching results, and similarity sorting is performed based on the shape matching results to obtain the best matching shape;

[0065] Calculate the comprehensive confidence by combining the best - matching shape and the human - like possibility score, obtain the confidence map of the human - like region, and perform region expansion on the confidence map of the human - like region to obtain the complete human - like region;

[0066] Extract the feature description vector from the complete human - like region, and filter the feature description vector to obtain the pseudo - label of the human - like region.

[0067] In this example, rectangular framing is performed on the preliminary target mask to obtain candidate rectangular regions. Rectangular framing can be achieved by finding continuous foreground pixel regions on the preliminary target mask and enclosing them with the minimum bounding rectangle to form candidate rectangular regions. Each rectangular region may contain a complete human figure or a partial human figure. According to the preset human body proportion range, the candidate rectangular regions are filtered to eliminate regions that do not conform to the human body proportion. Human body proportions usually include the ratio of height to shoulder width, the ratio of height to leg length, etc. Through ratio restrictions, noise regions that do not conform to the human body structure are effectively filtered out. Extract the gradient features of each preliminary human - like region. Gradient features reflect the direction and intensity of pixel brightness changes in the image and can capture edge information in the image. Histogram of Oriented Gradients (HOG) is a commonly used feature representation method. By calculating the gradient direction of each pixel in the image and statistically counting the gradient frequencies in different directions, a histogram is formed. For each preliminary human - like region, calculate its histogram of oriented gradients , where each entry represents the cumulative gradient value in the th direction. This histogram can effectively capture the directional distribution characteristics of edges within the region. To evaluate whether these regions are likely to contain human - like targets, calculate the feature similarity based on the histogram of oriented gradients. Feature similarity can be measured by calculating the distance between the histogram of oriented gradients of the current region and the preset human - like template histogram. Assume the histogram of oriented gradients of the human - like template is , then the feature similarity can be calculated by the following formula:

[0068]

[0069] where represents the similarity score, is the number of entries in the histogram. The similarity score The value ranges from 0 to 1. The closer the value is to 1, the more similar the current area is to the humanoid template. In this way, a humanoid possibility score is calculated for each preliminary humanoid area. Based on the calculated humanoid possibility scores, a secondary screening is performed on the preliminary humanoid areas. Only those areas with higher possibility scores will be retained to form a more refined humanoid area. To avoid multiple adjacent refined humanoid areas from repeatedly identifying the same target, the overlapping areas of these areas are merged. For two areas with a high degree of overlap, they are merged into a larger area to avoid redundant detection, resulting in candidate humanoid areas. Shape description data is extracted from the candidate humanoid areas. The shape description data usually includes geometric features of the contour, such as perimeter, area, rectangularity, etc. These description data can effectively represent the shape features of the objects in the image. To reduce the computational complexity, the extracted shape description data is compressed in features to generate a compact shape descriptor. Feature compression can be achieved through dimensionality reduction techniques such as principal component analysis, compressing the high-dimensional shape description data into a lower dimension while retaining the most representative feature information. The generated compact shape descriptor is matched with a preset standard shape template, and the shape matching result is obtained by calculating the similarity. The higher the similarity, the closer the shape of the candidate area is to the humanoid template. To determine the best-matching shape, the similarity of all shape matching results is sorted, and the result with the highest similarity is selected as the best-matching shape. Combining the best-matching shape and the previously calculated humanoid possibility scores, the confidence of each candidate area is comprehensively calculated. The confidence is a comprehensive index used to evaluate the possibility of each area being a humanoid target. The confidence information is used to generate a humanoid area confidence map. Each pixel value in the confidence map represents the probability that the pixel belongs to the humanoid target. On this basis, the humanoid area confidence map is expanded in area to fill in the local blanks caused by noise or other factors, resulting in a complete humanoid area. Feature description vectors are extracted from the complete humanoid area, containing the key feature information of the target area. By screening the feature description vectors, those areas that do not conform to the humanoid target features are removed, and the humanoid area pseudo-labels are obtained.

[0070] In one example, an initial humanoid recognition model is trained based on the humanoid area pseudo-labels and multi-scale feature extraction is performed to obtain a local spatial depth feature map and a target feature vector, including:

[0071] The humanoid area pseudo-labels are processed for data augmentation to obtain an augmented training set, and the augmented training set is standardized to obtain standardized training data;

[0072] A lightweight convolutional neural network backbone containing 5 convolutional blocks is constructed. Each convolutional block contains a convolutional layer, a batch normalization layer, and a ReLU activation function layer, and the preprocessed image is input through the convolutional neural network backbone for feature extraction to obtain a preliminary feature map;

[0073] Insert a discrete wavelet transform layer between the 3rd and 4th convolutional blocks of the convolutional neural network backbone, and perform discrete wavelet transform on the feature map passing through the 3rd convolutional block to obtain low-frequency approximation coefficients and high-frequency detail coefficients;

[0074] Perform adaptive weighted fusion on the low-frequency approximation coefficients and high-frequency detail coefficients to obtain a multi-scale feature map, and input the multi-scale feature map into the 4th convolutional block to obtain an enhanced feature representation;

[0075] Add a spatial pyramid pooling module after the last convolutional block of the convolutional neural network backbone; perform multi-scale pooling on the output feature map of the last convolutional block to obtain feature vectors of different scales;

[0076] Concatenate the feature vectors of different scales to obtain a multi-scale feature representation, and reduce the dimension of the multi-scale feature representation through 1x1 convolution to obtain a compressed feature vector;

[0077] Construct a fully connected layer network, including two fully connected layers and a Dropout layer, and input the compressed feature vector into the fully connected layer network to output a prediction result;

[0078] Use the cross-entropy loss function to calculate the error between the prediction result and the true label to obtain a loss value; optimize the network parameters through the backpropagation algorithm to obtain updated network parameters until the loss value converges or reaches the preset number of iterations to obtain an initial humanoid recognition model;

[0079] Input the new preprocessed image into the initial humanoid recognition model, and after processing by all convolutional blocks, discrete wavelet transform layers and spatial pyramid pooling modules, obtain a local spatial depth feature map, and perform global average pooling operation on the local spatial depth feature map to obtain a target feature vector.

[0080] In this example, perform data augmentation on the pseudo-labels of the humanoid region. Data augmentation is a method to increase the diversity of training data. By performing operations such as rotation, scaling, translation, flipping, and noise addition on existing samples, new training samples are generated. Assume that a humanoid image in the initial dataset is , and through data augmentation, multiple transformed images can be generated, such as 。The enhanced image is used to augment the training set to obtain more diverse training data. The augmented training set is normalized by scaling the pixel values of the image data to a unified range, usually by subtracting the mean of the data set from each pixel value and dividing by the standard deviation, thereby accelerating the convergence process of the model and obtaining normalized training data. A lightweight convolutional neural network backbone consisting of 5 convolutional blocks is constructed. Each convolutional block consists of a convolutional layer, a batch normalization layer, and a ReLU activation function layer. The role of the convolutional layer is to extract local features of the image, the batch normalization layer is used to accelerate network training and improve the stability of the model, and the ReLU activation function introduces non-linearity to enable the model to better fit complex data. The design of the convolutional neural network backbone is lightweight to adapt to the computing power of low-power monitoring devices while being able to effectively extract feature information from the input image. The input preprocessed image is processed by this convolutional neural network backbone to generate a preliminary feature map. A discrete wavelet transform layer is inserted between the 3rd and 4th convolutional blocks of the convolutional neural network backbone. Discrete wavelet transform can extract the low-frequency approximation coefficients and high-frequency detail coefficients of the image by decomposing the image signal. The low-frequency approximation coefficients mainly represent the overall structural information of the image, while the high-frequency detail coefficients capture the edge and texture details of the image. For example, assuming the output feature map of the 3rd convolutional block is , through discrete wavelet transform, the low-frequency approximation coefficients and the high-frequency detail coefficients are obtained. These two coefficients respectively retain different scale information of the image. Adaptive weighted fusion is performed on the low-frequency approximation coefficients and high-frequency detail coefficients. By combining feature information of different frequencies, a more rich multi-scale feature map is generated. The fused multi-scale feature map is represented by the formula:

[0081]

[0082] where and is an adaptive weight parameter, and the values of these two parameters are dynamically adjusted according to specific task requirements to balance the contributions of low-frequency and high-frequency information. The fused multi-scale feature map is input into the 4th convolutional block for feature extraction to obtain an enhanced feature representation. A spatial pyramid pooling module is added after the last convolutional block of the convolutional neural network backbone. The spatial pyramid pooling module generates multiple feature vectors of different scales by performing pooling operations on the feature map at different scales. These feature vectors retain information of different spatial scales and can adapt to input images of different sizes. Feature vectors of all scales are concatenated to form a multi-scale feature representation. To reduce the computational complexity of the model, the concatenated multi-scale feature representation is dimensionally reduced through 1x1 convolution to generate a compressed feature vector. The role of 1x1 convolution is to map high-dimensional features to a low-dimensional space through linear combination while retaining the most useful information. Assume the multi-scale feature representation is and the compressed feature vector is , then the 1x1 convolution can be expressed as:

[0083]

[0084] where is the weight matrix, is the bias term, is the finally generated low-dimensional feature representation. Construct a fully connected layer network. This network contains two fully connected layers and a Dropout layer, which are used to process the feature vectors and generate the final prediction results. The fully connected layer maps the feature vectors to the target output space through linear transformation and non-linear activation, while the Dropout layer is used to prevent overfitting. By randomly discarding some neurons, the generalization ability of the model is enhanced. During the training process, the cross-entropy loss function is used to calculate the error between the prediction result and the true label. Assume the prediction result is and the true label is , the cross-entropy loss function can be expressed as:

[0085]

[0086] where is the number of classes, and are respectively the The true labels and predicted probabilities of the classes. By minimizing the cross-entropy loss, the model parameters are optimized. The backpropagation algorithm is used to adjust the weights and biases in the network, gradually reducing the loss value. The training process continues until the loss value converges or reaches the preset number of iterations, obtaining the initial humanoid recognition model. After the model training is completed, the new preprocessed image is input into the initial humanoid recognition model. After being processed by all convolutional blocks, discrete wavelet transform layers, and spatial pyramid pooling modules, a local spatial depth feature map is generated. To aggregate the feature information, a global average pooling operation is performed on the local spatial depth feature map to generate the target feature vector. Global average pooling generates a fixed-length vector by taking the average value of each channel in the feature map, representing the features of the entire image.

[0087] In one example, the target feature vector is classified and weighted voted to obtain a preliminary humanoid recognition result, and the recognition results of consecutive frames are optimized to obtain a humanoid target tracking sequence, including:

[0088] The target feature vector is normalized to obtain a normalized feature vector, and the normalized feature vector is input into multiple preconfigured kernel extreme learning machines to obtain multiple sets of classification results;

[0089] The multiple sets of classification results are weighted averaged to obtain a comprehensive classification score, and the comprehensive classification score is binarized according to a preset classification threshold to obtain a preliminary humanoid recognition result;

[0090] Time series analysis is performed on the preliminary humanoid recognition results of consecutive multiple frames to obtain the temporal consistency score of the humanoid target, and the single-frame recognition result is corrected based on the temporal consistency score to obtain the temporally smoothed recognition result;

[0091] Spatial connectivity analysis is performed on the temporally smoothed recognition result to obtain the spatial consistency score, and the spatio-temporally consistent humanoid region is obtained by combining the spatial consistency score and the temporally smoothed recognition result;

[0092] Bounding box extraction is performed on the spatio-temporally consistent humanoid region to obtain the position and size information of the humanoid target, and an object state vector is constructed based on the position and size information to obtain the initial tracking state;

[0093] The Kalman filter is used to predict the initial tracking state to obtain the predicted position of the next frame, and the predicted position is associated with the recognition result of the next frame to obtain the updated tracking state;

[0094] Motion consistency constraints are imposed on the updated tracking state to eliminate abnormal tracking results that do not conform to the motion pattern, obtaining an optimized tracking sequence, and the optimized tracking sequence is smoothed to obtain the humanoid target tracking sequence.

[0095] In this example, the target feature vector is normalized to adjust the value range of the feature vector to a unified range, eliminate the scale differences between features, and obtain a standardized feature vector. The standardized feature vector is input into multiple pre-configured kernel extreme learning machines to obtain multiple sets of classification results. The kernel extreme learning machine is a fast learning algorithm based on kernel functions. It processes data through randomly generated hidden layers and kernel function mappings, and classifies by minimizing the weights of the output layer. The multiple sets of classification results are weighted and averaged. By combining the output results of different classifiers and assigning different weights to different classifiers, a more accurate comprehensive classification score is generated. According to the preset classification threshold, the comprehensive classification score is binarized to generate a preliminary humanoid recognition result. If the comprehensive classification score exceeds the threshold, the area is marked as a humanoid target; otherwise, it is marked as non-humanoid. Time series analysis is performed on the preliminary humanoid recognition results of multiple consecutive frames. By analyzing the consistency of the target in multiple consecutive frames, the temporal consistency score of the humanoid target is calculated. The temporal consistency score reflects the stability of the target over time. The higher the score, the greater the likelihood that the target is continuously recognized as a humanoid in multiple consecutive frames. Based on the temporal consistency score, the recognition result of a single frame is corrected to obtain a temporally smoothed recognition result. Spatial connectivity analysis is performed on the temporally smoothed recognition result to check the spatial consistency of the recognition result. By evaluating the coherence of the recognized area, the spatial consistency score is calculated. A high spatial consistency score means that the recognition result has high connectivity in space and low noise interference. By combining the temporal consistency score and the spatial consistency score, a spatio-temporally consistent humanoid area is obtained, which has high credibility in both time and space. For the recognized humanoid area, a bounding box is extracted to obtain the position and size information of each humanoid target. The bounding box is used to accurately locate the range of the humanoid target, and the position and size information are used to construct the target state vector. The target state vector usually contains dynamic information such as the position and speed of the target, describing the state of the target in the current frame. To predict the position of the humanoid target in the next frame, a Kalman filter is used to predict the initial tracking state. The Kalman filter is a classic recursive algorithm used to estimate the future state from the current state. Assuming the initial state vector of the target is , the predicted position in the next frame can be calculated using the following formula:

[0096]

[0097] where is the state transition matrix, which describes the state change of the target from the current frame to the next frame, is the control input, is the process noise. Through the prediction of the Kalman filter, the predicted position of the target in the next frame is obtained, and this predicted position is associated with the actual recognition result in the next frame to obtain the updated tracking state. In the updated tracking state, there may be incorrect results caused by noise or abnormal movement. To improve the tracking accuracy, motion consistency constraints are imposed on the updated tracking state to eliminate those abnormal tracking results that do not conform to the normal motion pattern. For example, if the speed or direction of the target suddenly changes abnormally, this result may be incorrect and needs to be eliminated. Through motion consistency constraints, an optimized tracking sequence is obtained. The optimized tracking sequence is smoothed to generate a stable humanoid target tracking sequence. The smoothing process generates a more smooth and continuous tracking path by removing short-term fluctuations to obtain the humanoid target tracking sequence.

[0098] In one example, the scene activity index is calculated based on the humanoid target tracking sequence, a multi-mode adaptive monitoring strategy is generated, and the low-power working scheme of the low-power monitoring device is dynamically adjusted based on the multi-mode adaptive monitoring strategy, including:

[0099] Count the number of targets in the humanoid target tracking sequence to obtain the target appearance frequency per unit time, and calculate the scene crowd density index based on the target appearance frequency to obtain the crowd density score;

[0100] Analyze the target movement speed in the humanoid target tracking sequence to obtain the average movement speed and the speed change rate, and calculate the scene dynamic index based on the average movement speed and the speed change rate to obtain the scene dynamic score;

[0101] Combine the crowd density score and the scene dynamic score, and calculate the comprehensive scene activity index through weighted summation;

[0102] Collect the current environmental light intensity data to obtain the light intensity value, and select appropriate image enhancement parameters based on the light intensity value to obtain the light adaptation processing strategy;

[0103] Read the current battery information of the device to obtain the battery percentage, and set the power consumption control level based on the battery percentage to obtain the battery self-adaptive control strategy;

[0104] Based on the scene activity index, the light adaptation processing strategy, and the battery self-adaptive control strategy, generate a multi-mode adaptive monitoring strategy;

[0105] Set the image acquisition frequency according to the multi-mode adaptive monitoring strategy to obtain a dynamic sampling scheme, and adjust the working parameters of the image sensor based on the dynamic sampling scheme to obtain an optimized image acquisition strategy;

[0106] Select the image processing complexity according to the multi-mode adaptive monitoring strategy to obtain the algorithm complexity control scheme, and formulate the low-power working scheme of the low-power monitoring device based on the algorithm complexity control scheme and the optimized image acquisition strategy.

[0107] In this example, analyze the number of targets in the human target tracking sequence. By counting the target appearance frequency per unit time, calculate the pedestrian flow density in the scene. Assume that in a time period the number of detected targets is , then the target appearance frequency can be expressed as:

[0108]

[0109] where represents the number of targets appearing per unit time. The target appearance frequency directly reflects the density of the pedestrian flow in the scene. According to this frequency, calculate the pedestrian flow density index. The pedestrian flow density index is a key parameter for measuring the crowd activities in the scene, and is usually obtained by combining the target appearance frequency with the scene area to get a more accurate evaluation. Assume that the scene area is , then the pedestrian flow density index can be expressed as:

[0110]

[0111] Based on the calculated pedestrian flow density index, generate a pedestrian flow density score to describe the crowd density in the scene. The higher the score, the more frequent the crowd activities and the greater the density in the scene. Analyze the target movement speed in the human target tracking sequence. Calculate the average movement speed of each target. Assume that a certain target moves a distance in the time period , then the average movement speed can be expressed as:

[0112]

[0113] where represents the average moving distance of the target per unit time. After calculating the average speed, analyze the speed change rate of the target, that is, the change of the target speed in different time periods. The speed change rate can be measured by calculating the standard deviation of the speed. Assume that the measured speed values in the time period are then the speed change rate can be expressed as:

[0114]

[0115] where Indicates the degree of speed fluctuation. Based on the average motion speed and the speed change rate, the dynamic index of the scene is calculated. The scene dynamic index is used to describe the intensity of the motion activity in the scene, reflecting the motion law of the target. According to this index, a scene dynamic score is generated to measure the activity level of the motion in the scene. Combining the pedestrian flow density score with the scene dynamic score, a comprehensive scene activity index is calculated by weighted summation. Assuming the pedestrian flow density score is , and the scene dynamic score is , then the comprehensive scene activity index can be expressed as:

[0116]

[0117] where, and is a weight parameter used to adjust the contribution ratio of the two scores to the scene activity level. The comprehensive scene activity index is used to reflect the overall activity level in the scene and is an important basis for subsequent monitoring strategy adjustment. Consider the impact of environmental light intensity on the monitoring effect. Collect the current environmental light intensity data through sensors to obtain the light intensity value. According to the light intensity value, select appropriate image enhancement parameters to ensure the stability of image quality under different lighting conditions. For example, in the case of low light, it may be necessary to increase the exposure time or gain to enhance the image brightness. Generate a light adaptation processing strategy based on these adjustments to adapt to different lighting conditions. At the same time, read the current power information of the device to obtain the power percentage. The power percentage is a key indicator to measure the remaining power of the device. Set different power consumption control levels according to the power percentage. The power consumption control level determines the operating mode of the device. For example, when the power is low, enter the energy-saving mode, reduce the image acquisition frequency and processing complexity, and extend the working time of the device. Generate a power adaptation control strategy based on the power information to ensure that the device can operate efficiently under different power states. Based on the scene activity index, light adaptation processing strategy, and power adaptation control strategy, generate a multi-mode adaptive monitoring strategy. The multi-mode adaptive monitoring strategy can dynamically adjust the operating mode of the monitoring device according to the real-time scene and device status. For example, in a high-scene activity, the system may increase the image acquisition frequency and processing complexity to ensure the accurate recognition and tracking of all targets in the scene; while in the case of low power, the system may reduce these parameters to reduce energy consumption. According to the multi-mode adaptive monitoring strategy, set the image acquisition frequency to obtain a dynamic sampling scheme. The dynamic sampling scheme optimizes the image acquisition process by adjusting the working parameters of the image sensor, such as frame rate, resolution, etc., to ensure that important information can be efficiently captured in different scenarios. For example, in a high pedestrian density and high-dynamic scene, the system may increase the frame rate to capture fast-moving targets; while in a low-activity scene, reduce the frame rate to save energy. Select the complexity of image processing according to the multi-mode adaptive monitoring strategy to obtain an algorithm complexity control scheme. The algorithm complexity control scheme involves selecting appropriate image processing algorithms and models. For example, in a high-activity scene, the system may use a more complex deep learning model to improve the recognition accuracy; while in the low-power mode, use a lightweight model to reduce the computational burden. Combine the algorithm complexity control scheme and the optimized image acquisition strategy to formulate a low-power working scheme for a low-power monitoring device.

[0118] In one example, incremental learning and model update are performed according to the human target tracking sequence to obtain a target human recognition model, and a multi-level wake-up control strategy is generated based on the target human recognition model and the low-power working scheme, including:

[0119] Crop and align the target images in the humanoid target tracking sequence to obtain standardized target samples, and extract features from the standardized target samples to obtain a new sample feature set;

[0120] Compare the new sample feature set with the historical sample feature library to obtain the sample novelty evaluation result, and screen out the incremental learning sample set according to the sample novelty evaluation result;

[0121] Fine-tune the last layer of the initial humanoid recognition model based on the incremental learning sample set to obtain updated model parameters, and perform weighted fusion of the updated model parameters and the original model parameters to obtain a temporarily updated model;

[0122] Evaluate the performance of the temporarily updated model, calculate the recognition accuracy and computational complexity to obtain the model update evaluation index, and decide whether to accept the update according to the model update evaluation index to obtain the target humanoid recognition model;

[0123] Set multiple confidence thresholds based on the target humanoid recognition model to obtain a multi-level humanoid recognition judgment criterion, and quickly screen the real-time image according to the multi-level humanoid recognition judgment criterion to obtain a preliminary wake-up trigger signal;

[0124] Combine the scene activity index in the low-power working scheme to adjust the weight of the preliminary wake-up trigger signal to obtain a scene-adaptive wake-up score, and set multi-level wake-up thresholds based on the scene-adaptive wake-up score;

[0125] According to the multi-level wake-up thresholds, divide the device wake-up process into multiple energy consumption levels to obtain a stepped wake-up strategy, and optimize the energy consumption of the stepped wake-up strategy to obtain an energy consumption-balanced multi-level wake-up scheme;

[0126] Combine the energy consumption-balanced multi-level wake-up scheme with the low-power working scheme to generate a multi-level wake-up control strategy.

[0127] In this example, operations of cropping and aligning the target images in the tracking sequence are performed. Regions containing humanoid targets are extracted from the video frames, and these regions are standardized to ensure the stability and consistency of the subsequent feature extraction process. The cropping operation cuts out the humanoid target part in the image by analyzing the bounding box of the target; the alignment operation adjusts the scale and corrects the rotation of the cropped image according to the position and pose of the target, so that all target samples are spatially consistent. Through these processes, standardized target samples are obtained. Feature extraction is performed on the standardized target samples to generate a new sample feature set. The image data is converted into high-dimensional feature vectors, which can effectively represent important information such as the shape and texture of the target. Common feature extraction methods include extracting deep features using a convolutional neural network or extracting edge features using the histogram of oriented gradients. Suppose a certain standardized target sample is , and through the feature extraction process, a feature vector is generated, where each represents a feature of the image. These feature vectors form a new sample feature set, representing the newly acquired target information. The new sample feature set is compared with the historical sample feature library to evaluate the similarity and novelty of these new samples. The historical sample feature library stores the target features learned in the past. By calculating the similarity between the new samples and these historical features, it is judged whether the new samples contain new target types or new poses. Novelty evaluation is usually achieved by calculating the Euclidean distance or cosine similarity between feature vectors. Suppose a sample feature in the historical sample feature library is , and the Euclidean distance between the new sample and it can be expressed as:

[0128]

[0129] where represents the distance between the two. The larger the distance, the greater the difference between the new sample and the historical sample, and the stronger the novelty. According to the novelty evaluation results, a part of the samples with higher novelty are selected to form an incremental learning sample set. Based on the selected incremental learning sample set, the last layer of the initial humanoid recognition model is fine-tuned. While retaining the original model features, it adapts to the new information brought by the new samples and improves the generalization ability of the model. The first few layers of the model are frozen, and only the parameters of the last layer are updated. The weights of the last layer are adjusted by the gradient descent method to generate updated model parameters. To balance the influence of new and old features, the updated model parameters and the original model parameters are weighted and fused to obtain a temporarily updated model. Suppose the original model parameters are , and the updated parameters are , then the fused temporarily updated model parameters can be expressed as:

[0130]

[0131] Among them, is the fusion coefficient, which determines the contribution ratio of the old and new parameters. Perform performance evaluation on the temporarily updated model, calculate the recognition accuracy and computational complexity, and obtain the model update evaluation index. The recognition accuracy measures the performance of the model on new data, while the computational complexity reflects the resources required for the model to run. Based on these evaluation indexes, judge whether to accept the update. If the performance of the temporarily updated model is better than that of the original model, accept the update and generate the final target humanoid recognition model. Based on the target humanoid recognition model, set multiple confidence thresholds to generate a multi-level humanoid recognition judgment criterion. The judgment criterion is used to quickly screen the real-time image to determine whether to trigger device wake-up. When the recognition confidence exceeds a certain threshold, generate a preliminary wake-up trigger signal to start a higher-precision detection module. Combine the preliminary wake-up trigger signal with the scene activity index in the low-power working scheme and adjust its weight. The scene activity index reflects the dynamics of the current monitored scene, such as the flow density and movement speed. By adjusting the weight of the preliminary wake-up trigger signal, generate a scene-adaptive wake-up score. The higher the scene-adaptive wake-up score, the more it indicates that an important event has occurred in the current scene and the stronger the necessity to wake up the device. Based on the scene-adaptive wake-up score, set multiple wake-up thresholds, divide the device wake-up process into multiple energy consumption levels, and form a stepped wake-up strategy. For example, when the wake-up score is low, the device may only wake up the low-power mode for basic detection; when the wake-up score is high, the device may start the high-precision mode for comprehensive monitoring. Optimize the energy consumption of the stepped wake-up strategy to obtain a multi-level wake-up scheme with balanced energy consumption. The goal of energy consumption optimization is to reasonably allocate energy consumption resources according to the battery state and task requirements of the device and extend the working time of the device. Combine the multi-level wake-up scheme with balanced energy consumption with the low-power working scheme to generate a comprehensive multi-level wake-up control strategy.

[0132] Referring to Figure 2 , this embodiment provides a low-power monitoring wake-up device based on humanoid recognition, including:

[0133] A preprocessing module 1, configured to perform downsampling and grayscale processing on the original image sequence collected by the low-power monitoring device, obtain a preprocessed image, and perform background difference to obtain a preliminary target mask;

[0134] A calculation module 2, configured to extract humanoid features from the preliminary target mask, obtain a candidate humanoid region, and perform shape descriptor calculation to obtain a humanoid region pseudo-label;

[0135] A training module 3, configured to train an initial humanoid recognition model based on humanoid region pseudo-labels and perform multi-scale feature extraction to obtain a local spatial depth feature map and a target feature vector;

[0136] A classification module 4, configured to classify and perform weighted voting on the target feature vector to obtain a preliminary humanoid recognition result, and optimize the recognition results of consecutive frames to obtain a humanoid target tracking sequence;

[0137] An adjustment module 5, configured to calculate a scene activity index according to the humanoid target tracking sequence, generate a multi-mode adaptive monitoring strategy, and dynamically adjust the low-power working scheme of the low-power monitoring device based on the multi-mode adaptive monitoring strategy;

[0138] A generation module 6, configured to perform incremental learning and model update according to the humanoid target tracking sequence to obtain a target humanoid recognition model, and generate a multi-level wake-up control strategy based on the target humanoid recognition model and the low-power working scheme.

[0139] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to the description in the above method embodiment, and details are not described herein again.

[0140] Refer to Figure 3 , in an embodiment of the present invention, a computer device is further provided. The computer device may be a server, and its internal structure may be as Figure 3 shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the above method.

[0141] Those skilled in the art can understand that Figure 3 the structure shown in

[0142] is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0143] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0144] It should be noted that in this article, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, apparatus, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, apparatus, article, or method including that element.

[0145] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall equally be included in the patent protection scope of the present invention.

Claims

1. A low-power monitoring and wake-up method based on humanoid recognition, characterized in that It includes the following steps: Downsample and grayscale the original image sequence collected by the low-power monitoring device, obtain a preprocessed image and perform background difference to obtain a preliminary target mask; Extract humanoid features from the preliminary target mask, obtain candidate humanoid regions and calculate shape descriptors to obtain humanoid region pseudo-labels; Train an initial humanoid recognition model based on the humanoid region pseudo-labels and perform multi-scale feature extraction to obtain a local spatial depth feature map and a target feature vector; Classify and perform weighted voting on the target feature vector to obtain a preliminary humanoid recognition result, and optimize the recognition results of consecutive frames to obtain a humanoid target tracking sequence; Calculate a scene activity index according to the humanoid target tracking sequence, generate a multi-mode adaptive monitoring strategy, and dynamically adjust the low-power working scheme of the low-power monitoring device based on the multi-mode adaptive monitoring strategy; Specifically, it includes: counting the number of targets in the humanoid target tracking sequence to obtain the target appearance frequency per unit time, and calculating a scene pedestrian flow density index based on the target appearance frequency to obtain a pedestrian flow density score; analyzing the target movement speed in the humanoid target tracking sequence to obtain the average movement speed and the speed change rate, and calculating a scene dynamic index according to the average movement speed and the speed change rate to obtain a scene dynamic score; combining the pedestrian flow density score and the scene dynamic score, and calculating a comprehensive scene activity index through weighted summation; collecting current environmental light intensity data to obtain a light intensity value, and selecting appropriate image enhancement parameters according to the light intensity value to obtain a light-adaptive processing strategy; reading the current battery information of the device to obtain a battery percentage, and setting a power consumption control level according to the battery percentage to obtain a battery-adaptive control strategy; generating a multi-mode adaptive monitoring strategy based on the scene activity index, the light-adaptive processing strategy and the battery-adaptive control strategy; setting an image acquisition frequency according to the multi-mode adaptive monitoring strategy to obtain a dynamic sampling scheme, and adjusting the working parameters of the image sensor based on the dynamic sampling scheme to obtain an optimized image acquisition strategy; selecting an image processing complexity according to the multi-mode adaptive monitoring strategy to obtain an algorithm complexity control scheme, and formulating the low-power working scheme of the low-power monitoring device based on the algorithm complexity control scheme and the optimized image acquisition strategy; Perform incremental learning and model update according to the humanoid target tracking sequence to obtain a target humanoid recognition model, and generate a multi-level wake-up control strategy based on the target humanoid recognition model and the low-power working scheme.

2. The low-power monitoring wake-up method based on humanoid recognition according to claim 1, characterized in that The downsampling and grayscaling of the original image sequence collected by the low-power monitoring device, obtaining a preprocessed image and performing background difference to obtain a preliminary target mask, includes: Reduce the resolution of the original image sequence collected by the low-power monitoring device to obtain a downsampled image, and perform grayscale conversion on the downsampled image to obtain a grayscale image; Perform image smoothing on the grayscale image to obtain a preprocessed image, and perform edge enhancement processing on the preprocessed image to obtain a sequence of preprocessed images; Select multiple frames of images from the sequence of preprocessed images for pixel-level averaging to obtain a background reference image, and perform a difference operation between the newly input preprocessed image and the background reference image to obtain a difference image; Perform threshold processing on the difference image, mark the pixels greater than the first threshold as foreground to obtain a binary difference image, and perform noise removal processing on the binary difference image to obtain an optimized moving region image; Perform edge detection on the optimized moving region image to obtain a target contour image, and extract the target contour line segments of the target contour image; Compare the target contour line segments with a preset human contour template, calculate a similarity score to obtain a human feature description, and compare the human feature description with a second threshold to screen and obtain a preliminary target mask.

3. The low-power monitoring wake-up method based on humanoid recognition according to claim 1, characterized in that Performing human feature extraction on the preliminary target mask to obtain a candidate human region and calculating a shape descriptor to obtain a human region pseudo-label, including: Perform rectangular framing on the preliminary target mask to obtain a candidate rectangular region, and screen the candidate rectangular region according to a preset human body proportion range to obtain a preliminary human region; Extract gradient features from the preliminary human region to obtain a histogram of gradient directions, and calculate a feature similarity according to the histogram of gradient directions to obtain a human likelihood score; Based on the human likelihood score, perform secondary screening on the preliminary human region to obtain a refined human region, and perform overlapping region merging on the refined human region to obtain a candidate human region; Extract shape description data from the candidate human region, and perform feature compression on the shape description data to obtain a compact shape descriptor; Match the compact shape descriptors to obtain a shape matching result, and perform similarity ranking according to the shape matching result to obtain the best matching shape; Combine the best matching shape and the human likelihood score to calculate a comprehensive confidence level to obtain a human region confidence map, and perform region expansion on the human region confidence map to obtain a complete human region; Extract a feature description vector from the complete human region, and screen the feature description vector to obtain a human region pseudo-label.

4. The low-power monitoring wake-up method based on humanoid recognition according to claim 1, characterized in that Training an initial human recognition model based on the human region pseudo-label and performing multi-scale feature extraction to obtain a local spatial depth feature map and a target feature vector, including: Perform data augmentation processing on the human region pseudo-label to obtain an augmented training set, and perform normalization processing on the augmented training set to obtain normalized training data; Construct a lightweight convolutional neural network backbone containing 5 convolutional blocks, each convolutional block contains a convolutional layer, a batch normalization layer, and a ReLU activation function layer, and perform feature extraction on the input preprocessed image through the convolutional neural network backbone to obtain a preliminary feature map; Insert a discrete wavelet transform layer between the 3rd convolutional block and the 4th convolutional block of the convolutional neural network backbone, and perform discrete wavelet transform on the feature map passing through the 3rd convolutional block to obtain low-frequency approximation coefficients and high-frequency detail coefficients; Perform adaptive weighted fusion on the low-frequency approximation coefficients and the high-frequency detail coefficients to obtain a multi-scale feature map, and input the multi-scale feature map into the 4th convolutional block to obtain an enhanced feature representation; Add a spatial pyramid pooling module after the last convolutional block of the convolutional neural network backbone; perform multi-scale pooling on the output feature map of the last convolutional block to obtain feature vectors of different scales; Concatenate the feature vectors of different scales to obtain a multi-scale feature representation, and perform dimensionality reduction on the multi-scale feature representation through 1x1 convolution to obtain a compressed feature vector; Construct a fully connected layer network, including two fully connected layers and a Dropout layer, and input the compressed feature vector into the fully connected layer network to output a prediction result; Use the cross-entropy loss function to calculate the error between the prediction result and the true label to obtain a loss value; optimize the network parameters through the backpropagation algorithm to obtain updated network parameters until the loss value converges or reaches a preset number of iterations to obtain an initial humanoid recognition model; Input a new preprocessed image into the initial humanoid recognition model, and after being processed by all convolutional blocks, the discrete wavelet transform layer and the spatial pyramid pooling module, obtain a local spatial depth feature map, and perform global average pooling operation on the local spatial depth feature map to obtain a target feature vector.

5. The low-power monitoring wake-up method based on humanoid recognition according to claim 1, wherein Classify and perform weighted voting on the target feature vector to obtain a preliminary humanoid recognition result, and optimize the recognition results of consecutive frames to obtain a humanoid target tracking sequence, including: Perform normalization processing on the target feature vector to obtain a normalized feature vector, and input the normalized feature vector into multiple pre-configured kernel extreme learning machines to obtain multiple sets of classification results; Perform weighted averaging on the multiple sets of classification results to obtain a comprehensive classification score, and perform binarization processing on the comprehensive classification score according to a preset classification threshold to obtain a preliminary humanoid recognition result; Perform time series analysis on the preliminary humanoid recognition results of consecutive multiple frames to obtain a time consistency score of the humanoid target, and correct the single-frame recognition result based on the time consistency score to obtain a time-smoothed recognition result; Perform spatial connectivity analysis on the time-smoothed recognition result to obtain a spatial consistency score, and combine the spatial consistency score and the time-smoothed recognition result to obtain a spatio-temporally consistent humanoid region; Extract the bounding box of the spatio-temporally consistent humanoid region to obtain the position and size information of the humanoid target, and construct a target state vector based on the position and size information to obtain an initial tracking state; Use a Kalman filter to predict the initial tracking state to obtain the predicted position of the next frame, and associate the predicted position with the recognition result of the next frame to obtain an updated tracking state; Perform motion consistency constraints on the updated tracking status, eliminate abnormal tracking results that do not conform to the motion pattern, obtain an optimized tracking sequence, and smooth the optimized tracking sequence to obtain a human target tracking sequence.

6. The low-power monitoring wake-up method based on humanoid recognition according to claim 1, characterized in that Perform incremental learning and model update based on the human target tracking sequence to obtain a target human recognition model, and generate a multi-level wake-up control strategy based on the target human recognition model and the low-power working scheme, including: Crop and align the target images in the human target tracking sequence to obtain standardized target samples, and perform feature extraction on the standardized target samples to obtain a new sample feature set; Compare the new sample feature set with the historical sample feature library to obtain a sample novelty evaluation result, and screen out an incremental learning sample set according to the sample novelty evaluation result; Fine-tune the last layer of the initial human recognition model based on the incremental learning sample set to obtain updated model parameters, and perform weighted fusion of the updated model parameters and the original model parameters to obtain a temporarily updated model; Perform performance evaluation on the temporarily updated model, calculate the recognition accuracy and computational complexity to obtain model update evaluation indicators, and decide whether to accept the update according to the model update evaluation indicators to obtain a target human recognition model; Set multiple confidence thresholds based on the target human recognition model to obtain multi-level human recognition judgment criteria, and quickly screen the real-time image according to the multi-level human recognition judgment criteria to obtain a preliminary wake-up trigger signal; Combine the scene activity index in the low-power working scheme to adjust the weight of the preliminary wake-up trigger signal to obtain a scene-adaptive wake-up score, and set multi-level wake-up thresholds based on the scene-adaptive wake-up score; According to the multi-level wake-up thresholds, divide the device wake-up process into multiple energy consumption levels to obtain a stepped wake-up strategy, and optimize the energy consumption of the stepped wake-up strategy to obtain a multi-level wake-up scheme with balanced energy consumption; Combine the multi-level wake-up scheme with balanced energy consumption and the low-power working scheme to generate a multi-level wake-up control strategy.

7. A low-power monitoring and wake-up device based on humanoid recognition, characterized in that For performing the steps of the method according to any one of claims 1 to 6, the device includes: A preprocessing module for downsampling and graying the original image sequence collected by the low-power monitoring device to obtain a preprocessed image and performing background difference to obtain a preliminary target mask; A calculation module for extracting human features from the preliminary target mask to obtain candidate human regions and calculating shape descriptors to obtain human region pseudo-labels; A training module for training an initial human recognition model based on the human region pseudo-labels and performing multi-scale feature extraction to obtain a local spatial depth feature map and a target feature vector; A classification module for classifying and weighted voting on the target feature vector to obtain a preliminary human recognition result, and optimizing the recognition results of consecutive frames to obtain a human target tracking sequence; An adjustment module, configured to calculate a scene activity index according to the humanoid target tracking sequence, generate a multi-mode adaptive monitoring strategy, and dynamically adjust the low-power working scheme of the low-power monitoring device based on the multi-mode adaptive monitoring strategy; A generation module, configured to perform incremental learning and model update according to the humanoid target tracking sequence to obtain a target humanoid recognition model, and generate a multi-level wake-up control strategy based on the target humanoid recognition model and the low-power working scheme.

8. A computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Video recording method and equipment of low-power-consumption video recording system

    CN118509702A