Monitoring method and device, nonvolatile storage medium and electronic device

By acquiring facial features in the monitoring system and using an ultra-wideband positioning chip to determine the location of target objects within the overall monitoring area, the problem of not being able to determine whether there are monitored objects in other areas in the existing technology is solved, and the efficient utilization of monitoring resources is achieved.

CN116778554BActive Publication Date: 2026-07-21TIANYI TELECOM TERMINALS
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANYI TELECOM TERMINALS
Filing Date
2023-06-19
Publication Date
2026-07-21

Smart Images

  • Figure CN116778554B_ABST
    Figure CN116778554B_ABST
Patent Text Reader

Abstract

The application discloses a kind of monitoring method and device, nonvolatile storage medium, electronic equipment.Therein, the method comprises: obtaining the face feature of monitoring object in current monitoring area, according to face feature, the target feature of monitoring object is identified;According to target feature, the target object indicated by target search instruction is searched in current monitoring area;In the case where target object is not found in current monitoring area, it is determined whether there is monitoring object in other monitoring area except current monitoring area in overall monitoring area;In the case where there is monitoring object in other monitoring area, target object is searched in other monitoring area.The application solves the technical problem of waste of monitoring resources caused by the fact that, in the related monitoring method, it cannot be determined whether there is monitoring object in other monitoring area except current monitoring area in overall monitoring area, and then it cannot search target object in other monitoring area in the case where there is no target object in current monitoring area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and more specifically, to a monitoring method and device, a non-volatile storage medium, and an electronic device. Background Technology

[0002] The development of artificial intelligence technology has made security monitoring of public areas via cameras more intelligent and efficient. This includes the following aspects: 1. Facial recognition technology: Facial images captured by cameras can be quickly and accurately recognized using algorithms such as deep learning, and compared with information stored in a database to achieve real-time monitoring, early warning, and tracing; 2. Behavioral analysis technology: In addition to facial recognition, video analysis algorithms can be used to detect and analyze behavioral characteristics, such as abnormal lingering, running, and pushing, enabling timely alarms and appropriate measures to be taken when abnormal situations occur.

[0003] In practical applications, there may be situations where there is no object being monitored within the area currently being monitored by the camera, but the camera continues to monitor the current monitoring area. Such situations will result in a significant waste of monitoring resources.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides a monitoring method and apparatus, a non-volatile storage medium, and an electronic device to at least solve the technical problem of wasted monitoring resources caused by the inability to determine whether a monitoring object exists in other monitoring areas besides the current monitoring area within the overall monitoring area in related monitoring methods, thus making it impossible to search for the target object in other monitoring areas when the target object does not exist in the current monitoring area.

[0006] According to one aspect of the embodiments of this application, a monitoring method is provided, comprising: acquiring facial features of a monitored object within a current monitoring area, and identifying target features of the monitored object based on the facial features; searching for a target object indicated by a target search instruction within the current monitoring area based on the target features; if no target object is found within the current monitoring area, determining whether a monitored object exists in other monitoring areas outside the current monitoring area within the overall monitoring area; and if a monitored object exists in other monitoring areas, searching for the target object in those other monitoring areas.

[0007] Optionally, obtaining facial features of monitored objects within the current monitoring area includes: obtaining monitoring video of the monitored objects within a preset duration; splitting the monitoring video into multiple frames for extracting facial features; selecting a target image with facial clarity within a first preset range from the multiple frames; inputting the target image into a trained facial feature extraction model; and outputting facial features through the facial feature extraction model, wherein the facial features include at least one of the following: eye features, nose features, and mouth features.

[0008] Optionally, identifying the target features of the monitored object based on facial features includes: inputting facial features into a first classifier in a trained convolutional neural network model to determine the gender of the monitored object; inputting facial features into a second classifier in a trained convolutional neural network model to determine the age of the monitored object; inputting facial features into a third classifier in a trained convolutional neural network model to determine the expression of the monitored object; and determining the target features of the monitored object based on the gender, age, and expression of the monitored object.

[0009] Optionally, if a target object is found within the current monitoring area, the target features corresponding to the target object are sent to the target database storing ID card information to identify the target object's identity information.

[0010] Optionally, if no target object is found in the current monitoring area, determine whether other monitoring areas outside the current monitoring area include the monitored object. This includes: obtaining anchor nodes with known locations and tag nodes with locations to be measured in other monitoring areas, wherein there is a correspondence between the tag nodes and the monitored object; measuring n distances between each of the n anchor nodes and the tag node, where n is a positive integer greater than 1; determining n circles by taking each of the n anchor nodes as the center and the distance between the anchor node and the tag node as the radius; and determining that other monitoring areas include the monitored object if there are intersections between the n circles, and determining the location of the monitored object corresponding to the tag node based on the intersections.

[0011] Optionally, if other monitoring areas do not contain the monitored object, the presence of the monitored object in other monitoring areas is determined at preset time intervals until the monitored object is found in other monitoring areas.

[0012] According to another aspect of the embodiments of this application, a monitoring device is also provided, comprising: an identification module, configured to acquire facial features of a monitored object within the current monitoring area, and identify target features of the monitored object based on the facial features; a first search module, configured to search for a target object indicated by a target search instruction within the current monitoring area based on the target features; a determination module, configured to determine whether a monitored object exists in other monitoring areas outside the current monitoring area if no monitored object is found within the current monitoring area; and a second search module, configured to search for the target object in other monitoring areas if a monitored object exists in other monitoring areas.

[0013] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, the storage medium including a stored program, wherein the program controls the device where the storage medium is located to execute the above monitoring method when it runs.

[0014] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the monitoring method described above is executed when the program is running.

[0015] In this embodiment, the method involves acquiring the facial features of the monitored object within the current monitoring area and identifying the target features of the monitored object based on the facial features; searching for the target object indicated by the target search instruction within the current monitoring area based on the target features; if no target object is found within the current monitoring area, determining whether a monitored object exists in other monitoring areas within the overall monitoring area besides the current monitoring area; and if a monitored object exists in other monitoring areas, searching for the target object in those other monitoring areas. By utilizing the ultra-wideband positioning chip in the camera to determine whether other monitoring areas within the overall monitoring area besides the current monitoring area include a monitored object, and searching for the target object in those other monitoring areas if a monitored object exists, the method avoids the situation where no monitored object exists within the current monitoring area. This achieves the technical effect of fully utilizing monitoring resources and solves the technical problem of wasted monitoring resources caused by the inability to determine whether a monitored object exists in other monitoring areas within the overall monitoring area besides the current monitoring area in related monitoring methods, thus making it impossible to search for the target object in other monitoring areas if no target object exists in the current monitoring area. Attached Figure Description

[0016] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0017] Figure 1This is a flowchart of a monitoring method according to an embodiment of this application;

[0018] Figure 2 This is a flowchart of another monitoring method according to an embodiment of this application;

[0019] Figure 3 This is a structural diagram of a monitoring device according to an embodiment of this application;

[0020] Figure 4 This is a hardware structure block diagram of a computer terminal (or electronic device) according to an embodiment of the present application for a monitoring method. Detailed Implementation

[0021] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0023] According to an embodiment of this application, a method embodiment for monitoring is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0024] Figure 1 This is a flowchart of a monitoring method according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:

[0025] Step S102: Obtain the facial features of the monitored objects within the current monitoring area, and identify the target features of the monitored objects based on the facial features.

[0026] According to some optional embodiments of this application, facial features are obtained using deep learning methods, specifically including the following steps:

[0027] 1. Data Acquisition and Preprocessing: Collect a large amount of facial image data and preprocess it, including cropping, resizing, and grayscale conversion.

[0028] 2. Construct a neural network model: Select a suitable deep learning model, such as a convolutional neural network or a residual network, and modify and optimize it according to the actual situation.

[0029] 3. Training the model: The preprocessed data is input into the neural network for training, and the weight parameters are continuously updated to improve the classification accuracy, and finally a trained neural network model is obtained.

[0030] 4. Feature extraction: Using a trained neural network model, newly input unknown samples are converted into fixed-length vectors to represent the information contained in the sample. These vectors are then used as the obtained facial features.

[0031] 5. Application Scenarios: The acquired facial features can be applied to various practical scenarios, such as recognition and verification. It should be noted that this embodiment identifies the target features of the monitored object based on facial features, such as facial expressions, age, and gender.

[0032] Optionally, machine learning algorithms or deep learning neural network models are used to train the extracted facial features and generate a classifier that can accurately determine different expression types (such as happy, angry, etc.). The facial images obtained from the region of interest in the new input image are input into the trained model for prediction testing to determine whether the face has certain specific reactions that can be described as basic behaviors such as "happy" and "angry". Finally, the facial expression features of the monitored object are output.

[0033] Optionally, data already labeled with gender attributes is used as training samples. A model is built using a classifier learning algorithm (such as SVM or KNN nearest neighbor algorithm), and a corresponding threshold is set according to the selected classifier to determine the category (i.e., male or female) of the test sample. Testing and evaluation: Testing is performed using unknown attribute labels (i.e., whether the image is male / female), and accuracy, recall, and F1-score are calculated to evaluate the recognition effect. If the recognition effect meets the requirements, the trained classification model is obtained. The extracted facial features are output to the trained classification model, which then outputs the gender characteristics of the monitored object.

[0034] Step S104: Search for the target object indicated by the target search instruction within the current monitoring area based on the target characteristics.

[0035] For example, if a target search command instructs the camera to search for a male aged 18-25 with a smiling expression, the camera will receive the target search command and, in response, search for the aforementioned target within the current monitoring area.

[0036] Step S106: If no target object is found in the current monitoring area, determine whether there is a monitoring object in other monitoring areas outside the current monitoring area in the overall monitoring area.

[0037] In some optional embodiments of this application, when no target object is found within the current monitoring area, the ultra-wideband (UWB) positioning chip in the camera is used to determine whether a monitoring object exists in other monitoring areas outside the current monitoring area within the overall monitoring area. UWB positioning technology is a high-precision positioning technology based on radio waves, characterized by extremely high time resolution and bandwidth. This technology utilizes the time delay difference of ultra-short pulse signals propagating in space to achieve position measurement, offering advantages such as low power consumption, good anti-interference performance, and high accuracy. The UWB positioning chip is a positioning chip based on UWB technology, suitable for precise positioning in indoor and urban environments. This chip employs a unique signal processing algorithm, enabling it to accurately calculate the distance between the receiver and transmitter in complex environments such as high noise and multipath interference, and obtain the target position through triangulation. The UWB positioning chip has the following advantages: 1. High precision: achieving centimeter-level accuracy; 2. Strong anti-interference: maintaining good stability even under high noise and multipath interference conditions; 3. Wide applicability: adaptable to accurate positioning in different scenarios and environments; 4. Relatively low cost: more reasonable compared to other technologies.

[0038] Step S108: If a monitored object exists in other monitored areas, search for the target object in those other monitored areas.

[0039] As some optional embodiments of this application, if there are monitorable people in other monitoring areas, the target object indicated by the target search instruction is searched in other monitoring areas, such as the male aged 18-25 years old with a smiling expression as exemplified in step S104.

[0040] Based on the above steps, by using the ultra-wideband positioning chip in the camera, it is determined whether other monitoring areas outside the current monitoring area include the monitoring object. If the monitoring object exists in other monitoring areas, the target object is searched in other monitoring areas. This achieves the goal of avoiding the situation where the monitoring object does not exist in the current monitoring area, thus realizing the technical effect of making full use of monitoring resources.

[0041] According to some optional embodiments of this application, obtaining facial features of a monitored object within the current monitoring area includes the following steps: obtaining monitoring video of the monitored object within a preset duration; splitting the monitoring video into multiple frames for extracting facial features; selecting a target image whose facial clarity falls within a first preset range from the multiple frames; inputting the target image into a trained facial feature extraction model, and outputting facial features through the facial feature extraction model, wherein the facial features include at least one of the following: eye features, nose features, and mouth features.

[0042] Understandably, facial recognition technology still has some limitations, such as being sensitive to factors like lighting and occlusion. Therefore, in practical applications, it is necessary to consider multiple factors to improve accuracy, which means that the acquired images need to be preprocessed. Generally, the following preprocessing operations are required:

[0043] 1. Image Denoising: Using image filtering algorithms (such as Gaussian filtering, median filtering, etc.) to reduce or eliminate noise in the image; 2. Image Enhancement: Improving the contrast, brightness, and clarity of the image through methods such as histogram equalization, making facial features more prominent; 3. Face Detection: Using face detection algorithms (such as Haar Cascade, HOG+SVM, etc.) to find all faces in the image and cut them out as separate images; 4. Face Alignment: Rotating and scaling all detected faces according to their eye positions to ensure they face the same direction and are the same size, facilitating subsequent processing; 5. Illumination Normalization: Using methods such as grayscale stretching and white balance adjustment to solve color distortion problems caused by different light sources.

[0044] In some optional embodiments of this application, the target features of a monitored object can be identified based on facial features through the following methods: inputting facial features into a first classifier in a trained convolutional neural network model to determine the gender of the monitored object; inputting facial features into a second classifier in a trained convolutional neural network model to determine the age of the monitored object; inputting facial features into a third classifier in a trained convolutional neural network model to determine the expression of the monitored object; and determining the target features of the monitored object based on its gender, age, and expression.

[0045] As some optional embodiments of this application, the classifier in a convolutional neural network model is a function used to map input data to different categories or labels. Typically, a classifier can be viewed as the output layer of the last layer of a convolutional neural network, where each node represents a possible category or label. During training, the classifier's performance is optimized by adjusting the weights and biases between neurons, enabling it to correctly classify new data. Common classification algorithms include softmax regression and support vector machines (SVM).

[0046] Softmax regression is a classification model used to map an input vector to a probability distribution of multiple classes. It is commonly used as the output layer of a convolutional neural network, where each node corresponds to a possible class. In softmax regression, a score or weight for each node (i.e., each possible class) is first calculated and converted to a non-negative value using an exponential function. These values ​​are then standardized by dividing by the sum of the scores of all nodes, making the sum equal to 1, thus obtaining a probability estimate for each class.

[0047] Support Vector Machine (SVM) is a commonly used classification and regression algorithm. Its workflow is as follows: 1. Data preprocessing: including data cleaning, missing value imputation, and feature selection; 2. Feature transformation: transforming the raw data into feature vectors in a high-dimensional space; 3. Splitting the training and test sets: dividing the data into training and test sets, for example, using cross-validation; 4. Selecting the kernel function: choosing an appropriate kernel function based on the specific problem, such as a linear kernel, polynomial kernel, or radial basis function kernel; 5. Training the model: training the model using the training data and improving model performance by adjusting hyperparameters; 6. Predicting results: evaluating the model using test data and calculating metrics such as accuracy and precision to evaluate model performance; 7. Parameter optimization and adjustment: further optimizing or adjusting SVM parameters based on the actual application scenario to achieve better classification results; 8. Model deployment: deploying the optimally configured SVM algorithm in the production or living environment.

[0048] In some optional embodiments of this application, when a target object is found in the current monitoring area, the target features corresponding to the target object are sent to the target database storing ID card information in order to identify the identity information of the target object.

[0049] In some optional embodiments, if no target object is found in the current monitoring area, it is determined whether other monitoring areas outside the current monitoring area include the monitored object. This is achieved by the following method: obtaining anchor nodes with known locations and tag nodes with locations to be measured in other monitoring areas, wherein there is a correspondence between the tag nodes and the monitored object; measuring n distances between each of the n anchor nodes and the tag node, where n is a positive integer greater than 1; determining n circles by taking each of the n anchor nodes as the center and the distance between the anchor node and the tag node as the radius; if there are intersections between the n circles, it is determined that other monitoring areas include the monitored object, and the location of the monitored object corresponding to the tag node is determined based on the intersections.

[0050] According to some preferred embodiments of this application, anchor nodes with known locations and tag nodes with unknown locations in other areas are obtained. Assuming that the positioning area includes n anchor nodes, the location coordinates of the tag node are (x, y), and the location coordinates of the i-th anchor node are (xi, yi), the distance ri between the i-th anchor node and the tag node is measured. On a two-dimensional plane, n circles are drawn with different anchor nodes as centers and the corresponding measured distances as radii. The intersection of these circles is the location of the tag node.

[0051] Ultra-wideband (UWB) positioning chips are positioning chips based on UWB technology, capable of precise positioning in indoor and urban environments. This chip employs a unique signal processing algorithm, enabling it to accurately calculate the distance between the receiver and transmitter in complex environments with high noise and multipath interference, and obtain the target position through triangulation. UWB positioning chips offer the following advantages: 1. High precision: achieving centimeter-level accuracy; 2. Strong anti-interference capability: maintaining good stability even under high noise and multipath interference conditions; 3. Wide applicability: adaptable to various scenarios and environments for accurate positioning; 4. Relatively low cost: more cost-effective compared to other technologies.

[0052] The workflow of an ultra-wideband positioning chip mainly consists of three steps: sending, receiving, and calculation. Specifically:

[0053] 1. Transmission: In a positioning system, at least one tag or node needs to transmit a signal. These signals can be different types of signals, such as electromagnetic waves or sound waves. Ultra-wideband technology typically uses short-pulse signals for transmission, which offer high controllability and accuracy, making them ideal for positioning applications.

[0054] 2. Reception: Once a signal is sent, nearby receivers begin accepting it. Each receiver records its arrival timestamp and sends it to the central processor. Because ultra-wideband technology offers high precision and microsecond-level time measurement capabilities, it can accurately measure the distance differences between nodes.

[0055] 3. Calculation: Data processing is performed in a centralized processor to determine the location of tags or nodes. By parsing arrival timestamps, the relative distance difference is calculated, and the triangle rule is used to determine the object's position. Simultaneously, other influencing factors such as multipath effects and obstacle interference must be considered to improve accuracy and resist error interference.

[0056] Optionally, if other monitoring areas do not contain the monitored object, the presence of the monitored object in other monitoring areas is determined at preset time intervals until the monitored object is found in other monitoring areas.

[0057] Figure 2 This is a flowchart of another monitoring method according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:

[0058] Step S202: Obtain the facial features of the monitored objects within the current monitoring area of ​​the camera, and use the trained facial feature extraction model to extract the facial features of the monitored objects within the monitoring area.

[0059] Optionally, the facial feature extraction model is trained using the following method:

[0060] 1. Obtain sample images including human faces;

[0061] 2. Input the sample image into the face feature extraction model, and the backbone network in the face feature extraction model extracts the face features of the sample image. For example, the face sample image is used as the input of the backbone network, and the output is the corresponding face feature. Here, the face feature refers to a d-dimensional vector, where d is the face feature dimension.

[0062] 3. Obtain the covariance matrix corresponding to the facial features, and obtain the first loss function based on the covariance matrix;

[0063] 4. Classify and recognize facial features, and obtain a second loss function based on the classification and recognition results;

[0064] 5. Adjust the model parameters of the face feature extraction model according to the first loss function and the second loss function until the training termination condition is met to obtain the target face feature extraction model.

[0065] Step S204: Identify the features of the monitored object based on its facial features. The features of the monitored object include, but are not limited to, age, gender, expression, and posture.

[0066] For example, if there are 50 faces in the current monitoring area, then 21 people are in the age range of [18-25], 16 people are in the age range of [26-45], 8 people are in the age range of [46-65], and the remaining 5 people are in other age ranges.

[0067] Step S206: Set a monitoring strategy, for example, to find people aged 18-25 in the current monitoring area, or to find male people in the current monitoring area, or to find people with a smiling expression in the current monitoring area.

[0068] Different monitoring strategies can be selected based on different usage scenarios. For example, when tracking a specific group of people, if only the gender, age range, and general facial features of the specific group are known, these features can be input into the monitoring strategy. If a suspicious person appears in the monitored area, he / she can be identified.

[0069] Step S208: Use UWB to locate people within the target area (overall monitoring area) and search for other areas within the target area but not within the current monitoring area to see if there are any monitorable objects.

[0070] For example, if the target range is a 360° all-around area within 30 meters, and the current monitoring area is a 60° area within the 360° range, then the location of people in other areas can be located using ultra-wideband positioning.

[0071] Step S210: If there is no target monitoring object specified by the monitoring strategy in the current monitoring area, but there is a monitoring object in other monitoring areas, turn the camera to other monitoring areas.

[0072] Optionally, if there are no monitored objects in the current monitoring area whose age range is [18-25], the camera is rotated to other monitoring areas outside the current monitoring area to monitor whether the other monitoring areas include monitored objects whose age range is [18-25].

[0073] In step S212, if the camera rotates to a new monitoring area, the next rotation strategy for the camera continues to be determined based on the content monitored by the camera and the content acquired by UWB.

[0074] Based on the steps described above, ultra-wideband human positioning allows the camera to detect people even when they are out of its reach. This enables the camera to rotate omnidirectionally, automatically positioning itself to the person's location. This avoids situations where someone is behind the camera and the camera cannot capture anything. The technology achieves the effect of enabling the camera to monitor and track all human figures and other targets in its vicinity at any time.

[0075] Figure 3 This is a structural diagram of a monitoring device according to an embodiment of this application, such as... Figure 3 As shown, the device includes:

[0076] The recognition module 30 is used to acquire the facial features of the monitored objects within the current monitoring area and to identify the target features of the monitored objects based on the facial features.

[0077] According to some optional embodiments of this application, facial features are obtained using deep learning methods, specifically including the following steps:

[0078] 1. Data Acquisition and Preprocessing: Collect a large amount of facial image data and preprocess it, including cropping, resizing, and grayscale conversion. 2. Neural Network Model Construction: Select a suitable deep learning model, such as a convolutional neural network or residual network, and modify and optimize it according to the actual situation. 3. Model Training: Input the preprocessed data into the neural network for training, and continuously update the weight parameters to improve classification accuracy, ultimately obtaining a trained neural network model. 4. Feature Extraction: Using the trained neural network model, newly input unknown samples are converted into fixed-length vectors to represent the information contained in the sample. These vectors are used as the obtained facial features. 5. Application Scenarios: Apply the obtained facial features to various practical scenarios, such as recognition and verification. In this embodiment, the target features of the monitored object are identified based on facial features, such as facial expression, age, and gender.

[0079] Optionally, machine learning algorithms or deep learning neural network models are used to train the extracted facial features and generate a classifier that can accurately determine different expression types (such as happy, angry, etc.). The facial images obtained from the region of interest in the new input image are input into the trained model for prediction testing to determine whether the face has certain specific reactions that can be described as basic behaviors such as "happy" and "angry". Finally, the facial expression features of the monitored object are output.

[0080] Optionally, data already labeled with gender attributes is used as training samples. A model is built using a classifier learning algorithm (such as SVM or KNN nearest neighbor algorithm), and a corresponding threshold is set according to the selected classifier to determine the category (i.e., male or female) of the test sample. Testing and evaluation: Testing is performed using unknown attribute labels (i.e., whether the image is male / female), and accuracy, recall, and F1-score are calculated to evaluate the recognition effect. If the recognition effect meets the requirements, the trained classification model is obtained. The extracted facial features are output to the trained classification model, which then outputs the gender characteristics of the monitored object.

[0081] The first search module 32 is used to search for the target object indicated by the target search instruction within the current monitoring area based on the target characteristics.

[0082] The determination module 34 is used to determine whether there is a monitored object in other monitoring areas outside the current monitoring area if no target object is found in the current monitoring area.

[0083] The second search module 36 is used to search for the target object in other monitoring areas when the monitored object exists in other monitoring areas.

[0084] It should be noted that the above Figure 3 The modules in can be program modules (e.g., a set of program instructions that implements a specific function) or hardware modules. For the latter, they can be represented in the following forms, but are not limited to these: each of the above modules is represented by a processor, or the functions of each of the above modules are implemented by a processor.

[0085] It should be noted that, Figure 3 Preferred embodiments of the shown examples can be found in [reference needed]. Figure 1 The relevant descriptions of the embodiments shown will not be repeated here.

[0086] Optionally, obtaining facial features of monitored objects within the current monitoring area includes: obtaining monitoring video of the monitored objects within a preset duration; splitting the monitoring video into multiple frames for extracting facial features; selecting a target image with facial clarity within a first preset range from the multiple frames; inputting the target image into a trained facial feature extraction model; and outputting facial features through the facial feature extraction model, wherein the facial features include at least one of the following: eye features, nose features, and mouth features.

[0087] Optionally, identifying the target features of the monitored object based on facial features includes: inputting facial features into a first classifier in a trained convolutional neural network model to determine the gender of the monitored object; inputting facial features into a second classifier in a trained convolutional neural network model to determine the age of the monitored object; inputting facial features into a third classifier in a trained convolutional neural network model to determine the expression of the monitored object; and determining the target features of the monitored object based on the gender, age, and expression of the monitored object.

[0088] Optionally, if a target object is found within the current monitoring area, the target features corresponding to the target object are sent to the target database storing ID card information to identify the target object's identity information.

[0089] Optionally, if no target object is found in the current monitoring area, determine whether other monitoring areas outside the current monitoring area include the monitored object. This includes: obtaining anchor nodes with known locations and tag nodes with locations to be measured in other monitoring areas, wherein there is a correspondence between the tag nodes and the monitored object; measuring n distances between each of the n anchor nodes and the tag node, where n is a positive integer greater than 1; determining n circles by taking each of the n anchor nodes as the center and the distance between the anchor node and the tag node as the radius; and determining that other monitoring areas include the monitored object if there are intersections between the n circles, and determining the location of the monitored object corresponding to the tag node based on the intersections.

[0090] Optionally, if other monitoring areas do not contain the monitored object, the presence of the monitored object in other monitoring areas is determined at preset time intervals until the monitored object is found in other monitoring areas.

[0091] Figure 4 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a monitoring method is shown. Figure 4 As shown, the computer terminal 40 (or mobile device 40) may include one or more processors 402 (shown as 402a, 402b, ..., 402n in the figure) (processor 402 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 404 for storing data, and a transmission module 406 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 4 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 40 may also include... Figure 4 The more or fewer components shown, or having the same Figure 4 The different configurations shown.

[0092] It should be noted that the aforementioned one or more processors 402 and / or other data processing circuitry are generally referred to herein as "data processing circuitry". This data processing circuitry may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 40 (or mobile device). As involved in the embodiments of this application, this data processing circuitry serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0093] The memory 404 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the monitoring method in this embodiment. The processor 402 executes various functional applications and data processing by running the software programs and modules stored in the memory 404, thereby realizing the aforementioned monitoring method. The memory 404 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 404 may further include memory remotely located relative to the processor 402, and these remote memories can be connected to the computer terminal 40 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0094] The transmission module 406 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 40. In one example, the transmission module 406 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission module 406 may be a radio frequency (RF) module, used for wireless communication with the Internet.

[0095] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 40 (or mobile device).

[0096] It should be noted here that, in some optional embodiments, the above... Figure 4The computer device (or electronic device) shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 4 This is only one instance of a particular specific instance, and is intended to illustrate the types of components that may exist in the aforementioned computer equipment (or electronic equipment).

[0097] It should be noted that, Figure 4 The electronic device shown is used to perform Figure 1 The monitoring method shown above, therefore, the relevant explanations in the execution method of the above commands also apply to this electronic device, and will not be repeated here.

[0098] This application also provides a non-volatile storage medium, which includes a stored program, wherein the program, when running, controls the device where the storage medium is located to execute the above monitoring method.

[0099] A non-volatile storage medium performs the following functions: acquires the facial features of the monitored objects within the current monitoring area, and identifies the target features of the monitored objects based on the facial features; searches for the target object indicated by the target search command within the current monitoring area based on the target features; if the target object is not found within the current monitoring area, determines whether there are monitored objects in other monitoring areas outside the current monitoring area within the overall monitoring area; if there are monitored objects in other monitoring areas, searches for the target object in those other monitoring areas.

[0100] This application also provides an electronic device, including a memory and a processor, wherein the processor is used to run a program stored in the memory, and the program executes the above-described monitoring method during runtime.

[0101] The processor is used to run programs that perform the following functions: acquire facial features of monitored objects within the current monitoring area, and identify target features of monitored objects based on facial features; search for the target object indicated by the target search instruction within the current monitoring area based on the target features; if the target object is not found within the current monitoring area, determine whether there are monitored objects in other monitoring areas outside the current monitoring area within the overall monitoring area; if there are monitored objects in other monitoring areas, search for the target object in those other monitoring areas.

[0102] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0103] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0104] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0105] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0106] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0107] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0108] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A monitoring method, characterized in that, include: Obtain the facial features of the monitored objects within the current monitoring area, and identify the target features of the monitored objects based on the facial features; Based on the target characteristics, search for the target object indicated by the target search command within the current monitoring area; If the target object is not found in the current monitoring area, determine whether the target object exists in other monitoring areas outside the current monitoring area within the overall monitoring area. This includes: acquiring anchor nodes with known locations and tag nodes with locations to be measured in the other monitoring areas, wherein the tag nodes correspond to the target object; measuring n distances between each of the n anchor nodes and the tag node, where n is a positive integer greater than 1; determining n circles by taking each of the n anchor nodes as the center and the distance between the anchor node and the tag node as the radius; and determining that the other monitoring area includes the target object if there is an intersection point between the n circles, and determining the location of the target object corresponding to the tag node based on the intersection point. If the target object is present in the other monitored area, the camera will be rotated to that other monitored area to locate the target object.

2. The method according to claim 1, characterized in that, Obtain facial features of monitored objects within the current monitoring area, including: Obtain monitoring video of the monitored object within a preset duration; The surveillance video is split into multiple frames for extracting the facial features; Select a target image whose facial clarity falls within a first preset range from the multiple frames of images; The target image is input into a trained facial feature extraction model, and the facial features are output through the facial feature extraction model. The facial features include at least one of the following: eye features, nose features, and mouth features.

3. The method according to claim 1, characterized in that, Identifying the target features of the monitored object based on the facial features includes: The facial features are input into the first classifier of the trained convolutional neural network model to determine the gender of the monitored object; The facial features are input into the second classifier in the trained convolutional neural network model to determine the age of the monitored object; The facial features are input into the third classifier in the trained convolutional neural network model to determine the expression of the monitored object; The target characteristics of the monitored object are determined based on the object's gender, age, and facial expression.

4. The method according to claim 3, characterized in that, The method further includes: If the target object is found within the current monitoring area, the target features corresponding to the target object are sent to the target database storing ID card information to identify the identity information of the target object.

5. The method according to claim 1, characterized in that, The method further includes: If the monitored object is not included in the other monitored areas, the presence of the monitored object in the other monitored areas is determined at preset time intervals until the monitored object is present in the other monitored areas.

6. A monitoring device, characterized in that, include: The recognition module is used to acquire the facial features of the monitored objects within the current monitoring area, and to identify the target features of the monitored objects based on the facial features; The first search module is used to search for the target object indicated by the target search instruction within the current monitoring area based on the target characteristics. The determination module is used to determine whether the monitored object exists in other monitoring areas outside the current monitoring area when the target object is not found in the current monitoring area. This includes: acquiring anchor nodes with known locations and tag nodes with locations to be measured in the other monitoring areas, wherein the tag nodes correspond to the monitored object; measuring n distances between each of the n anchor nodes and the tag node, where n is a positive integer greater than 1; determining n circles by taking each of the n anchor nodes as the center and the distance between the anchor node and the tag node as the radius; and determining that the other monitoring area includes the monitored object if there is an intersection point between the n circles, and determining the location of the monitored object corresponding to the tag node based on the intersection point. The second search module is used to rotate the camera to the other monitoring area when the monitored object exists in the other monitoring area, so as to search for the target object in the other monitoring area.

7. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored program, wherein, when the program is executed, it controls the device containing the non-volatile storage medium to perform the monitoring method according to any one of claims 1 to 5.

8. An electronic device, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, performs the monitoring method according to any one of claims 1 to 5.