Intelligent talkback terminal face recognition method based on external TF card and related equipment

The external TF card extends the functions of low-configured intelligent intercom devices and dynamically loads the face recognition function, solving the problem that the device is difficult to run high-complex functions, and achieving function expansion and cost savings.

CN120071452AActive Publication Date: 2025-05-30FUJIAN RUIYUNLIAN INTELLIGENT TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510555450.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Existing low-configuration intelligent intercom devices are difficult to run high-complexity functions such as face recognition, and the hardware upgrade cost is high and incompatible.

Method used

Extend device functions through external TF card, dynamically load or unload facial recognition functions, without hardware upgrades.

Benefits of technology

It realizes that without upgrading the device hardware, expanding the facial recognition function, saving costs, and avoiding the use of internal storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071452A_ABST
    Figure CN120071452A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent talkback terminal face recognition method based on an external TF card and related equipment, and is applied to the technical field of data processing. The method comprises the following steps: processing biological characteristic information of a target user to generate identity attribute information of the target user; processing the target TF card information to generate a target identity recognition model; processing the to-be-accessed application information, and generating application access purpose information of the target user and access permission information of the target user; processing the application access purpose information of the target user and the access permission information of the target user to generate access permission information of the target application; processing the access permission information of the target application and the identity information of the target user based on the target identity recognition model, and generating video image information of the target intercom device; and processing the emotion information of the target user and the video image information of the target user based on the video image information of the target intercom device to generate target event information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a face recognition method for an intelligent intercom terminal based on an external TF card and related devices. Background Art

[0002] In existing intelligent intercom devices, due to the limitations of processing power and storage space, low - configuration devices are difficult to run highly complex functions such as face recognition. Especially in scenarios where face recognition requires a large amount of storage space to load algorithm models, the hardware limitations of low - configuration devices make it difficult to meet the application requirements. Currently, most low - configuration intelligent intercom devices on the market still rely on traditional recognition methods (such as password input or card recognition) and cannot support more advanced biometric recognition, which brings limitations to the user experience. And if the hardware of these devices is upgraded, the cost will increase significantly, and the original devices are difficult to be compatible with such upgrades. Therefore, there is a lack of a method in the prior art that neither relies on internal storage upgrades nor can flexibly expand the face recognition function.

[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0004] The purpose of the present application is to provide a face recognition method for an intelligent intercom terminal based on an external TF card, related devices and systems, which can at least overcome the problems existing in the prior art to a certain extent. By externally connecting a TF card to expand the functions of low - configuration intelligent intercom devices, the device can dynamically load or unload the face recognition function according to requirements without hardware upgrade, saving costs. By using an external TF card, the device does not need to occupy the limited internal storage space, solving the problem that low - configuration devices cannot support large - scale algorithm models. Users only need to insert a TF card to use the face recognition function, which is simple and convenient, and can be unloaded or removed at any time when the function is not needed.

[0005] Other features and advantages of the present application will become apparent through the following detailed description, or will be learned in part through the practice of the present invention.

[0006] According to one aspect of the present application, a face recognition method for an intelligent intercom terminal based on an external TF card is provided, including: obtaining target TF card information, information of the application to be accessed, biometric information of the target user, and video image information of the target user; processing the biometric information of the target user to generate identity attribute information of the target user, where the identity attribute information of the target user includes the identity information of the target user and the emotional information of the target user; processing the target TF card information to generate a target identity recognition model; processing the information of the application to be accessed to generate application access purpose information of the target user and access permission information of the target user; processing the application access purpose information of the target user and the access permission information of the target user to generate access permission information of the target application; processing the access permission information of the target application and the identity information of the target user based on the target identity recognition model to generate video image information of the target intercom device; processing the emotional information of the target user and the video image information of the target user based on the video image information of the target intercom device to generate target event information, where the target event information is used to represent the voice call information between the target user and other users, and the other users are generated based on the video image information of the target intercom device.

[0007] Another aspect of the present application is an intelligent intercom terminal face recognition device based on an external TF card, characterized by including: an acquisition module for obtaining target TF card information, information of the application to be accessed, biometric information of the target user, and video image information of the target user; a processing module for processing the biometric information of the target user to generate identity attribute information of the target user, where the identity attribute information of the target user includes the identity information of the target user and the emotional information of the target user; processing the target TF card information to generate a target identity recognition model; processing the information of the application to be accessed to generate application access purpose information of the target user and access permission information of the target user; processing the application access purpose information of the target user and the access permission information of the target user to generate access permission information of the target application; processing the access permission information of the target application and the identity information of the target user based on the target identity recognition model to generate video image information of the target intercom device; processing the emotional information of the target user and the video image information of the target user based on the video image information of the target intercom device to generate target event information, where the target event information is used to represent the voice call information between the target user and other users, and the other users are generated based on the video image information of the target intercom device.

[0008] According to another aspect of the present application, an electronic device, characterized in that it includes: a first processor; and a memory for storing executable instructions of the first processor; wherein, the first processor is configured to execute the above-mentioned face recognition method for an intelligent intercom terminal based on an external TF card by executing the executable instructions.

[0009] According to another aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a second processor, the above-mentioned face recognition method for an intelligent intercom terminal based on an external TF card is implemented.

[0010] According to another aspect of the present application, there is provided a computer program product, including a computer program, characterized in that when the computer program is executed by a third processor, the above-mentioned face recognition method for an intelligent intercom terminal based on an external TF card is implemented.

[0011] The face recognition method for an intelligent intercom terminal based on an external TF card and related devices provided by the present application expand the functions of low-configuration intelligent intercom devices through an external TF card, enabling the device to dynamically load or unload the face recognition function according to requirements without hardware upgrade, saving costs. Through the external TF card, the device does not need to occupy the limited internal storage space, solving the problem that low-configuration devices cannot support large-scale algorithm models. Users only need to insert the TF card to use the face recognition function, which is simple and convenient, and can be unloaded or removed at any time when the function is not needed.

[0012] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 The flowchart showing a face recognition method for an intelligent intercom terminal based on an external TF card provided by an embodiment of the present application; Figure 2 The structural schematic diagram showing a face recognition device for an intelligent intercom terminal based on an external TF card provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0015] The following is combined with Figure 1Describe a face recognition method for an intelligent intercom terminal based on an external TF card according to an exemplary embodiment of the present application. It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard. On the contrary, the embodiments of the present application can be applied to any applicable scenario.

[0016] In one embodiment, the present application also proposes a face recognition method for an intelligent intercom terminal based on an external TF card and related devices. Figure 1 Schematically shows a flowchart of a face recognition method for an intelligent intercom terminal based on an external TF card according to an embodiment of the present application. As Figure 1 shown, this method is applied to a server and includes: S101, obtain target TF card information, information of the application to be accessed, biometric information of the target user, and video image information of the target user.

[0017] In one embodiment, in the field of intercom devices, the capacity of the target TF card is crucial. For example, for an intercom device used for large-scale event security, if it is necessary to store the face data of many security personnel to expand the face recognition function, assuming that the total capacity of the target TF card is 32 GB, 5 GB has been used, and 27 GB remains. The remaining space needs to consider not only accommodating the existing face recognition algorithm model (which may occupy several GB of space, depending on the complexity of the model), but also reserving enough space for the face data of subsequent security personnel that may be added. For example, if it is estimated that the face data and related feature information of each person occupy an average of 10 MB, then the remaining space can theoretically store about 2,700 people's data, but in practice, the cache and other temporary files occupied during system operation also need to be considered, so it is necessary to plan and use it reasonably. For some intercom devices that need to store voice records for a long time (such as important call records), the TF card capacity also determines the duration of the voice files that can be stored. Taking a common voice coding format as an example, a voice file may occupy about 1 MB of space per minute. If there is 27 GB of remaining space, then about 450 hours of voice records can be stored (without considering the influence of other factors).

[0018] In an intercom device, if a TF card in FAT32 format is used, its advantage is wide compatibility, and most intercom devices and related operating systems can recognize and read / write it. However, the maximum size of a single file cannot exceed 4GB, which may be limiting for some application scenarios that may need to store large-capacity face recognition model files (if still large after optimization) or large voice data packet files. The exFAT format supports larger file sizes and is more suitable for storing larger model or data files. However, not all older models of intercom devices can perfectly support the exFAT format, and device compatibility needs to be ensured before use. For example, in some new models of smart intercom devices, to better accommodate TF cards with large capacities and diverse formats, the system is optimized accordingly to ensure stable operation of both FAT32 and exFAT format cards. In the multi-card management scenario of intercom devices, the TF card serial number plays a significant role. For example, in an intercom device network in a large logistics park, there may be multiple TF cards distributed on devices in different areas. Through the TF card serial number, the specific location of each card in the intercom device can be accurately identified. For example, the TF card with the serial number "123456789" is installed in the intercom device at the entrance of the park, and its usage history, such as when it was inserted into the device to start working and whether there were any abnormal data read / write situations, can be recorded, facilitating the maintenance and management of the device and the TF card. At the same time, permission allocation is carried out according to the serial number. For example, a TF card with a specific serial number can only be used on intercom devices in certain designated areas to prevent the abuse of TF cards and data security issues.

[0019] Taking "Intercom Dispatching and Command Application" as an example, in the security protection scenario of large-scale events, the application name clarifies its core position among numerous intercom functions. Its version "V2.1.0" may have brought key improvements in this scenario. For example, in previous versions, when dealing with a large number of security personnel using intercoms for command and dispatch simultaneously, there may have been problems such as communication delays or untimely issuance of commands. The V2.1.0 version has improved the efficiency of command and dispatch by optimizing the communication protocol stack and algorithms. In terms of interface design, it may more intuitively display the personnel distribution and status in different security areas, facilitating commanders to make quick decisions. In the field of logistics transportation scheduling, the update of the application version may be related to adapting to new transportation management processes or hardware devices. For example, as new models of intercom devices are equipped with higher-resolution screens and more powerful processors, the V2.1.0 version has specifically optimized the interface display effect, enabling it to better display transportation routes, cargo status, etc. on the new devices, and improving the response speed of the application to ensure that dispatch commands can be promptly conveyed to drivers and warehouse personnel.

[0020] Group Management Module: In event security, the importance of the group management module is self-evident. For example, different groups with various functions and permissions can be created, such as "Indoor Venue Security Group", "Outdoor Venue Security Group", "Emergency Response Group", etc. For the "Emergency Response Group", its members are the emergency team members in each area, and the administrator's permission can be set to be able to call all members with one key and forcefully turn on the walkie-talkie microphones of all members in case of an emergency, ensuring rapid information transmission. Moreover, this group is set to a high priority, and the call will not be interfered with by other ordinary groups. Call Record and Playback Module: In the logistics transportation scenario, the call record and playback module is an important tool for resolving disputes and optimizing processes. Suppose a batch of goods is damaged during transportation. By playing back the call records between the driver and the warehouse shipping staff, as well as those with other relevant personnel during transportation, it is possible to accurately determine at which stage problems may have occurred, whether it was improper operation during loading or abnormal road conditions during transportation that caused the goods to be damaged. In enterprise production management, by playing back the call records between workers in different areas of the workshop, it is possible to analyze whether there are any communication breakdowns or operational mistakes in the production process, so as to optimize the production process. Location Tracking Module (if supported): In emergency rescue scenarios, such as mountain rescue operations, the location tracking module can display the positions of rescue personnel in real time. The command center can reasonably plan the rescue route based on the location information, avoiding repeated searches by rescue personnel or their entry into dangerous areas. At the same time, when rescue personnel encounter unexpected situations (such as being injured or getting lost), the command center can quickly locate them and dispatch nearby rescue forces for support. In the offshore operation scenario, for crew members equipped with walkie-talkies, the location tracking module combined with the ship positioning system can monitor the activity range of crew members on the ship in real time to ensure the safety of crew members. In case of emergencies such as a person falling into the water, the location of the incident can be quickly determined for rescue.

[0021] In the application of walkie-talkies in large enterprises, the role permissions of ordinary employee users may only allow the use of walkie-talkies to make voice calls with colleagues in the same department during working hours, and operations such as group creation and management are not permitted to prevent misoperations from affecting the overall communication order. In addition to normal calls, the role permissions of department supervisors can also view the call records of all employees in their department (for work supervision and understanding of work progress), and can create and manage temporary work groups within their department to facilitate project collaboration. The permissions of super administrators can comprehensively manage the walkie-talkie application system of the entire enterprise, including adding or deleting users, configuring system-level parameters (such as communication encryption methods, voice quality settings, etc.), viewing the call records of all departments, etc., to ensure the safe and efficient operation of the system. In the trial version of some walkie-talkie device applications, newly registered users may only be able to access basic voice call functions and cannot use functions such as location tracking (if the device supports it) and advanced group management. These advanced functions can only be unlocked after the user upgrades to a formal paid user and the enterprise where the user is located purchases the corresponding service package. For example, an enterprise purchases a service package including location tracking functions for some senior management personnel. When these management personnel use the walkie-talkie application, they can access and use the location tracking function after authentication to view the location information of their subordinate employees in real time and improve management efficiency.

[0022] In a project involving the use of cross-brand walkie-talkies, it is crucial for application developers to clearly list the supported walkie-talkie models. For example, a general walkie-talkie dispatching application supports models A, B, and C of brand X's walkie-talkies, as well as models D and E of brand Y's walkie-talkies, and there are specific hardware version requirements for each model. For instance, model A of brand X's walkie-talkie requires a hardware version of V3.0 or higher because the communication chip was upgraded in this version of the hardware, enabling better support for certain advanced functions of the application (such as high-definition voice calls, fast data transmission, etc.). If a walkie-talkie of this model with a version lower than V3.0 is used, problems such as degraded voice quality and malfunctioning of functions may occur. In practical applications, different brands and models of walkie-talkies may use different communication protocols. For example, some walkie-talkies use traditional analog communication protocols, while some new walkie-talkies support digital communication protocols (such as the DMR protocol). The walkie-talkie dispatching application needs to have strong communication protocol adaptation capabilities. When the application connects to a walkie-talkie using an analog communication protocol, through the built-in analog communication protocol conversion module, the application's instructions and data are converted into analog signals for transmission; when connecting to a walkie-talkie that supports a digital communication protocol, direct use of the digital communication protocol is made for efficient data interaction. At the same time, the application also needs to be able to automatically identify the type of communication protocol of the walkie-talkie to ensure seamless communication between different devices. For example, in a scenario where both old analog walkie-talkies and new digital walkie-talkies are used, the application can intelligently switch and adapt to different protocols to ensure that all walkie-talkies can be normally connected to the dispatching command system for unified management and command.

[0023] The biometric information of the target user includes the following: Facial feature information: If it is an intelligent walkie-talkie device extended based on the face recognition function, the face image data of the target user will be obtained, and the feature vector of the face will be extracted through the face recognition algorithm. For example, the facial feature vector may be a numerical vector of 128 dimensions or higher, which can represent the unique features of the user's face, such as the quantitative representation of features like the distance between the eyes, the shape of the nose, and the facial contour. Fingerprint feature information (if the device supports the fingerprint recognition function extension): The fingerprint image of the target user is obtained through the fingerprint recognition sensor, and then the feature point information of the fingerprint is extracted, such as the position, direction, and type of feature points like the ridges, valleys, and bifurcation points of the fingerprint. This feature point information forms the fingerprint feature template for subsequent fingerprint comparison and identity recognition.

[0024] On intelligent intercom devices, cameras are usually equipped. When a user approaches the device for authentication, the camera captures the user's face image and then transmits the image to the face recognition module. The face recognition module uses a pre-trained face recognition algorithm to process the image. First, face detection is performed to determine the position and pose of the face in the image, and then feature extraction is carried out to convert the face image into a feature vector. For example, a deep learning-based face recognition algorithm is adopted, such as using a convolutional neural network (CNN) to perform multi-layer convolution and pooling operations on the face image, and finally a feature vector is obtained. The fingerprint recognition sensor on the device collects the user's fingerprint image. After the fingerprint image undergoes preprocessing (such as filtering, enhancing contrast, etc.), a fingerprint feature extraction algorithm is used to extract the feature point information. Common fingerprint feature extraction algorithms include minutiae-based algorithms, which determine the position and type of feature points by detecting the changes in ridges and valleys in the fingerprint image and construct a fingerprint feature template.

[0025] The video image information of the target user includes the following: Video frame data: This is the digital representation of the original video signal input by the user through the camera of the intelligent intercom device. It records the pixel information of the video image at different time points and is the most basic form of video data. For example, the video signal is sampled at a certain frame rate (such as 25fps) and resolution (such as 640×480 pixels) to obtain a series of discrete video frames, and these frames constitute the video frame data. Video image content (obtained through image recognition technology): Image recognition algorithms are used to analyze the video frame data to identify elements such as people, objects, and scenes in it for semantic understanding and event judgment operations. For example, when the user appears at the community entrance in the video, the image recognition system recognizes key elements such as the user's image and the community gate, which can help the system understand the user's environment and behavior and further perform permission judgment and business processing.

[0026] The camera of the intelligent intercom device continuously captures the user's video signal. After the analog-to-digital conversion (A / D conversion) and necessary image signal processing (such as filtering, enhancing contrast, etc.), video frame data is obtained. With the support of the device's video driver and related image processing libraries, the video frame data is stored in the buffer, waiting for subsequent processing. An image recognition engine is used to process the video frame data. The image recognition engine usually adopts an algorithm based on a deep learning model (such as a convolutional neural network, CNN). First, feature extraction is performed on the video frame data, such as extracting the color features, texture features, shape features, etc. of the image, and then these features are input into the image recognition model. The model outputs the understanding result of the video image content according to the training data and scene knowledge. In the intelligent intercom device, a mature image recognition SDK (Software Development Kit) can be integrated to implement the analysis function of the video image.

[0027] S102. Process the biometric information of the target user to generate the identity attribute information of the target user.

[0028] In one implementation, the biometric information of the target user is processed to generate the initial face image information and the identity information of the target user. When the target user approaches a device with biometric recognition function (such as an intelligent access control device or an intercom with face recognition), the camera on the device will capture the user's face image, which is the face image part of the obtained biometric information. Then, the face image is preliminarily processed through an image processing algorithm. First, face detection is performed to determine the position and pose of the face in the image. For example, the cascade classifier algorithm based on Haar features or the face detection algorithm based on deep learning (such as SSD-Face, etc.) is used to find the face area in the image and separate it from the background to obtain the initial face image information. At the same time, the extracted face image is compared with the face template pre-stored in the database. The database stores the face images of registered users and their corresponding identity information (such as name, job number, user ID, etc.). Through the feature matching algorithm (such as feature matching based on Euclidean distance or feature vector comparison using deep learning), the similarity between the input face image and the face template in the database is calculated. If the similarity exceeds the set threshold, it is considered a successful match, and thus the identity information of the target user is obtained.

[0029] Assume that the resolution of the original image captured by the camera is 1280×720 pixels. After face detection and cropping, the obtained initial face image information may be a grayscale image of 256×256 pixels (for example only), which only contains the face part and removes irrelevant information such as the background. If the database stores the face template of a user named "Zhang San" with a job number of "001", when the target user Zhang San approaches the device and the comparison is successful, the obtained identity information of the target user is "Name: Zhang San, Job number: 001".

[0030] Process the initial facial image information to generate a data feature set, where the data feature set is used to represent the facial features corresponding to different expression information. Use a feature extraction algorithm to process the initial facial image information. For example, use traditional feature extraction methods such as the Local Binary Pattern (LBP) algorithm or the Scale-Invariant Feature Transform (SIFT) algorithm, or use a Convolutional Neural Network (CNN) based on deep learning to extract facial features. For different expressions, such as happy, angry, sad, surprised, etc., there are obvious changes in facial features in areas such as the eyes, eyebrows, and mouth. Through a large amount of training data (including facial images with different expressions and their labeled expression categories), the feature extraction algorithm can learn the relationship between these expressions and facial features. Combine the extracted facial features to form a data feature set, where each element corresponds to the facial feature representation of an expression.

[0031] Take a simple example based on LBP features. For a happy expression, features such as narrowed eyes (corresponding to a specific pattern of LBP feature values around the eyes) and upturned mouth corners (corresponding to a specific pattern of LBP feature values in the mouth area) may be extracted; for an angry expression, features such as frowning eyebrows (a specific pattern of LBP feature values in the eyebrow area) and widened eyes (a specific pattern of LBP feature values in the eye area) may be extracted. The features of these different expressions are combined together to form a data feature set. Suppose the data feature set is represented by a matrix, and each row represents the feature vector of an expression. For example, the feature vector of a happy expression is [0.1, 0.3, 0.2, …] (these are just example values here), and the feature vector of an angry expression is [0.4, 0.1, 0.5, …], etc.

[0032] Perform classification processing on the data feature set to generate target face classification information. Use a classification algorithm to classify the data feature set. Common classification algorithms include Support Vector Machine (SVM), decision tree, neural network, etc. Take each feature vector in the data feature set as input, and the classification algorithm, according to the trained model, determines which expression category it belongs to. For example, for a system using an SVM classifier, the SVM model learns the boundaries between different expression feature vectors and divides the input feature vector into the corresponding expression category, thereby generating target face classification information, that is, determining the expression category that the current facial image most likely belongs to. Suppose the classification result is that the target face classification information is "happy". This means that after classification processing, it is judged that the expression corresponding to the current input facial image is closest to the happy category. If the system also gives the confidence level of the classification, it may be "happy (confidence level: 0.85)", indicating that the system has a high confidence in this classification result. The confidence level value is calculated by the classification algorithm based on factors such as the distance between the feature vector and various category models.

[0033] Process the classification information of the target face part to generate facial feature information and light information that matches the facial feature information. According to the classification information of the target face part (such as expression categories like happy, angry, etc.), further analyze the facial feature information in the face image. For the extraction of facial features, a geometric model-based method can be used. For example, determine the position, shape, size and other parameters of facial features such as eyes, nose, mouth, etc. For instance, for a happy expression, the eyes may be in a squinted state. Describe the eye features by calculating geometric features such as the opening degree of the eyes (parameters such as the ratio of the eye width to the height), the curvature of the eye corners, etc.; the mouth is upturned, and describe the mouth features by measuring parameters such as the angle of the mouth corners, the change in the thickness of the lips, etc.; similarly, determine the relative position and shape change of the nose under the expression change (such as the degree of nasal wing expansion, etc.).

[0034] At the same time, analyze the light information in the image. The light information includes light intensity, light direction, etc. Estimate the light intensity through methods such as the brightness histogram of the image. For example, if the area with higher pixel values in the brightness histogram accounts for a larger proportion, the light intensity may be stronger; infer the light direction by analyzing the shadow conditions of different parts of the face. For example, if there is a shadow on one side of the nose, the light may come from the other side. Match the extracted facial feature information with the light information, that is, determine the specific feature pattern presented by the facial features under the current lighting conditions. The facial feature information may be expressed as: eye features (opening degree: 0.3, eye corner curvature: 0.5), mouth features (mouth corner angle: 45 degrees, change in lip thickness: +0.2 cm), nose features (degree of nasal wing expansion: 0.1), etc. (the values here are only examples, and the actual calculation will be obtained according to specific algorithms and image data). The light information may be expressed as: light intensity (average brightness value: 150, brightness range: 100 - 200), light direction (horizontal direction angle: 45 degrees, vertical direction angle: 30 degrees). These facial feature information and light information together reflect the facial feature state of the target user under the current expression and lighting conditions.

[0035] The facial feature information and the light information that matches the facial feature information are processed to generate the influencing factors of the target user's expression features. Based on the pre-established expression feature model, the model considers the influence of facial features, light information, and the interaction between them on the expression features. For example, a model is obtained through a large amount of experimental data and machine learning algorithm training, and the model knows which expression a specific facial feature combination is more likely to correspond to under a certain lighting direction and intensity. The facial feature information and light information are used as input and substituted into the expression feature model for calculation. The model analyzes the changing trend of facial features under the current lighting conditions and the degree of correlation between these changes and different expressions, so as to determine the key factors affecting the expression features. For example, if under side lighting conditions, the shadow part of the eyes is large and the shape of the mouth changes greatly, the model may determine that the light has a certain influence on the recognition of the expression, and the changes of the eyes and mouth in the facial features are the main influencing factors of the expression features. It is assumed that the generated influencing factors of the target user's expression features are: facial features (mainly changes in eyes and mouth), light influence (side lighting causes the recognition of some facial features to decrease). This shows that in the current situation, the morphological changes of the eyes and mouth are the main features for expressing facial expressions, and while the lighting conditions affect the accurate recognition of these features to a certain extent, the overall expression is still mainly determined by the facial features.

[0036] Process the influencing factors of the target user's facial expression features based on a preset emotion recognition model to generate target face image information, where the target face image information includes the expression change information of the target user. The preset emotion recognition model is usually a deep learning model trained with a large amount of training data, such as a convolutional neural network (CNN) or a recurrent neural network (RNN), etc. Preprocess the influencing factors of facial expression features (including facial feature information, light information, etc.) to make them meet the input requirements of the model. For example, encode the facial features and light information and convert them into a numerical vector form that the model can accept. Then input the preprocessed information into the preset emotion recognition model. The model, based on the input information and the relationship between expressions and emotions it has learned, predicts the expression change information of the target user. The model will output a description of the expression change, such as the change process from calm to happy, including the amplitude of the expression change (such as the gradual increase in the degree of happiness) and the speed of change (such as the time it takes for the expression to change from the start to the most obvious state). These information constitute the expression change part in the target face image information. Suppose the expression change information in the target face image information output by the model is: the expression starts from calm and gradually becomes happy within 2 seconds, and the degree of happiness reaches 80% (here the degree of happiness can be quantified by a numerical value, such as a value between 0 and 1, where 1 represents the highest degree of happiness). At the same time, the model may also output some detailed features of the expression change, such as the speed at which the eyes gradually narrow and the trend of the angle change of the upturned corners of the mouth, to more detailedly describe the dynamic change process of the expression.

[0037] Process the target face image information to generate the emotion information of the target user. Based on the expression change information in the target face image information, further map it to the corresponding emotion category and intensity. This may require an emotion mapping table or a machine learning-based emotion classification model. For example, if the expression change is from calm to happy quickly and the degree of happiness is relatively high, through the emotion mapping table, the corresponding emotion can be determined as "Joy (Intensity: High)". If using an emotion classification model, the model will match the expression change information with the patterns in the training data and output the most likely emotion category and intensity. The emotion categories can include joy, anger, sadness, surprise, fear, etc., and the intensity can be represented by low, medium, high, or quantified by a specific numerical range (such as a value between 0 and 1). Suppose the generated emotion information of the target user is "Joy (Intensity: 0.8)", indicating that the target user is currently in a relatively high degree of joy emotion state. This emotion information can be used in subsequent application scenarios. For example, in customer service, if it is detected that the customer is in a joy emotion, the service staff can provide more positive interactions; in security monitoring, if it is found that a person shows abnormal anger or fear emotions, further attention may be required.

[0038] Generate the identity attribute information of the target user based on the identity information and emotional information of the target user. Integrate the identity information (such as name, employee number, user ID, etc.) and emotional information (such as emotional categories and intensities like joy, anger, sadness, etc.) of the target user. A data structure can be established to store this information, such as a structure or object containing identity fields and emotional fields. Further process the integrated information according to business requirements. For example, in internal enterprise management, if an employee frequently shows anger during working hours, they may be marked as an emotionally unstable employee, and this mark is associated with the employee's identity information to form the identity attribute information of the target user. In marketing, if a customer shows joy when browsing a product, they may be marked as a potential highly satisfied customer, and together with the customer's identity information, it serves as the identity attribute information of the target user for subsequent precision marketing or customer relationship management.

[0039] Assume the identity information of the target user is "Name: Li Si, Employee Number: 002", and the emotional information is "Anger (Intensity: 0.6)". The generated identity attribute information of the target user may be "Employee Li Si (Employee Number 002) is emotionally unstable in the current scenario, with a relatively high degree of anger, and it is necessary to pay attention to their work status". Or in the customer scenario, the identity information of the target user is "Customer ID: 12345", and the emotional information is "Joy (Intensity: 0.8)". The identity attribute information may be "Customer 12345 is in a highly satisfied state and is a potential high-quality customer. Relevant high-end products or services can be recommended".

[0040] S103, Process the target TF card information to generate a target identity recognition model.

[0041] In one implementation, the target TF card information is processed to generate the original data of the target TF card and a preset deep learning model based on the multi-layer perceptron (MLP). When the target TF card is inserted, the system first reads the file system information of the TF card, including the file directory structure, file types, etc. For example, it identifies the specific folder path storing the face recognition data in the TF card. By traversing these paths, various data files stored in the TF card are obtained, such as face image files (possibly in common image formats like JPEG, PNG, etc.), potential related configuration files (containing information such as data format, resolution, etc.), and the pre-stored deep learning model file based on the MLP (the model file format may be a specific framework format, such as the SavedModel format of TensorFlow or the.pth format of PyTorch, etc.). These constitute the original data of the target TF card. At the same time, the preset deep learning model based on the MLP is loaded from the TF card into the system memory for subsequent training or application. The model is already constructed in the TF card, including the definition of the MLP network structure (such as the number of neurons and connection methods in the input layer, hidden layers, and output layer) and the initial parameter values (obtained through previous training or predefined).

[0042] Suppose there is a folder named "face_data" in the TF card, which contains 1000 face image files of different people, with an image resolution of 128×128 pixels and in JPEG format. Additionally, there is a folder named "model" that stores the deep learning model file based on MLP named "mlp_face_model.pb". By reading the information of these files and folders, the original data of the target TF card obtained is the collection of these image files and model files. The loaded MLP model may have a structure where the number of neurons in the input layer is 128×128 (corresponding to the number of image pixels), there are two hidden layers with 256 and 128 neurons respectively, and the output layer is the number of categories representing different people's identities (such as 100 categories corresponding to 100 different people), and the model parameters are the initial values obtained from previous training.

[0043] Preprocess the original data of the target TF card to generate target TF card data. For the obtained face image files (a part of the original data of the target TF card), perform image normalization processing. For example, normalize the pixel values of the image to between 0 and 1, which helps improve the stability and accuracy of subsequent processing. This can be achieved by dividing each pixel value by the maximum pixel value of the image (for 8-bit images, the maximum pixel value is 255). Crop and scale the image to ensure that all images have the same size and pose. For example, crop the image to a standard size of 112×112 pixels, and through image transformation techniques (such as affine transformation), make the key parts such as eyes, nose, and mouth in all face images be in similar positions and angles. If the original data contains a configuration file, read the relevant information in it, such as whether data augmentation operations (such as random flipping, rotation, adding noise, etc.) are required, and perform corresponding processing according to the configuration. Data augmentation can increase the diversity of training data and improve the generalization ability of the model. For example, if the configuration file specifies that a random horizontal flip operation is to be performed, then horizontally flip some of the images to generate new training samples. After preprocessing, the pixel values of 1000 face images are all between 0 and 1, and the size is uniformly 112×112 pixels. Assuming that data augmentation operations are performed, and 500 images are randomly horizontally flipped, then the final generated target TF card data contains 1500 preprocessed and data-augmented face images, and these images can be more conveniently used for subsequent feature extraction and model training.

[0044] Extract features from the target TF card data to generate face feature vectors. Use the convolutional layer in the deep learning model (if the MLP model contains a convolutional layer, or you can first use a separate convolutional neural network for feature extraction and then input it into the MLP) to extract features from the preprocessed face images. The convolutional layer slides a convolutional kernel over the image to extract local features of the image, such as edges, textures, etc. For example, use multiple convolutional kernels with different sizes and parameters to perform convolutional operations on a 112×112 pixel face image to obtain a series of feature maps. Perform pooling operations on the feature maps output by the convolutional layer, such as max pooling or average pooling, to reduce the resolution of the feature maps, reduce the amount of data, and at the same time retain the main features. For example, use a 2×2 max pooling operation to halve the size of the feature map. Flatten the pooled feature maps to convert them into one-dimensional vectors, and this vector is the face feature vector. For example, if a 10×10×64 (width×height×number of channels) feature map is obtained after convolution and pooling, after flattening, a face feature vector with a length of 10×10×64 = 6400 is obtained. For each preprocessed 112×112 pixel face image, through the above feature extraction process, a face feature vector with a length of 6400 (for example only) is generated. These 1500 images (the target TF card data after data augmentation) will respectively generate 1500 face feature vectors, and these vectors can more effectively represent the features of each face. Compared with the original image data, they have a lower dimension and contain more key recognition information.

[0045] Process the face feature vectors to generate a training set and a validation set. Divide the face feature vectors into a training set and a validation set according to a certain ratio (such as 80:20 or 70:30, etc.). For example, if there are 1500 face feature vectors and they are divided according to the 80:20 ratio, 1200 vectors are used for the training set and 300 vectors are used for the validation set. During the division process, ensure that the distribution of the training set and the validation set is representative, that is, it contains feature vectors of different people, and the ratio between various categories is relatively balanced (if it is used for a classification task, such as identifying different people's identities). You can use random sampling for the division, but pay attention to ensuring randomness and independence to avoid data leakage and bias. At the same time, for a classification task, save the person identity labels (such as 0 to 99 representing 100 different people) corresponding to the face feature vectors in the training set and the validation set together with the feature vectors for use when training and validating the model. The training set contains 1200 face feature vectors and the corresponding 1200 person identity labels, and the validation set contains 300 face feature vectors and the corresponding 300 identity labels. These data sets will be used to train and evaluate a deep learning model based on a multi-layer perceptron (MLP) to ensure that the model can accurately recognize face features and classify person identities.

[0046] Process a preset deep learning model based on a multi-layer perceptron (MLP) using a training set and a validation set to generate a target identity recognition module. Input the face feature vectors and corresponding identity labels of the training set into the preset deep learning model based on the MLP for training. In the MLP model, the input layer receives the face feature vectors, performs non-linear transformation and combination of features through the hidden layer, and finally outputs the predicted person identity category at the output layer. Calculate the loss function between the model prediction result and the true identity label, such as the cross-entropy loss function (for classification tasks). Then, according to the value of the loss function, use the backpropagation algorithm to adjust the parameters (such as weights and biases) in the MLP model to minimize the loss function and improve the accuracy of the model. This process will be repeated for multiple training epochs in the training set until the model converges or reaches the preset training stop conditions (such as the maximum number of training epochs, the loss function value is lower than a certain threshold, etc.).

[0047] During the training process, periodically evaluate the model using the validation set. Input the face feature vectors of the validation set into the training model and calculate evaluation metrics such as the accuracy and recall rate of the model on the validation set. If it is found that the performance of the model on the validation set no longer improves or even deteriorates (overfitting may occur), training can be stopped early, and the model parameters at the best performance are selected as the final model parameters. After training and evaluation, the optimized deep learning model based on the MLP is the target identity recognition module, which can be used to predict the identity of new face feature vectors. Suppose that during the training process, after 100 training epochs, the accuracy of the model on the validation set reaches 95% (for example only), and the accuracy no longer improves significantly in subsequent training. At this time, select the model parameters as the parameters of the final target identity recognition module. When new face feature vectors are input, the target identity recognition module can calculate and output the predicted person identity category according to these parameters, thus realizing the face recognition function. For example, in an access control system, it can determine whether the person coming is an authorized person, and in a monitoring system, it can identify the identity of the person in the picture, etc.

[0048] S104. Process the to-be-accessed application information to generate the application access purpose information of the target user and the access permission information of the target user.

[0049] In one implementation, the application information to be accessed is processed to generate the access information of the target user and the access feedback information of the target user. When the target user attempts to access an application named "Intercom Dispatching and Command Application", the system first obtains its relevant information. Suppose the application version number is 3.0 and the application source is custom-developed within the enterprise and distributed through the enterprise's dedicated server. The system checks the user's intercom device and finds that the application is installed, but the version is 2.5. At the same time, the device operating system is an Android system customized for a specific intercom device, with a version of 8.1, and the available memory is 200MB, while the application requires at least 500MB of available memory. Also, the firmware version of the device's communication module is V1.2, and the application requires a minimum communication module firmware version of V1.5 (to better support certain advanced dispatching functions). The generated access information of the target user is: The application is installed but the version is outdated, the device operating system meets the requirements but the available memory is insufficient, and the firmware version of the communication module is too low. The corresponding access feedback information is: "The version of your 'Intercom Dispatching and Command Application' is too low. Please go to the enterprise internal server to download and update it to version 3.0. At the same time, your device has insufficient available memory. It is recommended to close some unnecessary background applications or clear the cache to free up memory. In addition, the firmware version of your device's communication module is too low. You need to upgrade it to version V1.5 or higher to fully use all the functions of the application. You can upgrade the firmware in the system update option in the device settings (if available), or contact the device administrator for assistance in upgrading."

[0050] The access information of the target user is processed to generate the access permission information of the target user. The system queries relevant rules in the user permission management system according to the above access information of the target user. Suppose the target user is an ordinary security guard in the security department. For the "Intercom Dispatching and Command Application", the permission management system stipulates that: Ordinary security guards can only use the intercom equipped with a camera for group video calls during the specified working hours (such as from 8 am to 8 pm), and can only join the groups in their own security area (such as "Security Group in Area A of the Venue", "Security Group in Area B of the Venue", etc.), and cannot create or manage groups; they can receive emergency notifications and instructions from the superior dispatcher, but cannot send global broadcasts; they can only view the security plan documents related to their own positions and cannot modify or delete any documents. The generated access permission information of the target user is: Permitted operations - Group video calls (during working hours, only within the groups in their own security area), receiving emergency notifications and instructions, viewing relevant security plan documents; Prohibited operations - Group creation and management, sending global broadcasts, modifying or deleting documents.

[0051] Process the access feedback information of the target user to generate the application access purpose information of the target user. The system performs semantic analysis and intent recognition on the above access feedback information (prompting to update the application version, release memory, and upgrade the communication module firmware). Since the prompts mainly revolve around application version updates and device resource and firmware upgrades, the system infers that the main purpose of the user using this application is to perform efficient video communication and collaboration in a security work scenario. The user may need to receive dispatch instructions in a timely manner and cooperate with colleagues to handle security incidents. However, due to device resource and firmware issues, the use of some advanced functions of the application may be affected, such as high-definition video calls, real-time location sharing (if relying on the communication module firmware function), etc., and the video image analysis functions (such as personnel recognition, behavior analysis, etc.) may also be restricted. The generated application access purpose information of the target user is: in a security work scenario, the main purpose of the user is to perform real-time video communication and collaboration (such as receiving instructions, communicating with colleagues about security situations, etc.). However, due to the current state of the device, the advanced functions of the application may not be fully utilized, and it is necessary to update the application, release memory, and upgrade the communication module firmware as soon as possible to obtain a better working experience. This information helps the system to prioritize ensuring the stability of the basic video communication function during the subsequent use of the user, prompt the user to upgrade relevant functions when device resources permit, or guide the user to experience new security dispatch functions (such as more accurate personnel positioning based on the upgraded communication module firmware and abnormal behavior detection based on video image analysis) after the application is updated and the device is optimized.

[0052] S105. Process the application access purpose information of the target user and the access permission information of the target user to generate the access permission information of the target application.

[0053] In one implementation, the access permission information of the target user is processed to generate the access time period information of the target user, where the access time period information of the target user is used to characterize the single access duration and the accessible time period of the target user. In the walkie-talkie usage management system of an enterprise, the access permission information of the target user is stored in a dedicated permission database. For example, for ordinary security personnel, the system stipulates that they can use a walkie-talkie equipped with a camera for video communication during their daily work (Monday to Friday, 7:00 am to 7:00 pm), and the single continuous usage duration shall not exceed 1.5 hours to avoid excessive battery consumption and long-term occupation of the device, which may affect the use of other personnel. After the system reads the permission record from the database, it parses out the accessible time period (Monday to Friday, 7:00 - 19:00) and the single access duration limit (1.5 hours) of the ordinary security personnel, and generates the access time period information. During special events, such as large-scale exhibitions, some security personnel are assigned to specific areas to perform special tasks and need to obtain additional walkie-talkie usage permissions during the exhibition period (such as Saturday and Sunday, 8:00 am to 10:00 pm). The system will merge these special time periods into the access time period information of these security personnel according to the temporary permission settings. Taking the security personnel "Zhao Liu" as an example, his regular access time period is Monday to Friday, 7:00 - 19:00, and the single access duration is 1.5 hours. The temporary permission obtained due to the exhibition is Saturday and Sunday, 8:00 - 22:00. Then the generated access time period information of "Zhao Liu" is: the accessible time period is Monday to Friday, 7:00 - 19:00 and Saturday and Sunday, 8:00 - 22:00; the single access duration limit is 1.5 hours.

[0054] Process the application access usage information of the target user to generate the access priority of the target user. The system obtains the application access usage information of the target user, and the sources of this information are diverse. For example, in the enterprise's walkie-talkie dispatching application, if a security guard selects the "Emergency Incident Response in Patrol Area" task when logging in, the system will determine that the application access usage of this security guard has a relatively high urgency. Or, if the system analyzes that a certain security guard has frequently used the walkie-talkie to contact the monitoring center recently, involving operations such as the investigation of suspicious persons, it can be inferred that the application access usage of this security guard is to handle important security monitoring matters and should be given a relatively high priority. At the same time, consider the department and position characteristics of the user. For example, the access priority of the emergency rescue team during the execution of rescue tasks is higher than that of ordinary security guards during daily patrol communication. The system assigns a priority value or level to the target user according to the preset priority rule algorithm. The priority can be set from level 1 to level 5. Users performing key tasks such as emergency rescue and handling emergency security incidents may be assigned a priority level of 1, while ordinary patrol security guards may be assigned a priority level of 3 or 4. Suppose the security guard "Sun Qi" belongs to the emergency rescue team. When he logs in to the walkie-talkie dispatching application, he selects the "Earthquake Rescue Site Communication" task, and the system analyzes that he has mainly participated in various emergency rescue drills and actual rescue work recently. According to the priority rules, emergency rescue is a high-priority service during the execution of tasks. The system generates an access priority of level 1 for "Sun Qi", indicating that he should be given priority to obtain walkie-talkie resources and communication responses during the current access to ensure the timely transmission of rescue instructions and the rapid feedback of rescue situations.

[0055] Process the access priorities of target users to generate real-time access channel information for the users to be accessed. There are multiple predefined access channels in the walkie-talkie system, and the resource allocation and functional characteristics of each channel are different. For example, there are high-definition video channels, ordinary video channels, and dedicated emergency call channels, etc. (assuming here that the walkie-talkie supports channels for video communication of different qualities). According to the access priorities of target users, the system assigns users to corresponding real-time access channels. High-priority users (such as those with priorities 1 and 2) are usually assigned to high-definition video channels to ensure clear and stable video communication quality, which is convenient for efficiently handling emergency rescue, important scheduling, and other matters, and the high-definition video channels can support clearer identification of personnel and on-site conditions. Medium-priority users (such as level 3) may be assigned to ordinary video channels, and low-priority users (such as levels 4 and 5) are assigned to channels with lower resource occupancy (such as channels that only provide basic video communication functions) to reasonably allocate system resources and avoid communication obstruction for high-priority users. The system monitors the load conditions of each channel in real time. If there are too many high-priority users in the high-definition video channel, resulting in resource tension, the system will, according to the load balancing strategy, dynamically adjust some high-priority users to other channels at the same level with sufficient resources. Taking the security guard "Sun Qi" as an example, his access priority is level 1. The system assigns him to the high-definition video channel for real-time access. At this time, there may be other team members or command personnel participating in the earthquake rescue in the high-definition video channel. The system records the user list of this channel in real time to form the real-time access channel information of the users to be accessed. For example, the users to be accessed in the high-definition video channel include "Sun Qi", "Zhou Ba" (the team leader of the rescue team), etc., and the current resource utilization rate of this channel is 80% (indicating that there is still a certain amount of resources to accommodate more high-priority users, but it is close to saturation).

[0056] Process the real-time channel information of the users to be accessed to generate the access time period information of the users to be accessed. Based on the characteristics and usage rules of the real-time access channel where the users to be accessed are located, the system determines their access time period information. For example, due to high resource occupancy, the high-definition video channel stipulates that the single continuous usage duration of each user shall not exceed 45 minutes, and during large-scale events or emergency incidents (such as the above-mentioned earthquake rescue scenario), this channel is open 24 hours a day; during non-emergency periods (such as daily enterprise operations), it is only open during working hours (such as from 8:00 am to 6:00 pm), and is closed or switched to a low-resource occupancy mode during non-working hours. Taking the security guard "Sun Qi" in the high-definition video channel as an example, if it is currently during the earthquake rescue period and he enters the channel at 9:00 am. According to the channel rules, his single access duration shall not exceed 45 minutes, and the channel is open 24 hours a day. Therefore, the generated access time period information of "Sun Qi" is: the accessible time period is from 9:00 am to 9:45 am (if continuously accessing), and the overall available time period of the channel is 0:00 - 24:00 (during the rescue period).

[0057] Process the access time period information of the target user based on the access time period information of the user to be accessed, and generate the access permission information of the target application. The system comprehensively considers the original access time period information of the target user (based on permission settings) and the access time period information in the real-time channel (based on channel rules). For example, the original access time period information of the security guard "Sun Qi" is that he can access the walkie-talkie at any time during an emergency (such as in an earthquake rescue scenario) (based on permission settings), and the access time period information in the high-definition video channel is from 9:00 am to 9:45 am (based on channel rules). Take the intersection of the two, and finally determine that the access time period of "Sun Qi" in the current situation is from 9:00 am to 9:45 am. At the same time, the system re-evaluates the permissions according to the adjusted access time period. If "Sun Qi" originally had the permission to perform comprehensive operations on the walkie-talkie during the rescue period (such as switching channels, adjusting the volume, initiating group video calls, etc.), but during this shortened time period, the system may temporarily restrict some of his non-critical operations (such as prohibiting random switching to non-rescue-related channels) to ensure that the communication focuses on the rescue task and prevent misoperations from affecting the rescue efficiency. According to the comprehensive processing results, generate the access permission information of the target application (walkie-talkie dispatching application), and clarify the specific operation permissions and accessible time range of "Sun Qi" for the walkie-talkie in the current situation. If he still needs to use the walkie-talkie to participate in the rescue after 9:45 am, he needs to re-evaluate the permissions or re-apply for access according to the channel resource situation (such as when channel resources permit, the access time period and corresponding permissions can be re-allocated).

[0058] S106. Process the access permission information of the target application and the identity information of the target user based on the target identity recognition model to generate the video image information of the target walkie-talkie device.

[0059] In one implementation, the access permission information of the target application is processed based on the target identity recognition model to generate video channel allocation information. Suppose the target application is an intercom dispatching and command system within an enterprise. Its access permission information includes the user's identity role (such as ordinary security personnel, security supervisor, system administrator, etc.), function permissions (such as whether they can initiate an emergency call, whether they can conduct a video call with personnel in a specific area, etc.), and data access levels (such as whether they can obtain the video channels associated with the surveillance information of sensitive areas, etc.). The target identity recognition model first accurately identifies and verifies the user's identity through methods such as face recognition, fingerprint recognition, or password verification (combined with the identity recognition function of the intercom device) to ensure a complete match with the identity information stored in the system. Then, based on the user's identity and access permission information, the video channel allocation information is determined according to pre-set complex rules and algorithms. For example, for the ordinary security personnel "Li Si", his permission only allows him to conduct group video calls with security personnel in the same area within his responsible patrol area (such as Area A of the factory campus). The target identity recognition model will retrieve the predefined list of security group video channels in Area A of the campus in the system and assign "Li Si" to the corresponding channel, such as the "Daily Patrol Video Channel in Area A". If "Li Si" attempts to access channels in other areas or manage channels, the system will reject his access. In addition to being able to access the security group video channels in their own area, the security supervisor can also access the dedicated command and dispatching video channels for management to coordinate security work across regions and make emergency decision-making commands. The system administrator has the permission to access all video channels for maintenance and management operations when the system fails or global settings are required. At the same time, for security personnel with access permissions to sensitive area surveillance data, the model ensures that they can only access the video channels that match their own permissions to prevent the leakage of sensitive information. For example, the security personnel responsible for monitoring the data room can only access the warning and emergency handling video channels related to the data room and cannot enter other irrelevant channels to obtain information.

[0060] Process the identity information of the target user based on the video channel allocation information to generate the video image information of the target user. The system obtains the detailed identity information of the target user, including name, employee number, department or position (such as security department - patrol post in Area A), etc., as well as the video channel allocation information generated in the previous step. For "Li Si", the system closely associates and records the identity information such as his name "Li Si", employee number "00234", and position "patrol post in Area A" with the allocated "Daily Patrol Video Channel in Area A". And, according to the default permissions of ordinary security personnel in this channel, clarify the specific operation permissions of "Li Si" in the channel, such as the basic permissions to normally turn on the camera to report the patrol situation via video, view the video images of other security personnel in the channel, view the online member list of the current channel, etc., but cannot perform channel management operations (such as adding or deleting channel members, modifying the channel name or settings, etc.). These information are integrated to form the video image information of the target user "Li Si". In addition, it may also include some personalized settings, such as the video notification sound being the default warning sound effect (to attract attention in a noisy environment), and the initial value of the camera resolution being 640×480 (which can be adjusted by the user according to the actual usage environment) to optimize the user experience in this channel.

[0061] Process the video image information of the target user to generate the video image information of the target intercom device. When "Li Si" uses the equipped intercom device to log in to the intercom dispatching and command application, the application quickly sends the video image information of "Li Si" to the intercom device. After receiving this information, the intercom device immediately performs a series of intelligent setting and configuration operations. It will clearly display on the device screen the video channel names available for "Li Si" to choose from, that is, "Daily Patrol Video Channel in Area A". At the same time, according to the permission settings of "Li Si" in this channel, it intelligently enables or disables the corresponding operation buttons on the device. For example, the "Start Video Call" button is in an available state, facilitating "Li Si" to communicate with colleagues via video at any time to report the situation; the "Channel Management" button is in an unavailable state (gray display or hidden) to prevent misoperation. Meanwhile, the intercom device automatically establishes a stable connection with the application server and performs precise video parameter settings based on the channel parameters provided by the server. For example, set the video encoding format to H.264 (a commonly used efficient video encoding format), the frame rate to 25fps (to reduce the data transmission volume while ensuring video smoothness), and the resolution to 640×480 (which can be adjusted according to actual requirements and network conditions) to ensure clear, smooth and low-latency video interaction with other intercom devices in the "Daily Patrol Video Channel in Area A", meeting the strict requirements of real-time communication in security patrol work. At this time, the video image information of the intercom device has been fully configured according to the permissions and requirements of "Li Si", providing a solid guarantee for his efficient video communication in this channel.

[0062] S107. Process the emotional information of the target user and the video image information of the target user based on the video image information of the target intercom device to generate target event information.

[0063] In one implementation, the emotional information of the target user and the video image information of the target user are processed based on the video image information of the target intercom device to generate video call keyword information. Suppose the target intercom device is being used in the production workshop scheduling scenario of an enterprise, and the video image information indicates that the current channel is the "Workshop Production Scheduling Video Channel". The target user is a workshop worker who looks anxious and acts flustered during the video call (obtained through video image analysis). The system first performs behavior and speech recognition processing on the video image information of the target user, converts the speech signal into text, and analyzes the meaning of his behavior actions (for example, the worker frequently points to a certain machine, and his intention can be understood in combination with the speech). For example, the worker says, "This machine is malfunctioning and has been making strange noises. I can't continue working. Hurry up and find someone to fix it!" At the same time, the emotional information of the target user is analyzed using emotion recognition technology, and it is determined that he is in an anxious and angry emotional state (by analyzing multi-modal features such as facial expressions, body movements, and the pitch, speed, and volume of speech). Then, according to the preset keyword extraction rules, key information is extracted from the recognized text and the information contained in the behavior actions. In this example, the extracted keywords may include "Machine malfunction", "Strange noise", "Find someone to fix", etc. At the same time, combined with the analysis of the behavior actions in the video, there may also be "Specific machine location (determined according to the direction pointed by the worker)". These keywords are combined to form video call keyword information, which reflects the core content of the current call, the emotional state of the user, and the relevant behavior action information. Video call keyword information: ["Machine malfunction", "Strange noise", "Find someone to fix", "Anxiety-related features (such as tense facial expressions, flustered body movements, high pitch, fast speech rate, etc.)", "Specific machine location (such as Machine Tool No. 3 in Area A of the workshop)"].

[0064] Process the keyword information in the video call to generate a target warning value and event information that matches the target warning value. There is a pre-set keyword classification and weight system in the system. For example, keywords are classified into equipment failure categories, safety accident categories, personnel requirement categories, etc. In this example, "machine failure" belongs to the equipment failure category keyword. For the keyword "machine failure", assume that its basic weight ω (assume it is 8) in the equipment failure category keyword set is relatively high because it directly indicates that there is a problem with the production equipment and may affect the production progress. Its occurrence frequency f in the video call (here it is 1 time, but if the worker emphasizes it repeatedly, the frequency will increase accordingly). Since the current video channel is the "workshop production scheduling video channel", the correlation coefficient r (assume it is 0.9) between the equipment failure category keyword and this channel is relatively high because this channel is mainly used to handle workshop production-related matters. The target user is in an anxious and angry mood, and the mood influence factor e (assume it is 1.5, indicating a relatively strong mood). At the same time, considering additional information such as the "specific machine location" extracted from the video, it will also be an important reference factor when calculating the warning value and determining the event information. For example, if the location where the machine is located is a key link in the production process, it may further increase the warning value.

[0065] The method also includes a calculation formula for obtaining the target warning value. The calculation formula is: ; where m represents the number of categories in the keyword set, represents the number of keywords included in the j-th category keyword set, represents the basic weight of the i-th keyword in the j-th category keyword set, represents the occurrence frequency of the i-th keyword in the j-th category keyword set in the video call, represents the correlation coefficient between the j-th category keyword set and the current video channel, and e represents the user mood influence factor. Calculate the target warning value V (assuming there is only this one keyword) according to the formula: V = r×(ω×f)×e = 0.9×(8×1)×1.5 = 10.8. At the same time, according to the keyword "machine failure", the event information matched by the system is "Production equipment failure event, need to arrange maintenance personnel to go for handling". For example, target warning value: 10.8, event information: "Production equipment failure event, need to arrange maintenance personnel to go for handling".

[0066] Process the target warning value based on a preset event warning table to generate a real-time event warning value. Different processing methods and warning levels corresponding to different warning value ranges are set in the preset event warning table. For example, a warning value between 0 and 10 is a low risk, between 10 and 20 is a medium risk, and above 20 is a high risk. For the calculated target warning value of 32.4, look up the corresponding processing method and warning level in the preset event warning table. Determine that it is at the high risk level, and operations such as immediately notifying the supervisor of the maintenance department and dispatching more resources may be required. Adjust or convert the target warning value (if necessary) according to the rules in the warning table to generate a real-time event warning value. For example, 32.4 may be converted into a specific high-risk code (such as "HR-03", indicating a more serious high-risk situation) for subsequent unified processing and identification by the system. Real-time event warning value: "HR-03" (indicating a high-risk production equipment failure event).

[0067] If the real-time event warning value is greater than the preset warning threshold, generate an event warning message. Assume that the preset warning threshold is set to 20 (i.e., a warning is triggered when the real-time event warning value is greater than 20). Since the actual warning value corresponding to the real-time event warning value "HR-03" is 32.4, which is greater than the preset warning threshold, the system generates an event warning message. The event warning message may include the warning level (high risk), event type (production equipment failure), occurrence location (specific location in the workshop, which can be located through an intercom device or manually reported by workers, known here as production line 3 in area A of the workshop), relevant personnel (target user and the team they belong to, i.e., worker Zhang San and the team of production line 3 in area A), etc. At the same time, the system may send the warning message in various ways, such as popping up a prominent warning notice on the display screen in the workshop dispatching center, sending text messages or push notifications to the mobile phones of relevant managers, etc. Event warning message: {"warning level": "high risk", "event type": "production equipment failure", "occurrence location": "production line 3 in area A of the workshop", "relevant personnel": "worker Zhang San and the team of production line 3 in area A"}.

[0068] Process the event warning information to generate target event information, which is used to represent the video call information between the target user and other users. The other users are generated based on the video image information of the target intercom device. The system further improves and integrates relevant data according to the event warning information to generate target event information. In addition to the content in the event warning information, it may also add the timestamp of the event (accurate to seconds, recording the time when the equipment failure occurred, assumed to be 2023-11-10 14:30:15), the detailed description of the event (such as the specific manifestations of the machine failure, supplementing "the machine made abnormal metal friction sounds, stopped running, and there was a smoking phenomenon" according to the description of the worker, etc.), and the possible impact assessment (such as the expected duration of the impact on the production schedule is 3 hours, the possible product loss is 200 pieces, etc.). The target event information will be stored in the system database for convenient subsequent query, statistics and analysis, and at the same time provides a more comprehensive basis for further decision-making and processing. For example, the maintenance department can prepare corresponding maintenance tools and spare parts according to the target event information, and the production management department can adjust the production plan according to the impact assessment, such as arranging other production lines to work overtime temporarily to make up for the lost output. For example, the target event information: {"event time": "2023-11-10 14:30:15", "warning level": "high risk", "event type": "production equipment failure", "occurrence location": "Production line No. 3 in Area A of the workshop", "related personnel": "Worker Zhang San and the team of Production line No. 3 in Area A", "detailed event description": "The machine made abnormal metal friction sounds, stopped running, and there was a smoking phenomenon", "possible impact assessment": "Expected to affect the production schedule for 3 hours, possible product loss of 200 pieces"}.

[0069] This application obtains the target TF card information (including capacity, format, serial number, etc.), the application information to be accessed (such as the version, function modules, permission requirements, compatibility, etc. of the intercom dispatching and command application), the biometric information of the target user (face or fingerprint features), and the video image information (waveform data and text content) from the server. Then, process the biometric information to generate identity attribute information, including generating identity information through detection and comparison, and then obtaining emotion information through feature extraction, classification, etc. Process the TF card information to generate a target identity recognition model, which involves reading the original data, preprocessing, feature extraction, dividing the data set, and training the model. Process the application information to be accessed to obtain the application access purpose and access permission information, such as generating access feedback information according to the device situation, and then inferring the access purpose based on this.

[0070] Then, the target application access permission information is generated by combining the application access purpose and permission information, including determining the access time period and priority, allocating a real-time access channel, and finally determining the permission. The video image information of the intercom device is generated based on the target identity recognition model and relevant information, covering the allocation of video channels and device configuration. Finally, the target event information is generated by processing the user's emotion and video image information according to the video image information of the intercom device. First, keywords are extracted to calculate the warning value, and then it is processed by the warning table. If the threshold is exceeded, event warning information is generated, and the target event information (including time, location, personnel, impact assessment, etc.) is further improved for subsequent query, decision-making, and processing. The overall method helps to improve the security, functionality, and management efficiency of the intelligent intercom terminal.

[0071] In one implementation, as Figure 2 shown, the present application further provides a face recognition device for an intelligent intercom terminal based on an external TF card, including: An acquisition module 201, configured to acquire target TF card information, information of an application to be accessed, biometric information of a target user, and video image information of the target user; A processing module 202, configured to process the biometric information of the target user to generate identity attribute information of the target user, where the identity attribute information of the target user includes identity information of the target user and emotion information of the target user; process the target TF card information to generate a target identity recognition model; process the information of the application to be accessed to generate application access purpose information of the target user and access permission information of the target user; process the application access purpose information of the target user and the access permission information of the target user to generate access permission information of the target application; process the access permission information of the target application and the identity information of the target user based on the target identity recognition model to generate video image information of the target intercom device; process the emotion information of the target user and the video image information of the target user based on the video image information of the target intercom device to generate target event information, where the target event information is used to represent voice call information between the target user and other users, and the other users are generated based on the video image information of the target intercom device.

Claims

1. A face recognition method for an intelligent intercom terminal based on an external TF card, characterized in that: include: Obtain target TF card information, application information to be accessed, target user's biometric information, and target user's video image information; Processing the biometric information of the target user to generate identity attribute information of the target user, wherein the identity attribute information of the target user includes the identity information of the target user and the emotional information of the target user; Process the target TF card information to generate a target identity recognition model; Processing the application information to be accessed to generate application access purpose information of the target user and access permission information of the target user; Processing the target user's application access purpose information and the target user's access permission information to generate the target application's access permission information; The target application's access permission information and the target user's identity information are processed based on the target identity recognition model to generate video image information of the target intercom device; The target user's emotion information and the target user's video image information are processed based on the target intercom device's video image information to generate target event information.

2. The method according to claim 1, characterized in that Process the target user's biometric information to generate the target user's identity attribute information, including: Processing the target user's biometric information to generate initial facial image information and the target user's identity information; Processing the initial facial image information to generate a data feature set, wherein the data feature set is used to characterize facial features corresponding to different expression information; Classify the data feature set to generate target face classification information; Processing the target face classification information to generate facial feature information and light information matching the facial feature information; Processing facial feature information and light information matching the facial feature information to generate influencing factors of the target user's facial expression features; Processing the influencing factors of the target user's expression characteristics based on a preset emotion recognition model to generate target face image information, wherein the target face image information includes the expression change information of the target user; Process the target face image information to generate the target user's emotional information; The identity attribute information of the target user is generated based on the identity information of the target user and the emotional information of the target user.

3. The method according to claim 1, characterized in that Process the target TF card information to generate a target identity recognition model, including: Process the target TF card information to generate the original data of the target TF card and a preset deep learning model based on the multi-layer perceptron MLP; Preprocess the original data of the target TF card to generate the target TF card data; Perform feature extraction on the target TF card data to generate a facial feature vector; Process the facial feature vector to generate training set and validation set; Based on the training set and the validation set, the preset deep learning model based on the multi-layer perceptron MLP is processed to generate a target identity recognition module.

4. The method according to claim 1, characterized in that The application information to be accessed is processed to generate the application access purpose information of the target user and the access permission information of the target user, including: Process the access application information to generate the target user's access information and the target user's access feedback information; Processing the target user's access information to generate the target user's access permission information; The access feedback information of the target user is processed to generate the application access usage information of the target user.

5. The method according to claim 4, characterized in that The target user's application access purpose information and the target user's access permission information are processed to generate access permission information for the target application, including: Processing the access permission information of the target user to generate access time period information of the target user, wherein the access time period information of the target user is used to characterize the single access duration and accessible time period of the target user; Process the target user's application access purpose information to generate the target user's access priority; Process the access priority of the target user and generate real-time access channel information of the user to be accessed; Processing the real-time channel information of the user to be visited, generating the visiting time period information of the user to be visited; The access time period information of the target user is processed based on the access time period information of the user to be accessed, and access permission information of the target application is generated.

6. The method according to claim 5, characterized in that The target application's access permission information and the target user's identity information are processed based on the target identity recognition model to generate video image information of the target intercom device, including: Processing the access permission information of the target application based on the target identity recognition model to generate video channel allocation information; Processing the identity information of the target user based on the video channel allocation information to generate video image information of the target user; The video image information of the target user is processed to generate the video image information of the target intercom device.

7. The method according to claim 6, characterized in that The target user's emotional information and the target user's video image information are processed based on the target intercom device's video image information to generate target event information, including: Processing the target user's emotional information and the target user's video image information based on the target intercom device to generate call video keyword information; Processing the call video keyword information to generate a target warning value and event information matching the target warning value; Process the target warning value based on the preset event warning table to generate a real-time event warning value; If the real-time event warning value is greater than the preset warning threshold, event warning information is generated; Process the event warning information to generate target event information; The method also includes a calculation formula for obtaining the target warning value, the calculation formula is: ; Among them, m represents the number of categories in the keyword set, represents the number of keywords contained in the j-th keyword set, represents the basic weight of the i-th keyword in the j-th keyword set, represents the frequency of occurrence of the i-th keyword in the j-th keyword set in the call video, represents the correlation coefficient between the j-th keyword set and the current video channel, and e represents the user emotion influencing factor.

8. A face recognition device for an intelligent intercom terminal based on an external TF card, characterized in that: The device includes: The acquisition module is used to obtain the target TF card information, the application information to be accessed, the biometric information of the target user and the video image information of the target user; The processing module is used to process the biometric information of the target user to generate the identity attribute information of the target user, wherein the identity attribute information of the target user includes the identity information of the target user and the emotional information of the target user; process the target TF card information to generate a target identity recognition model; process the application information to be accessed to generate the application access purpose information of the target user and the access permission information of the target user; process the application access purpose information of the target user and the access permission information of the target user to generate the access permission information of the target application; process the access permission information of the target application and the identity information of the target user based on the target identity recognition model to generate the video image information of the target intercom device; process the emotional information of the target user and the video image information of the target user based on the video image information of the target intercom device to generate the target event information.

9. An electronic device, characterized in that: include: a first processor; and a memory for storing executable instructions for the first processor; The first processor is configured to execute the face recognition method for a smart intercom terminal based on an external TF card of any one of claims 1 to 7 by executing executable instructions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the second processor, the face recognition method of the smart intercom terminal based on an external TF card according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Identity authentication system with biological characteristic recognition function and authentication method thereof

    CN101986597A

  • Palm vein identification intelligent building visible intercom system

    CN105208320A

  • Acquiring biometric print by means of smartcard

    CN110826674A

  • Information processing method and device, electronic equipment and storage medium

    CN119007754A

  • Two-dimensional code recognition unlocking method suitable for entrance machine and related equipment

    CN119251939A