A face recognition method and related equipment for intelligent intercom terminal based on external TF card

The face recognition function of the intelligent intercom terminal is expanded through external TF cards, which solves the problem that low-configuration devices cannot support high-complexity facial recognition, and achieves convenient function expansion and cost savings.

CN120071452BActive Publication Date: 2025-08-08FUJIAN RUIYUNLIAN INTELLIGENT TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510555450.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-08
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Low-configured intelligent intercom devices are difficult to run high-complex facial recognition functions, and the hardware upgrade cost is high. The existing technology lacks a method to flexibly expand facial recognition functions.

Method used

The functions of the intelligent intercom terminal are expanded through external TF cards, dynamically load or unload the face recognition function, without hardware upgrades, and the TF card stores algorithm models and biometric information to achieve identity and emotion recognition.

Benefits of technology

It realizes the convenient facial recognition function expansion of low-configuration devices, saves costs, and can be uninstalled at any time when not needed, solving the problem of storage space limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071452B_ABST
    Figure CN120071452B_ABST
Patent Text Reader

Abstract

The present invention provides a face recognition method and related equipment for an intelligent intercom terminal based on an external TF card, which is applied to the field of data processing technology. The application processes the biometric information of a target user to generate the identity attribute information of the target user; processes the target TF card information to generate a target identity recognition model; processes the application information to be accessed to generate the application access purpose information of the target user and the access permission information of the target user; processes the application access purpose information of the target user and the access permission information of the target user to generate the access permission information of the target application; processes the access permission information of the target application and the identity information of the target user based on the target identity recognition model to generate the video image information of the target intercom device; processes the emotional information of the target user and the video image information of the target user based on the video image information of the target intercom device to generate target event information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a face recognition method for an intelligent intercom terminal based on an external TF card and related equipment. Background Art

[0002] Among existing smart intercom devices, low-configuration devices have difficulty running highly complex functions such as face recognition due to limitations in processing power and storage space. Especially in scenarios where face recognition requires a large amount of storage space to load algorithm models, the hardware limitations of low-configuration devices make it difficult to meet application requirements. Currently, most low-configuration smart intercom devices on the market still rely on traditional recognition methods (such as password input or card recognition) and cannot support more advanced biometric recognition, which limits the user experience. If the hardware of these devices is upgraded, the cost will increase significantly, and the original equipment will be difficult to be compatible with such upgrades. Therefore, the existing technology lacks a method that does not rely on internal storage upgrades and can flexibly expand face recognition functions.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0004] The purpose of this application is to provide a face recognition method for an intelligent intercom terminal based on an external TF card and related equipment and systems, which at least to a certain extent overcome the problems existing in the prior art. By expanding the functions of low-configuration intelligent intercom devices through external TF cards, the device can dynamically load or unload the face recognition function according to demand, without the need for hardware upgrades, thus saving costs. By using an external TF card, the device does not need to occupy the limited internal storage space, solving the problem that low-configuration devices cannot support large-scale algorithm models. Users only need to insert the TF card to use the face recognition function, which is simple and convenient, and can be uninstalled or removed at any time when the function is not needed.

[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the invention.

[0006] According to one aspect of the present application, a face recognition method for an intelligent intercom terminal based on an external TF card is provided, including: obtaining target TF card information, information of an application to be accessed, biometric information of a target user, and video image information of the target user; processing the biometric information of the target user to generate identity attribute information of the target user, wherein the identity attribute information of the target user includes the identity information of the target user and the emotional information of the target user; processing the target TF card information to generate a target identity recognition model; processing the application information to be accessed to generate application access purpose information of the target user and access permission information of the target user; processing the application access purpose information of the target user and the access permission information of the target user to generate access permission information of the target application; processing the access permission information of the target application and the identity information of the target user based on the target identity recognition model to generate video image information of a target intercom device; processing the emotional information of the target user and the video image information of the target user based on the video image information of the target intercom device to generate target event information, wherein the target event information is used to characterize voice call information between the target user and other users, and other users are generated based on the video image information of the target intercom device.

[0007] Another aspect of the present application is a face recognition device for an intelligent intercom terminal based on an external TF card, characterized in that it includes: an acquisition module for acquiring target TF card information, information of an application to be accessed, biometric information of a target user, and video image information of the target user; a processing module for processing the biometric information of the target user to generate identity attribute information of the target user, wherein the identity attribute information of the target user includes the identity information of the target user and the emotional information of the target user; processing the target TF card information to generate a target identity recognition model; processing the application information to be accessed to generate application access purpose information of the target user and access permission information of the target user; processing the application access purpose information of the target user and the access permission information of the target user to generate access permission information of the target application; processing the access permission information of the target application and the identity information of the target user based on the target identity recognition model to generate video image information of a target intercom device; processing the emotional information of the target user and the video image information of the target user based on the video image information of the target intercom device to generate target event information, wherein the target event information is used to represent voice call information between the target user and other users, and other users are generated based on the video image information of the target intercom device.

[0008] According to another aspect of the present application, an electronic device is characterized in that it includes: a first processor; and a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the above-mentioned face recognition method of the smart intercom terminal based on an external TF card by executing the executable instructions.

[0009] According to another aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a second processor, the above-mentioned face recognition method of the smart intercom terminal based on the external TF card is implemented.

[0010] According to another aspect of the present application, a computer program product is provided, including a computer program, characterized in that when the computer program is executed by a third processor, the computer program implements the above-mentioned face recognition method for the smart intercom terminal based on an external TF card.

[0011] This application provides a facial recognition method and related equipment for smart intercom terminals based on an external TF card. By using an external TF card to expand the functionality of low-configuration smart intercom devices, the device can dynamically load or unload facial recognition functions as needed, eliminating the need for hardware upgrades and saving costs. By using an external TF card, the device does not need to occupy limited internal storage space, solving the problem of low-configuration devices being unable to support large-scale algorithm models. Users only need to insert a TF card to use the facial recognition function, which is simple and convenient, and can be uninstalled or removed at any time when the function is no longer needed.

[0012] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 A flowchart of a face recognition method for an intelligent intercom terminal based on an external TF card provided in one embodiment of the present application is shown;

[0014] Figure 2 A structural diagram of a face recognition device for an intelligent intercom terminal based on an external TF card provided in one embodiment of the present application is shown. DETAILED DESCRIPTION

[0015] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0016] The following combination Figure 1The following describes a face recognition method for an intelligent intercom terminal based on an external TF card according to an exemplary embodiment of the present application. It should be noted that the following application scenarios are only provided to facilitate understanding of the spirit and principles of the present application, and the embodiments of the present application are not limited in this respect. On the contrary, the embodiments of the present application can be applied to any applicable scenario.

[0017] In one embodiment, the present application also proposes a face recognition method and related equipment for an intelligent intercom terminal based on an external TF card. Figure 1 The following schematically shows a flow chart of a face recognition method for an intelligent intercom terminal based on an external TF card according to an embodiment of the present application. Figure 1 As shown, the method is applied to the server and includes:

[0018] S101, obtaining target TF card information, information of the application to be accessed, biometric information of the target user, and video image information of the target user.

[0019] In one implementation, target TF card capacity is crucial in the intercom field. For example, consider an intercom system used for security at large events. To expand facial recognition functionality, consider a 32GB TF card with 5GB already used, leaving 27GB. The remaining space must accommodate not only the existing facial recognition algorithm model (which may occupy several GB, depending on model complexity) but also sufficient space for the facial data of additional security personnel. For example, if each person's facial data and related feature information is estimated to occupy an average of 10MB, theoretically, the remaining space can store data for approximately 2,700 individuals. However, in practice, system runtime cache and other temporary file usage must also be considered, so careful planning is crucial. For intercom systems that require long-term storage of voice recordings (such as important call logs), TF card capacity also determines the duration of voice file storage. For example, using common voice coding formats, each minute of voice file may occupy approximately 1MB of space. With 27GB of remaining space, approximately 450 hours of voice recordings can be stored (not accounting for other factors).

[0020] Using FAT32 formatted TF cards in intercom devices offers wide compatibility, allowing most intercom devices and related operating systems to recognize and read / write. However, this format limits individual files to 4GB, which may be limiting for applications that require storing large facial recognition model files (even after optimization, they remain large) or large volumes of voice data files. The exFAT format, on the other hand, supports larger file sizes and is more suitable for storing larger models or data files. However, not all older intercom models fully support the exFAT format, so device compatibility must be ensured before use. For example, some newer smart intercoms implement system optimizations to ensure compatibility with large-capacity and diverse TF cards, ensuring stable operation with both FAT32 and exFAT formats. The TF card serial number is particularly useful in managing multiple cards in intercom devices. For example, in an intercom network within a large logistics park, multiple TF cards may be distributed across devices in different areas. The TF card serial number can accurately identify the specific intercom device where each card resides. For example, a TF card with serial number "123456789" is installed in the intercom equipment at the entrance of the park. Its usage history, including when the device was inserted and started working, and whether there were any data reading and writing anomalies, is recorded to facilitate maintenance and management of the equipment and TF card. At the same time, permissions are assigned based on the serial number. For example, TF cards with specific serial numbers can only be used on intercom equipment in certain designated areas, preventing TF card abuse and data security issues.

[0021] Take the "walkie-talkie dispatch and command application" as an example. In large-scale event security scenarios, the application's name clearly reflects its core role among numerous walkie-talkie functions. Its version "V2.1.0" may bring key improvements in this context. For example, previous versions could experience communication delays or delayed instructions when handling large numbers of security personnel using walkie-talkies simultaneously for dispatch and command. Version V2.1.0 improves dispatch and command efficiency by optimizing the communication protocol stack and algorithms. The interface design may more intuitively display the distribution and status of personnel in different security areas, facilitating quick decision-making by dispatchers. In the logistics and transportation dispatch sector, application version updates may be related to new transportation management processes or hardware device adaptation. For example, as newer walkie-talkie models are equipped with higher-resolution screens and more powerful processors, version V2.1.0 specifically optimizes the interface display, enabling better display of information such as transportation routes and cargo status on newer devices. It also improves the application's responsiveness, ensuring that dispatch instructions are promptly communicated to drivers and warehouse staff.

[0022] Group Management: The importance of group management in event security is self-evident. For example, groups with different functions and permissions can be created, such as "In-venue Security Group," "Out-venue Security Group," and "Emergency Response Group." The "Emergency Response Group" consists of members of the emergency response team in each area. Administrator permissions can be set to allow a single-click call to all members in an emergency, forcing all members' intercom microphones to turn on, ensuring rapid communication. This group is also given high priority, preventing calls from being interrupted by other regular groups. Call Recording and Replay: In logistics and transportation scenarios, the call recording and playback module is a crucial tool for resolving disputes and optimizing processes. For example, suppose a shipment of goods is damaged during transportation. By replaying conversations between the driver and warehouse shipping personnel, as well as other relevant personnel during transportation, it is possible to accurately determine the likely cause of the problem, such as improper loading procedures or unusual road conditions during transportation. In enterprise production management, by replaying conversations between workers in different areas of the workshop, it is possible to analyze any communication issues or operational errors within the production process, thereby optimizing the process. Location tracking module (if supported): In emergency rescue scenarios, such as mountain rescue operations, the location tracking module can display the location of rescuers in real time. The command center can rationally plan rescue routes based on location information to prevent rescuers from repeated searches or entering dangerous areas. At the same time, when rescuers encounter emergencies (such as injury or getting lost), the command center can quickly locate and dispatch nearby rescue forces to provide support. In offshore operation scenarios, for crew members equipped with walkie-talkies, the location tracking module, combined with the ship positioning system, can monitor the crew's activities on the ship in real time to ensure the crew's safety. In the event of an emergency such as someone falling into the water, the location of the incident can be quickly determined for rescue.

[0023] In large enterprise intercom applications, a standard employee user role might only allow voice calls with colleagues within their department during work hours, and may not be able to create or manage groups, preventing accidental operations from disrupting overall communication. A department manager role, on the other hand, can access the call logs of all employees in their department (for monitoring work progress) and create and manage temporary work groups within their department to facilitate project collaboration. Super administrator privileges allow comprehensive management of the entire enterprise intercom application, including adding or removing users, configuring system-level parameters (such as communication encryption and voice quality settings), and viewing call logs for all departments, ensuring system security and efficiency. In the trial version of some intercom applications, newly registered users may only have access to basic voice call features, without access to features like location tracking (if supported) and advanced group management. These advanced features are unlocked only after the user upgrades to a full paid user and the company purchases the corresponding service plan. For example, a company purchased a service package that includes a location tracking function for some senior managers. When these managers use the walkie-talkie application, they can access and use the location tracking function after identity verification, view the location information of their subordinates in real time, and improve management efficiency.

[0024] In a project that uses walkie-talkies across multiple brands, it's crucial for application developers to clearly list the supported models. For example, a general walkie-talkie dispatch and command application might support walkie-talkie models A, B, and C from brand X, and models D and E from brand Y, with specific hardware version requirements for each model. For example, brand X's walkie-talkie model A requires hardware version V3.0 or higher because the communication chip in this version has been upgraded to better support certain advanced features of the application, such as high-definition voice calls and fast data transmission. Using a model with a version lower than V3.0 may result in degraded voice quality and malfunctions. In practice, different brands and models of walkie-talkies may utilize different communication protocols. For example, some walkie-talkies use traditional analog protocols, while newer models support digital protocols such as DMR. Walkie-talkie dispatch and command applications require robust communication protocol adaptability. When the application is connected to a walkie-talkie using an analog communication protocol, the built-in analog communication protocol conversion module converts the application's commands and data into analog signals for transmission. When connected to a walkie-talkie supporting a digital communication protocol, the digital communication protocol is directly used for efficient data exchange. At the same time, the application also needs to be able to automatically identify the communication protocol type of the walkie-talkie to ensure seamless communication between different devices. For example, in a scenario where both old analog walkie-talkies and new digital walkie-talkies are used in a mixed environment, the application can intelligently switch and adapt to different protocols to ensure that all walkie-talkies can normally access the dispatch and command system, achieving unified management and command.

[0025] The target user's biometric information includes the following: Facial feature information: If it is a smart intercom device based on the facial recognition function extension, it will obtain the target user's facial image data and extract the facial feature vector through the facial recognition algorithm. For example, the facial feature vector may be a 128-dimensional or higher-dimensional numerical vector, which can represent the unique features of the user's face, such as the distance between the eyes, the shape of the nose, the quantitative representation of facial contours and other features. Fingerprint feature information (if the device supports the fingerprint recognition function extension): The fingerprint image of the target user is obtained through the fingerprint recognition sensor, and then the fingerprint feature point information is extracted, such as the position, direction and type of the fingerprint's ridges, valleys, bifurcation points and other feature points. These feature point information constitute the fingerprint feature template, which is used for subsequent fingerprint comparison and identity recognition.

[0026] Smart intercom devices are typically equipped with a camera. When a user approaches the device for authentication, the camera captures the user's face and transmits the image to the facial recognition module. The facial recognition module processes the image using a pre-trained facial recognition algorithm. It first detects the face to determine the position and posture of the face in the image, then performs feature extraction, converting the facial image into a feature vector. For example, deep learning-based facial recognition algorithms, such as convolutional neural networks (CNNs), perform multi-layer convolution and pooling operations on the facial image to ultimately generate a feature vector. The fingerprint recognition sensor on the device captures the user's fingerprint image. After preprocessing the fingerprint image (such as filtering and contrast enhancement), a fingerprint feature extraction algorithm is used to extract feature point information. Common fingerprint feature extraction algorithms include minutiae-based algorithms, which detect changes in the ridges and valleys in the fingerprint image to determine the location and type of feature points and construct a fingerprint feature template.

[0027] The target user's video image information includes the following: Video frame data: This is a digital representation of the original video signal input by the user through the camera of the smart intercom device. It records the pixel information of the video image at different time points and is the most basic form of video data. For example, the video signal is sampled at a certain frame rate (such as 25fps) and resolution (such as 640×480 pixels) to obtain a series of discrete video frames, which constitute the video frame data. Video image content (obtained through image recognition technology processing): The image recognition algorithm is used to analyze the video frame data to identify elements such as people, objects, and scenes in it, so as to perform operations such as semantic understanding and event judgment. For example, a user appears in the video standing at the entrance of a community. The image recognition system identifies key elements such as the user's image and the community gate. This can help the system understand the user's environment and behavior, and further perform permission judgment and business processing.

[0028] The camera in a smart intercom device captures the user's video signal in real time. After analog-to-digital conversion (A / D conversion) and necessary image signal processing (such as filtering and contrast enhancement), the resulting video frame data is stored in a buffer, supported by the device's video driver and relevant image processing libraries, pending further processing. The video frame data is then processed using an image recognition engine. This engine typically utilizes algorithms based on deep learning models (such as convolutional neural networks (CNNs)). It first extracts features from the video frame data, such as color, texture, and shape. These features are then fed into the image recognition model, which, based on training data and scene knowledge, outputs an understanding of the video image content. Smart intercom devices can integrate a mature image recognition SDK (software development kit) to implement video image analysis.

[0029] S102: Process the target user's biometric information to generate the target user's identity attribute information.

[0030] In one embodiment, the target user's biometric information is processed to generate initial facial image information and the target user's identity information. When the target user approaches a device with biometric recognition capabilities (such as a smart access control system or an intercom with facial recognition), the device's camera captures the user's facial image, which represents the facial image portion of the acquired biometric information. This facial image is then initially processed using an image processing algorithm. First, face detection is performed to determine the position and pose of the face in the image. For example, a cascade classifier algorithm based on Haar features or a deep learning-based face detection algorithm (such as SSD-Face) is used to locate the face region in the image and separate it from the background, generating initial facial image information. Simultaneously, the extracted facial image is compared with facial templates pre-stored in a database. The database contains facial images of registered users and their corresponding identity information (such as name, employee number, and user ID). A feature matching algorithm (such as feature matching based on Euclidean distance or feature vector matching using deep learning) is used to calculate the similarity between the input facial image and the facial templates in the database. If the similarity exceeds the set threshold, the match is considered successful, and the identity information of the target user is obtained.

[0031] Assuming the original image resolution captured by the camera is 1280×720 pixels, after face detection and cropping, the initial facial image information may be a 256×256 pixel grayscale image (for example only), which only contains the face, removing irrelevant information such as the background. If the database stores a face template for a user named "Zhang San" and employee number "001," when the target user Zhang San approaches the device, a successful comparison is performed, and the target user's identity information obtained is "Name: Zhang San, Employee Number: 001."

[0032] The initial facial image information is processed to generate a data feature set, which is used to characterize the facial features corresponding to different facial expressions. A feature extraction algorithm is then used to process the initial facial image information. For example, traditional feature extraction methods such as the Local Binary Pattern (LBP) algorithm or the Scale-Invariant Feature Transform (SIFT) algorithm, or a deep learning-based convolutional neural network (CNN) can be used to extract facial features. For different expressions, such as happiness, anger, sadness, and surprise, facial features in areas such as the eyes, eyebrows, and mouth vary significantly. By using a large amount of training data (including facial images with different expressions and their annotated expression categories), the feature extraction algorithm can learn the relationship between these expressions and facial features. The extracted facial features are combined to form a data feature set, where each element corresponds to a facial feature representation of a specific expression.

[0033] For example, based on LBP features, for a happy expression, features such as narrowed eyes (corresponding to a specific pattern of LBP eigenvalues around the eyes) and raised corners of the mouth (corresponding to a specific pattern of LBP eigenvalues in the mouth region) might be extracted. For an angry expression, features such as furrowed eyebrows (a specific pattern of LBP eigenvalues in the eyebrow region) and widened eyes (a specific pattern of LBP eigenvalues in the eye region) might be extracted. These features of different expressions are combined to form a data feature set. Assume that the data feature set is represented by a matrix, with each row representing a feature vector for a different expression. For example, the feature vector for a happy expression is [0.1, 0.3, 0.2, …] (this is just an example), and the feature vector for an angry expression is [0.4, 0.1, 0.5, …], and so on.

[0034] The data feature set is classified to generate target facial classification information. A classification algorithm is used to classify the data feature set. Common classification algorithms include support vector machines (SVMs), decision trees, and neural networks. Each feature vector in the data feature set is input, and the classification algorithm uses a trained model to determine which expression category it belongs to. For example, in a system using an SVM classifier, the SVM model learns the boundaries between different expression feature vectors and classifies the input feature vectors into corresponding expression categories, thereby generating target facial classification information. This determines the most likely expression category for the current facial image. Suppose the classification result is "happy." This means that the classification process determines that the expression corresponding to the current facial image is most similar to the happy category. If the system also provides a classification confidence score, perhaps "happy (confidence: 0.85)," this indicates that the system has a high degree of confidence in the classification result. The confidence score is calculated by the classification algorithm based on factors such as the distance between the feature vector and the model for each category.

[0035] The target facial classification information is processed to generate facial feature information and lighting information that matches the facial feature information. Based on the target facial classification information (e.g., expression categories such as happiness and anger), the facial feature information in the facial image is further analyzed. Geometric model-based methods can be used to extract facial features, such as determining parameters such as the position, shape, and size of facial features such as the eyes, nose, and mouth. For example, in a happy expression, the eyes may appear squinting. Geometric features such as the degree of eye openness (e.g., the ratio of eye width to height) and the curvature of the eye corners can be used to describe the eye features. For an upturned mouth, parameters such as the angle of the mouth corners and changes in lip thickness can be used to describe the mouth features. Similarly, the relative position and shape changes of the nose (e.g., the degree of expansion of the nostrils) can be determined as the expression changes.

[0036] The image also analyzes lighting information. Lighting information includes intensity and direction. Light intensity is estimated using methods such as the image's brightness histogram. For example, a larger proportion of areas with higher pixel values in the brightness histogram indicates stronger light intensity. Light direction is also inferred by analyzing shadows in different parts of the face. For example, a shadow on one side of the nose may indicate light coming from the other side. The extracted facial feature information is matched with the lighting information to determine the specific characteristic patterns exhibited by the facial features under the current lighting conditions. Facial feature information may include: eye characteristics (opening degree: 0.3, eye corner curvature: 0.5), mouth characteristics (mouth corner angle: 45 degrees, lip thickness change: +0.2 cm), nose characteristics (nose wing flare: 0.1), etc. (The values here are examples only; actual calculations will depend on the specific algorithm and image data). Lighting information may include: light intensity (average brightness value: 150, brightness range: 100-200) and light direction (horizontal angle: 45 degrees, vertical angle: 30 degrees). These facial feature information and light information together reflect the facial feature status of the target user under the current expression and lighting conditions.

[0037] The facial feature information and the lighting information that matches it are processed to generate factors influencing the target user's facial expressions. This model, based on a pre-established facial feature model, considers the impact of facial features, lighting, and their interactions on facial features. For example, a model trained with extensive experimental data and machine learning algorithms determines which facial feature combinations are more likely to correspond to which expressions under certain lighting directions and intensities. The facial feature information and lighting information are fed into the facial feature model for calculation. The model analyzes how facial features change under current lighting conditions and the degree of correlation between these changes and different expressions, thereby identifying the key factors influencing facial features. For example, if side lighting results in significant eye shadows and significant changes in mouth shape, the model may determine that lighting has a certain impact on facial recognition, with changes in the eyes and mouth being the primary factors influencing facial features. Assume that the generated factors influencing the target user's facial features are: facial features (primarily changes in the eyes and mouth) and lighting effects (side lighting reduces the recognition of some facial features). This shows that in the current situation, the morphological changes of the eyes and mouth are the main features for expressing facial expressions, and the lighting conditions affect the accurate recognition of these features to a certain extent, but the overall expression is still mainly determined by the facial features.

[0038] Based on a preset emotion recognition model, the factors influencing the target user's facial features are processed to generate target facial image information, which includes information about the target user's facial expression changes. The preset emotion recognition model is typically a deep learning model, such as a convolutional neural network (CNN) or a recurrent neural network (RNN), trained with a large amount of training data. Factors influencing facial features (including facial features and lighting information) are preprocessed to meet the model's input requirements. For example, facial features and lighting information are encoded and converted into numerical vectors acceptable to the model. This preprocessed information is then input into the preset emotion recognition model. Based on this input information and its learned relationship between facial expressions and emotions, the model predicts the target user's facial expression changes. The model outputs a description of the facial expression change, such as the transition from calm to happiness. This information includes the magnitude of the change (e.g., the gradual increase in happiness) and the speed of the change (e.g., the time it takes for the expression to reach its peak). This information constitutes the facial expression change component of the target facial image information. Suppose the model outputs the following information about the facial expression change in the target face image: the expression starts out calm, then gradually changes to happiness over 2 seconds, reaching a happiness level of 80%. (The happiness level here can be quantified using a number, such as a value between 0 and 1, with 1 representing the highest happiness.) The model may also output detailed features about the expression change, such as the speed at which the eyes gradually narrow and the trend of the upward angle of the mouth corners, to more fully describe the dynamic process of the expression.

[0039] The target facial image is processed to generate emotional information about the target user. Based on the facial expression changes in the target facial image, the corresponding emotion category and intensity are mapped. This may require an emotion map or a machine learning-based emotion classification model. For example, if the expression changes rapidly from calm to joyful with a high degree of joy, the corresponding emotion can be determined as "joy (intensity: high)" using an emotion map. If an emotion classification model is used, the model matches the expression change information with patterns in the training data and outputs the most likely emotion category and intensity. Emotion categories include joy, anger, sadness, surprise, fear, etc., and intensity can be expressed as low, medium, or high, or quantified using a specific numerical range (such as 0-1). Assuming the generated target user emotion information is "joy (intensity: 0.8)," it indicates that the target user is currently in a high degree of joy. This emotional information can be used in subsequent application scenarios. For example, in customer service, if a customer's joy is detected, service personnel can provide more positive interactions. In security monitoring, if a person displays unusual anger or fear, further attention may be warranted.

[0040] Generate the target user's identity attribute information based on their identity information and emotional information. Integrate the target user's identity information (such as name, employee number, user ID) and emotional information (such as the type and intensity of emotions such as joy, anger, and sadness). This information can be stored in a data structure, such as a structure or object containing identity and emotion fields. This integrated information can be further processed based on business needs. For example, in internal management, if an employee frequently expresses anger during work hours, they might be labeled as emotionally unstable. This label is then associated with the employee's identity information to form the target user's identity attribute information. In marketing, if a customer expresses joy while browsing a product, they might be labeled as a potentially highly satisfied customer. This information, combined with the customer's identity information, can be used as the target user's identity attribute information for subsequent precision marketing or customer relationship management.

[0041] For example, if the target user's identity information is "Name: Li Si, Employee ID: 002" and the emotion information is "Angry (Intensity: 0.6)", the generated target user's identity attribute information might be "Employee Li Si (Employee ID 002) is emotionally unstable and highly angry in the current scenario, and their work status requires attention." Alternatively, in a customer scenario, if the target user's identity information is "Customer ID: 12345" and the emotion information is "Joyful (Intensity: 0.8)", the generated target user's identity attribute information might be "Customer 12345 is highly satisfied and is a potential high-quality customer, and relevant high-end products or services can be recommended."

[0042] S103: Process the target TF card information to generate a target identity recognition model.

[0043] In one embodiment, the target TF card information is processed to generate the original data of the target TF card and a preset deep learning model based on a multi-layer perceptron MLP. When the target TF card is inserted, the system first reads the file system information of the TF card, including the file directory structure, file type, etc. For example, the specific folder path of the TF card storing face recognition data is identified. These paths are traversed to obtain various data files stored in the TF card, such as face image files (which may be common image formats such as JPEG, PNG, etc.), possible related configuration files (containing data format, resolution, etc.) and pre-stored deep learning model files based on a multi-layer perceptron MLP (the model file format may be a specific framework format, such as TensorFlow's SavedModel format or PyTorch's .pth format, etc.). These constitute the original data of the target TF card. At the same time, the preset deep learning model based on the multi-layer perceptron MLP is loaded from the TF card into the system memory for subsequent training or application. The model is already built in the TF card and includes the MLP network structure definition (such as the number of neurons and connection methods in the input layer, hidden layer, and output layer) and the initial parameter values (through previous training or pre-defined).

[0044] Suppose there is a folder named "face_data" on a TF card, containing 1,000 facial images of different people, with a resolution of 128×128 pixels and JPEG format. There is also a folder named "model," which contains an MLP-based deep learning model file named "mlp_face_model.pb." Reading these files and folders, the raw data retrieved from the target TF card is a combination of these image files and model files. The loaded MLP model might have an input layer with 128×128 neurons (corresponding to the number of image pixels), two hidden layers with 256 and 128 neurons, respectively, and an output layer with the number of categories representing different identities (e.g., 100 categories corresponding to 100 different people). The model parameters are initialized from previously trained values.

[0045] Preprocess the raw data from the target TF card to generate target TF card data. Perform image normalization on the acquired facial image files (a portion of the raw data from the target TF card). For example, normalize the image pixel values to a range between 0 and 1. This helps improve the stability and accuracy of subsequent processing. This can be achieved by dividing each pixel value by the image's maximum pixel value (255 for 8-bit images). Crop and scale the images to ensure that all images have the same size and pose. For example, crop the images to a standard size of 112×112 pixels and use image transformation techniques (such as affine transformation) to ensure that key facial features such as the eyes, nose, and mouth are in similar positions and angles across all facial images. If the raw data contains a configuration file, read the relevant information, such as whether data augmentation (such as random flipping, rotation, or noise addition) is required, and perform the corresponding processing according to the configuration. Data augmentation can increase the diversity of training data and improve the generalization ability of the model. For example, if the configuration file specifies random horizontal flipping, horizontal flipping is performed on some images to generate new training samples. After preprocessing, the pixel values of the 1,000 face images are all between 0 and 1, and the size is uniformly 112 × 112 pixels. Assuming that data augmentation is performed, 500 of the images are randomly horizontally flipped. The resulting target TF card data will contain 1,500 preprocessed and data augmented face images, which can be more conveniently used for subsequent feature extraction and model training.

[0046] Perform feature extraction on the target TF card data to generate a facial feature vector. Use the convolutional layer in the deep learning model (if the MLP model includes a convolutional layer, or use a separate convolutional neural network for feature extraction before inputting it into the MLP) to extract features from the preprocessed facial image. The convolutional layer slides a convolution kernel over the image to extract local features, such as edges and texture. For example, a 112×112 pixel facial image is convolved using multiple convolution kernels of different sizes and parameters to generate a series of feature maps. Pooling operations, such as max pooling or average pooling, are performed on the feature maps output by the convolutional layer to reduce the resolution and data size while preserving key features. For example, a 2×2 max pooling operation can be used to reduce the size of the feature map by half. The pooled feature map is then flattened to convert it into a one-dimensional vector, which is the facial feature vector. For example, if convolution and pooling result in a 10×10×64 (width×height×number of channels) feature map, flattening it yields a facial feature vector of length 10×10×64=6400. For each preprocessed 112×112 pixel face image, the feature extraction process described above generates a facial feature vector of length 6400 (for example only). These 1500 images (target TF card data after data augmentation) will each generate 1500 facial feature vectors. These vectors more effectively represent the characteristics of each face, have lower dimensionality, and contain more critical recognition information than the original image data.

[0047] Process the facial feature vectors to generate training and validation sets. Divide the facial feature vectors into training and validation sets according to a specific ratio (e.g., 80:20 or 70:30). For example, if there are 1500 facial feature vectors, divide them into 80:20, with 1200 vectors used for the training set and 300 vectors used for the validation set. During the division process, ensure that the distribution of the training and validation sets is representative, meaning that they contain feature vectors from different individuals and that the proportions between categories are relatively balanced (for classification tasks, such as identifying different individuals). Random sampling can be used for division, but care must be taken to ensure randomness and independence to avoid data leakage and bias. For classification tasks, the person identity labels corresponding to the facial feature vectors in the training and validation sets (e.g., 0 to 99 represents 100 different individuals) are stored along with the feature vectors for use in model training and validation. The training set contains 1200 facial feature vectors and their corresponding 1200 person identity labels, while the validation set contains 300 facial feature vectors and their corresponding 300 identity labels. These datasets will be used to train and evaluate deep learning models based on multi-layer perceptrons (MLPs) to ensure that the models can accurately recognize facial features and classify people's identities.

[0048] A pre-set deep learning model based on a multi-layer perceptron (MLP) is processed based on the training and validation sets to generate a target identity recognition module. The facial feature vectors and corresponding identity labels from the training set are input into the pre-set deep learning model based on a multi-layer perceptron (MLP) for training. In the MLP model, the input layer receives the facial feature vectors, the hidden layers perform nonlinear transformations and combinations of the features, and the output layer outputs the predicted identity category. A loss function, such as the cross-entropy loss function (for classification tasks), is calculated between the model's prediction and the true identity label. Based on the loss function, the backpropagation algorithm is used to adjust the parameters of the MLP model (such as weights and biases) to minimize the loss function and improve model accuracy. This process is repeated for multiple training cycles (epochs) on the training set until the model converges or a pre-set training stopping condition (such as the maximum number of training cycles or the loss function value falling below a certain threshold) is reached.

[0049] During training, the model is regularly evaluated using the validation set. The facial feature vectors from the validation set are fed into the training model, and evaluation metrics such as precision and recall are calculated on the validation set. If the model's performance on the validation set stops improving or even deteriorates (possibly due to overfitting), training can be stopped early, and the model parameters that achieved the best performance can be selected as the final model parameters. After training and evaluation, the optimized deep learning model based on the multi-layer perceptron (MLP) is the target identity recognition module, which can be used to predict identity for new facial feature vectors. Assume that after 100 training cycles, the model achieves 95% accuracy on the validation set (for example only), and if the accuracy does not improve significantly with subsequent training, these model parameters are selected as the final parameters for the target identity recognition module. When a new facial feature vector is fed into the target identity recognition module, the module can calculate and output a predicted identity category based on these parameters, thereby implementing facial recognition capabilities. For example, in access control systems, this module can determine whether a user is authorized, or in surveillance systems, identify the identity of a person in a video.

[0050] S104: Process the application information to be accessed to generate application access purpose information of the target user and access permission information of the target user.

[0051] In one embodiment, the application information to be accessed is processed to generate the target user's access information and the target user's access feedback information. When the target user attempts to access an application called "walkie-talkie dispatch and command application", the system first obtains its relevant information. Assume that the application version number is 3.0, and the application source is customized and developed within the enterprise and distributed through the enterprise's dedicated server. The system checks the user's walkie-talkie device and finds that the application has been installed, but the version is 2.5. At the same time, the device operating system is an Android system customized for specific walkie-talkie devices, version 8.1, with 200MB of available memory, while the application requires at least 500MB of available memory to run, and the device's communication module firmware version is V1.2, and the application requires a minimum communication module firmware version of V1.5 (to better support certain advanced dispatch functions). The generated target user's access information is: the application has been installed but the version is outdated, the device operating system meets the requirements but the available memory is insufficient, and the communication module firmware version is too low. The corresponding access feedback information is: "Your version of the 'Intercom Dispatch and Command Application' is too low. Please go to the company's internal server to download and update to version 3.0. At the same time, your device has insufficient available memory. It is recommended to close some unnecessary background applications or clear the cache to free up memory. In addition, the firmware version of your device's communication module is too low and needs to be upgraded to V1.5 or above to fully use all the functions of the application. The firmware upgrade can be performed in the system update option in the device settings (if available), or contact the device administrator for assistance with the upgrade."

[0052] The target user's access information is processed to generate access rights information for the target user. Based on the target user's access information, the system queries the user rights management system for relevant rules. For example, let's assume the target user is a general security officer in the security department. For the "walkie-talkie dispatch and command application," the rights management system stipulates that general security officers can only use camera-equipped walkie-talkies for group video calls during designated working hours (e.g., 8:00 AM to 8:00 PM). They can only join groups within their security zone (e.g., "Venue Area A Security Group" and "Venue Area B Security Group") and cannot create or manage groups. They can receive emergency notifications and instructions from higher-level dispatchers but cannot send global broadcasts. They can only view safety plan documents related to their position and cannot modify or delete any documents. The generated access rights for the target user are: Permitted operations—group video calls (during working hours, limited to the group within their security zone), receiving emergency notifications and instructions, and viewing relevant safety plan documents; Prohibited operations—creating and managing groups, sending global broadcasts, and modifying or deleting documents.

[0053] The target user's access feedback is processed to generate information about the target user's application access purpose. The system performs semantic analysis and intent recognition on this access feedback (prompts to update the application version, free up memory, and upgrade the communication module firmware). Since the prompts primarily focus on application version updates and device resource and firmware upgrades, the system infers that the user's primary purpose for using the application is to enable efficient video communication and collaboration in security work scenarios. The user may need to receive dispatch instructions promptly and collaborate with colleagues on security incidents. However, device resource and firmware issues may affect the use of some advanced features of the application, such as HD video calls and real-time location sharing (if these rely on the communication module firmware). Video image analysis features (such as person recognition and behavioral analysis) may also be limited. The generated application access purpose information indicates that the user's primary purpose in the security work scenario is to enable real-time video communication and collaboration (such as receiving instructions and communicating with colleagues about security situations). However, due to the current device state, the user may not be able to fully utilize the application's advanced features. Therefore, the user needs to update the application, free up memory, and upgrade the communication module firmware as soon as possible to achieve a better work experience. This information helps the system prioritize the stability of basic video communication functions during subsequent user use, prompt users to upgrade related functions when device resources permit, or guide users to experience new security dispatch functions after application updates and device optimization (such as more accurate personnel positioning based on upgraded communication module firmware and abnormal behavior detection based on video image analysis).

[0054] S105 : Process the target user's application access purpose information and the target user's access permission information to generate access permission information for the target application.

[0055] In one embodiment, the target user's access permission information is processed to generate access time period information for the target user. This access time period information represents the duration of a single access and the time period during which the target user can access the device. In an enterprise's intercom usage management system, the target user's access permission information is stored in a dedicated permissions database. For example, for ordinary security personnel, the system stipulates that they can use camera-equipped intercoms for video communication during their daily work hours (Monday to Friday, 7:00 AM to 7:00 PM), with a single continuous usage limit of no more than 1.5 hours to avoid excessive battery drain and prolonged device usage that could affect other personnel. After reading the permission record from the database, the system parses the security personnel's access time period (Monday to Friday, 7:00 AM to 7:00 PM) and the single access time limit (1.5 hours) to generate the access time period information. During special events, such as large-scale exhibitions, some security personnel may be assigned to specific areas for special tasks and may need additional intercom access during the exhibition period (e.g., Saturdays and Sundays, 8:00 AM to 10:00 PM). Based on the temporary permissions, the system will incorporate these special time periods into the security personnel's access time information. For example, security personnel "Zhao Liu" has regular access from 7:00 AM to 7:00 PM, Monday to Friday, with a single access time limit of 1.5 hours. Temporary permissions granted for the exhibition extend to Saturdays and Sundays from 8:00 AM to 10:00 PM. The generated access time information for "Zhao Liu" is: Accessible from 7:00 AM to 7:00 PM, Monday to Friday, and from 8:00 AM to 10:00 PM, Saturdays and Sundays; a single access time limit of 1.5 hours.

[0056] The system processes the target user's application access purpose information to generate their access priority. This information can be obtained from a variety of sources. For example, in a company's walkie-talkie dispatch application, if a security officer selects the "Patrol Area Emergency Response" task upon login, the system will determine that their application access purpose is highly urgent. Alternatively, if the system analyzes a security officer's recent frequent use of the walkie-talkie to contact the monitoring center, including operations such as suspicious person screening, it can be inferred that their application access purpose is to handle important security monitoring tasks and should be given a higher priority. The system also considers the user's department and position. For example, the access priority of an emergency rescue team performing a rescue mission is higher than that of an ordinary security officer conducting routine patrol communications. The system assigns a priority value or level to the target user based on a pre-set priority rule algorithm. Priorities can be set from 1 to 5. Users performing critical tasks such as emergency rescue and handling urgent security incidents might be assigned priority 1, while ordinary patrol security personnel might be assigned priority 3 or 4. For example, security guard "Sun Qi" works for the emergency rescue team. When logging into the intercom dispatch application, he selects the "Earthquake Rescue Site Communications" task. The system analyzes his recent involvement in various emergency rescue drills and actual rescue operations. Based on priority rules, emergency rescue is a high-priority task. Therefore, the system assigns "Sun Qi" a priority level of 1, indicating that he should have priority access to intercom resources and communication responses during this access, ensuring timely delivery of rescue instructions and rapid feedback on the rescue situation.

[0057] The access priority of the target user is processed to generate real-time access channel information for the target user. The intercom system has multiple predefined access channels, each with different resource allocations and functional features. For example, there are HD video channels, standard video channels, and dedicated emergency call channels (assuming the intercom supports channels with different video communication quality). Based on the target user's access priority, the system assigns the user to the corresponding real-time access channel. High-priority users (such as those with priority levels 1 and 2) are typically assigned to HD video channels to ensure clear and stable video communication quality, facilitating efficient processing of emergency rescue and critical dispatch tasks. HD video channels also provide clearer identification of personnel and on-site conditions. Medium-priority users (such as those with priority levels 3) may be assigned to standard video channels, while low-priority users (such as those with priority levels 4 and 5) may be assigned to channels with lower resource usage (such as those that only provide basic video communication functions). This ensures the optimal allocation of system resources and avoids communication interruptions for high-priority users. The system monitors the load of each channel in real time. If an HD video channel becomes overpopulated with high-priority users, causing resource constraints, the system will dynamically relocate some high-priority users to other channels of the same level with more sufficient resources based on load balancing strategies. For example, security guard "Sun Qi" has access priority level 1. The system assigns him to the HD video channel for real-time access. At this point, the HD video channel may also be occupied by other earthquake rescue team members or command personnel. The system records the channel's user list in real time, generating real-time channel access information for users waiting to access the channel. For example, users waiting to access the HD video channel include "Sun Qi" and "Zhou Ba" (the rescue team leader). The channel's current resource utilization rate is 80%, indicating that there is still some capacity to accommodate more high-priority users, but the channel is nearing capacity.

[0058] The system processes the real-time channel information of the user being accessed and generates the user's access time period. This access time period is determined based on the characteristics and usage rules of the user's real-time channel. For example, due to high resource usage, the HD video channel has a single continuous usage limit of 45 minutes per user. During major events or emergencies (such as the earthquake rescue scenario described above), the channel is open 24 hours a day. During non-emergency periods (such as daily business operations), the channel is only open during business hours (e.g., 8:00 AM to 6:00 PM) and is closed or switched to a low-resource usage mode during non-business hours. For example, consider security guard "Sun Qi" accessing the HD video channel at 9:00 AM during earthquake rescue operations. According to the channel rules, his single access time limit is 45 minutes, and the channel is open 24 hours a day. Therefore, the generated access time period for "Sun Qi" is: accessible from 9:00 AM to 9:45 AM (if accessed continuously), and the overall channel availability period is from 0:00 AM to 12:00 AM (during the rescue period).

[0059] The target user's access time information is processed based on the access time information of the user to be accessed, generating access rights information for the target application. The system considers the target user's original access time information (based on permission settings) and their access time information in real-time channels (based on channel rules). For example, security guard "Sun Qi" originally had access to the intercom at any time during an emergency (such as an earthquake rescue scenario) (based on permission settings), while their access time information in the HD video channel was from 9:00 AM to 9:45 AM (based on channel rules). Taking the intersection of these two, Sun Qi's current access time is determined to be from 9:00 AM to 9:45 AM. The system also reassesses permissions based on the access time adjustment. If Sun Qi originally had full access to intercom functions (such as switching channels, adjusting volume, and initiating group video calls) during the rescue period, the system may temporarily restrict some non-critical operations (such as prohibiting switching to non-rescue-related channels) during this shortened time period to ensure communication is focused on the rescue mission and prevent misoperations that affect rescue efficiency. Based on the comprehensive processing results, access permissions for the target application (the walkie-talkie dispatch application) are generated, clarifying Sun Qi's specific radio operation permissions and access time range in the current situation. If he still needs to use the walkie-talkie to participate in the rescue after 9:45 AM, his permissions will be reassessed or he will need to reapply for access based on channel resources. (If channel resources allow, the access time period and corresponding permissions can be reassigned.)

[0060] S106: Process the access permission information of the target application and the identity information of the target user based on the target identity recognition model to generate video image information of the target intercom device.

[0061] In one implementation, the target application's access rights information is processed based on a target identity recognition model to generate video channel allocation information. For example, the target application is an internal enterprise intercom dispatch and command system. Access rights information includes the user's role (e.g., general security officer, security supervisor, system administrator), functional permissions (e.g., ability to initiate emergency calls, ability to make video calls with personnel in a specific area), and data access level (e.g., ability to access video channels associated with surveillance information in sensitive areas). The target identity recognition model first accurately identifies and verifies the user's identity through methods such as facial recognition, fingerprint recognition, or password verification (combined with the intercom device's identity recognition capabilities), ensuring a perfect match with the identity information stored in the system. Then, based on the user's identity and access rights information, video channel allocation information is determined according to pre-defined complex rules and algorithms. For example, general security officer "Li Si" is only authorized to make group video calls with security personnel in his or her assigned patrol area (e.g., Area A of the factory). The target identity recognition model retrieves the predefined list of video channels for the security group in Area A of the factory and assigns "John Doe" to the corresponding channel, such as the "Area A Daily Patrol Video Channel." If "John Doe" attempts to access channels in other areas or the management channel, the system will deny access. Security supervisors, in addition to accessing the security group video channels for their area, also have access to the management-only command and dispatch video channel for cross-regional security coordination and emergency decision-making. System administrators have access to all video channels, allowing them to perform maintenance and management operations in the event of system failures or when global settings are required. Furthermore, for security personnel with access to sensitive area surveillance data, the model ensures that they can only access video channels that match their permissions, preventing the leakage of sensitive information. For example, security personnel responsible for monitoring the data room can only access the warning and emergency response video channels related to the data room and are unable to access information from other unrelated channels.

[0062] Based on the video channel assignment information, the target user's identity information is processed to generate the target user's video image information. The system obtains the target user's detailed identity information, including name, employee number, department or position (e.g., Security Department - Area A Patrol Post), as well as the video channel assignment information generated in the previous step. For "Li Si," the system closely associates his name (Li Si), employee number (00234), and position (Area A Patrol Post) with his assigned "Area A Daily Patrol Video Channel." Furthermore, based on the default permissions granted to ordinary security personnel within the channel, the system specifies Li Si's specific operational permissions within the channel. These permissions include the ability to open the camera to report patrol status, view the video feeds of other security personnel within the channel, and view the list of currently online channel members. However, they are not allowed to perform channel management operations (e.g., adding or removing channel members, modifying the channel name or settings, etc.). This information is integrated to form the video image information of the target user "Li Si." In addition, it may also include some personalized settings, such as the video prompt sound is the default warning sound (to attract attention in noisy environments), the initial camera resolution is 640×480 (which can be adjusted by the user according to the actual usage environment), etc., to optimize the user experience in the channel.

[0063] The target user's video image information is processed to generate video image information for the target intercom device. When "John" logs into the intercom dispatch and command application using his intercom device, the application quickly transmits John's video image information to the intercom device. Upon receiving this information, the intercom device immediately performs a series of intelligent setup and configuration operations. The device clearly displays the name of the video channel available for John to select, namely "Area A Daily Patrol Video Channel," on the device screen. Simultaneously, based on John's permissions for that channel, the corresponding operation buttons on the device are intelligently enabled or disabled. For example, the "Start Video Call" button is enabled, allowing John to communicate and report on situations with colleagues via video at any time. The "Channel Management" button is disabled (grayed out or hidden) to prevent accidental operation. Simultaneously, the intercom device automatically establishes a stable connection with the application server and accurately sets video parameters based on the channel parameters provided by the server. For example, the video encoding format is set to H.264 (a commonly used high-efficiency video encoding format), the frame rate is set to 25fps (to ensure video smoothness while reducing data transmission volume), and the resolution is set to 640×480 (adjustable based on actual needs and network conditions) to ensure clear, smooth, and low-latency video interaction with other intercom devices in the "Area A Daily Patrol Video Channel," meeting the strict requirements of real-time communication in security patrol work. At this point, the intercom device's video image information is fully configured according to "Li Si's" permissions and needs, providing a solid guarantee for his efficient video communication in this channel.

[0064] S107 , processing the target user's emotional information and the target user's video image information based on the target intercom device's video image information to generate target event information.

[0065] In one embodiment, the target user's emotional information and video image information are processed based on the target intercom device's video image information to generate video call keyword information. Assume that the target intercom device is used in a company's production workshop scheduling scenario, and the video image information indicates the current channel is the "Workshop Production Scheduling Video Channel." The target user is a workshop worker who appears anxious and flustered during the video call (as determined by video image analysis). The system first performs behavioral and voice recognition on the target user's video image information, converting the voice signal into text and analyzing the meaning of their actions (for example, if a worker frequently points at a machine, their intention can be understood by combining the voice information). For example, if a worker says, "This machine is broken and keeps making strange noises. I can't continue working! I need someone to fix it!" Simultaneously, emotion recognition technology is used to analyze the target user's emotional information, determining that they are anxious or angry (by analyzing facial expressions, body movements, and multimodal voice features such as pitch, speed, and volume). Then, based on pre-set keyword extraction rules, key information is extracted from the recognized text and the information contained in the actions. In this example, extracted keywords might include "machine failure," "strange noise," and "repair assistance." Combined with behavioral analysis in the video, these keywords might also include "specific machine location (determined by the direction the worker is pointing)." These keywords are combined to form video call keyword information, reflecting the core content of the call, the user's emotional state, and related behavioral information. Video call keyword information: ["machine failure," "strange noise," "repair assistance," "anxiety-related characteristics (such as tense facial expressions, frantic body movements, high-pitched voice, rapid speech)," "specific machine location (such as machine tool No. 3 in workshop area A)"].

[0066] The video call keyword information is processed to generate target warning values and event information matching the target warning values. The system has a pre-defined keyword classification and weighting system. For example, keywords are categorized into categories such as equipment failure, safety incidents, and personnel needs. In this example, "machine failure" belongs to the equipment failure category. Assume that the keyword "machine failure" has a high base weight ω (assuming 8) within the set of equipment failure keywords because it directly indicates a problem with production equipment, potentially impacting production progress. Its frequency of appearance in the video call, f, is 1 (here, but this frequency will increase if the worker repeatedly emphasizes it). Since the current video channel is the "Workshop Production Scheduling Video Channel," the correlation coefficient r (assuming 0.9) between the equipment failure keyword and this channel is high, as this channel is primarily used for handling workshop production-related matters. The target user is experiencing anxiety and anger, with an emotional impact factor e (assuming 1.5, indicating strong emotions). At the same time, additional information such as the "specific machine location" extracted from the video will also serve as an important reference factor when calculating the warning value and determining event information. For example, if the location of the machine is at a critical link in the production process, the warning value may be further improved.

[0067] The method also includes a calculation formula for obtaining the target warning value, which is:

[0068] ; Where m represents the number of categories in the keyword set, Indicates the number of keywords contained in the j-th keyword set, represents the basic weight of the i-th keyword in the j-th keyword set, represents the frequency of occurrence of the i-th keyword in the j-th keyword set in the call video, represents the correlation coefficient between the j-th keyword set and the current video channel, and e represents the user sentiment influence factor. The target warning value V (assuming this is the only keyword) is calculated using the formula: V = r × (ω × f) × e = 0.9 × (8 × 1) × 1.5 = 10.8. Furthermore, based on the keyword "machine failure," the system matches the event information "Production equipment failure event, maintenance personnel must be dispatched to address it." For example, the target warning value is 10.8, and the event information is: "Production equipment failure event, maintenance personnel must be dispatched to address it."

[0069] The target warning value is processed based on the preset event warning table to generate a real-time event warning value. The preset event warning table specifies the corresponding handling methods and warning levels for different warning value ranges. For example, warning values between 0 and 10 are low risk, between 10 and 20 are medium risk, and above 20 are high risk. For the calculated target warning value of 32.4, the corresponding handling method and warning level are searched in the preset event warning table. Determining it as a high risk level may require immediate notification to the maintenance department manager and the dispatch of additional resources. Based on the rules in the warning table, the target warning value is adjusted or converted (if necessary) to generate a real-time event warning value. For example, 32.4 may be converted to a specific high-risk code (such as "HR-03," indicating a more severe high-risk situation) to facilitate subsequent unified processing and identification by the system. Real-time event warning value: "HR-03" (indicating a high-risk production equipment failure event).

[0070] If the real-time event warning value exceeds the preset warning threshold, an event warning message is generated. Assume the preset warning threshold is set to 20 (i.e., an alert is triggered when the real-time event warning value exceeds 20). Since the actual warning value of 32.4 corresponding to the real-time event warning value "HR-03" exceeds the preset warning threshold, the system generates an event warning message. The event warning message may include the warning level (high risk), event type (production equipment failure), location (specific location in the workshop, which can be located via intercom equipment or manually reported by a worker; in this case, it is production line 3 in area A of the workshop), and relevant personnel (the target user and their team, namely, worker Zhang San and the team of production line 3 in area A). The system may also send warning messages through various means, such as a prominent pop-up alert on the display screen of the workshop dispatch center, text messages, or push notifications to the mobile phones of relevant managers. Event warning message: {"Warning Level": "High Risk", "Event Type": "Production Equipment Failure", "Location": "Production Equipment Failure", "Related Personnel": "Worker Zhang San and the team of production line 3 in area A"}.

[0071] The event warning information is processed to generate target event information. Target event information represents the target user's video call information with other users. Other users are generated based on the video image information of the target intercom device. The system further refines and integrates relevant data based on the event warning information to generate target event information. In addition to the content in the event warning information, it may also include the event timestamp (accurate to the second, recording the time the equipment failure occurred, for example, 2023-11-10 14:30:15), a detailed description of the event (e.g., the specific symptoms of the machine failure, such as "the machine emitted an abnormal metallic friction sound, stopped operating, and smoke was emitted," based on the worker's description), and an assessment of the potential impact (e.g., the estimated duration of the production schedule impact is 3 hours, the potential product loss is 200 units, etc.). Target event information is stored in the system database for subsequent query, statistics, and analysis, providing a more comprehensive basis for further decision-making and processing. For example, the maintenance department can prepare appropriate repair tools and parts based on the target event information, and the production management department can adjust production plans based on the impact assessment, such as arranging temporary overtime on other production lines to compensate for lost production. For example, target event information: {“Event time”: “2023-11-10 14:30:15”, “Warning level”: “High risk”, “Event type”: “Production equipment failure”, “Occurrence location”: “Production line 3, area A, workshop”, “Related personnel”: “Worker Zhang San and the team of production line 3, area A”, “Detailed description of the event”: “The machine made an abnormal metal friction sound, stopped running, and there was smoke”, “Possible impact assessment”: “It is expected to affect the production progress for 3 hours, and may cause the loss of 200 products”}.

[0072] This application requires the server to obtain target TF card information (including capacity, format, serial number, etc.), information about the application to be accessed (such as the version, functional modules, permission requirements, compatibility, etc. of the walkie-talkie dispatch command application), target user biometric information (face or fingerprint features), and video image information (waveform data and text content). Next, the biometric information is processed to generate identity attribute information, including generating identity information through detection and comparison, and then obtaining emotional information through feature extraction and classification. The TF card information is processed to generate a target identity recognition model, which involves reading raw data, preprocessing, feature extraction, data set division, and model training. The application information to be accessed is processed to obtain application access purpose and access permission information, such as generating access feedback information based on the device situation, and then inferring the access purpose based on this information.

[0073] Then, the target application access permission information is generated by combining the application access purpose and permission information. This includes determining the access time period and priority, allocating real-time access channels, and ultimately determining permissions. Intercom device video image information is generated based on the target identity recognition model and related information, including allocating video channels and configuring devices. Finally, target event information is generated by processing user emotions and video image information based on the intercom device video image information. Keywords are first extracted to calculate the warning value, which is then processed through the warning table. If the threshold is exceeded, event warning information is generated. Further refinement is performed to obtain the target event information (including time, location, personnel, impact assessment, etc.) for subsequent query, decision-making, and processing. This overall approach helps improve the security, functionality, and management efficiency of smart intercom terminals.

[0074] In one embodiment, Figure 2 As shown, the present application also provides a face recognition device for an intelligent intercom terminal based on an external TF card, comprising:

[0075] Acquisition module 201, used to obtain target TF card information, information of the application to be accessed, biometric information of the target user, and video image information of the target user;

[0076] The processing module 202 is used to process the biometric information of the target user to generate the identity attribute information of the target user, wherein the identity attribute information of the target user includes the identity information of the target user and the emotional information of the target user; process the target TF card information to generate a target identity recognition model; process the application information to be accessed to generate the application access purpose information of the target user and the access permission information of the target user; process the application access purpose information of the target user and the access permission information of the target user to generate the access permission information of the target application; process the access permission information of the target application and the identity information of the target user based on the target identity recognition model to generate the video image information of the target intercom device; process the emotional information of the target user and the video image information of the target user based on the video image information of the target intercom device to generate target event information, wherein the target event information is used to represent the voice call information between the target user and other users, and the other users are generated based on the video image information of the target intercom device.

Claims

1. A face recognition method for an intelligent intercom terminal based on an external TF card, characterized in that: include: Obtain target TF card information, application information to be accessed, target user's biometric information, and target user's video image information; Processing the target user's biometric information to generate the target user's identity attribute information, wherein the target user's identity attribute information includes the target user's identity information and the target user's emotional information; Process the target TF card information and generate a target identity recognition model; Processing the application information to be accessed to generate the target user's application access purpose information and the target user's access permission information, including processing the application information to be accessed to generate the target user's access information and the target user's access feedback information, where the application to be accessed is a walkie-talkie dispatch and command application; processing the target user's access information to generate the target user's access permission information, where the target user's access permission information is set based on the target user's job position permissions; processing the target user's access feedback information to generate the target user's application access purpose information, where the application access purpose information is used to indicate the need to receive dispatch instructions in a timely manner and collaborate with colleagues to handle security incidents; Processing the target user's application access purpose information and the target user's access permission information to generate the target application's access permission information, including processing the target user's access permission information to generate the target user's access time period information. The target user's access permission information is dynamically adjusted based on the target user's job responsibilities. Processing the target user's application access purpose information to generate the target user's access priority. Processing the target user's access priority to generate real-time access channel information for the user to be accessed. The access channel information includes high-definition video channels, ordinary video channels, and emergency call-dedicated channels. The target application's access permission information and the target user's identity information are processed based on the target identity recognition model to generate video image information of the target intercom device; The target user's emotional information and the target user's video image information are processed based on the target intercom device's video image information to generate target event information.

2. The method according to claim 1, wherein Process the target user's biometric information to generate the target user's identity attribute information, including: Processing the target user's biometric information to generate initial facial image information and the target user's identity information; Processing the initial facial image information to generate a data feature set, wherein the data feature set is used to represent facial features corresponding to different expression information; Classify the data feature set to generate target face classification information; Processing the target face classification information to generate facial feature information and light information matching the facial feature information; Processing facial feature information and light information that matches the facial feature information to generate factors influencing the target user's facial expression characteristics; Processing the influencing factors of the target user's facial features based on a preset emotion recognition model to generate target facial image information, wherein the target facial image information includes the target user's facial expression change information; Process the target face image information to generate the target user's emotional information; The target user's identity attribute information is generated based on the target user's identity information and the target user's emotional information.

3. The method according to claim 1, wherein Process the target TF card information to generate a target identity recognition model, including: Process the target TF card information to generate the target TF card's raw data and a preset deep learning model based on the multi-layer perceptron (MLP); Preprocess the original data of the target TF card to generate target TF card data; Perform feature extraction on the target TF card data to generate a facial feature vector; Process the facial feature vector to generate training set and validation set; The preset deep learning model based on the multi-layer perceptron MLP is processed based on the training set and the validation set to generate a target identity recognition module.

4. The method according to claim 1, wherein The target user's application access purpose information and the target user's access permission information are processed to generate the target application's access permission information, including: Processing the real-time channel information of the user to be visited to generate the visiting time period information of the user to be visited; The access time period information of the target user is processed based on the access time period information of the user to be accessed, and access permission information of the target application is generated.

5. The method according to claim 4, wherein The target application's access permission information and the target user's identity information are processed based on the target identity recognition model to generate video image information of the target intercom device, including: Processing the access permission information of the target application based on the target identity recognition model to generate video channel allocation information; Processing the target user's identity information based on the video channel allocation information to generate the target user's video image information; The video image information of the target user is processed to generate the video image information of the target intercom device.

6. The method according to claim 5, wherein The target user's emotional information and the target user's video image information are processed based on the target intercom device's video image information to generate target event information, including: Processing the target user's emotional information and the target user's video image information based on the target intercom device's video image information to generate call video keyword information; Processing call video keyword information to generate target warning values and event information matching the target warning values; Process the target warning value based on the preset event warning table to generate a real-time event warning value; If the real-time event warning value is greater than the preset warning threshold, an event warning message is generated; Process event warning information to generate target event information; The method also includes a calculation formula for obtaining the target warning value, which is: ; Among them, m represents the number of categories of keyword sets, Indicates the number of keywords contained in the j-th keyword set, represents the basic weight of the i-th keyword in the j-th keyword set, represents the frequency of occurrence of the i-th keyword in the j-th keyword set in the call video, represents the correlation coefficient between the j-th keyword set and the current video channel, and e represents the user emotion influencing factor.

7. A face recognition device for an intelligent intercom terminal based on an external TF card, characterized in that: For implementing the method according to claim 1, the apparatus comprises: The acquisition module is used to obtain the target TF card information, the application information to be accessed, the biometric information of the target user and the video image information of the target user; The processing module is used to process the biometric information of the target user to generate the identity attribute information of the target user, wherein the identity attribute information of the target user includes the identity information of the target user and the emotional information of the target user; process the target TF card information to generate a target identity recognition model; process the application information to be accessed to generate the application access purpose information of the target user and the access permission information of the target user; process the application access purpose information of the target user and the access permission information of the target user to generate the access permission information of the target application; process the access permission information of the target application and the identity information of the target user based on the target identity recognition model to generate video image information of the target intercom device; process the emotional information of the target user and the video image information of the target user based on the video image information of the target intercom device to generate target event information.

8. An electronic device, characterized in that: include: a first processor; and a memory for storing executable instructions for the first processor; The first processor is configured to execute the face recognition method for an intelligent intercom terminal based on an external TF card according to any one of claims 1 to 6 by executing executable instructions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the second processor, the face recognition method for the smart intercom terminal based on the external TF card described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Identity authentication system with biological characteristic recognition function and authentication method thereof

    CN101986597A

  • Information processing method and device, electronic equipment and storage medium

    CN119007754A

  • Two-dimensional code recognition unlocking method suitable for entrance machine and related equipment

    CN119251939A