Visual SOP management system

By using multimodal fusion analysis of facial expressions, voice, and operational behavior, the system can identify operators' confusion in real time and automatically push auxiliary cases, solving the problems of incomplete status perception and untimely auxiliary response in the existing SOP management system, and realizing real-time control of operational procedures and efficiency improvement.

CN121504385APending Publication Date: 2026-02-10MINGWU SHUZHI TECH RES INST (NANJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511690118.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing SOP management systems lack the ability to capture operators' emotional state and operational fluency in real time, making it difficult to identify abnormal states, which increases the risk of operational errors. Furthermore, they lack automatic auxiliary support mechanisms, resulting in low response efficiency and difficulty in achieving precise control over operational procedures.

Method used

It employs facial expression capture, voice data capture, and operation behavior capture modules, combined with a multimodal fusion analysis model, to analyze the operator's level of confusion in real time. When a preset threshold is reached, it automatically pushes videos of similar operation cases, and performs real-time comparison and early warning in conjunction with SOP standards.

Benefits of technology

It enables accurate identification of operators' confused states, improves response speed and operational standardization, reduces learning costs, minimizes operational deviations, and enhances work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504385A_ABST
    Figure CN121504385A_ABST
Patent Text Reader

Abstract

The invention discloses a visual SOP management system, which belongs to the field of computer systems, integrates three dimensions of facial expression recognition, voice emotion analysis and operation fluency detection, realizes quantitative scoring of the confusion degree of an operator through a multi-modal fusion analysis model, and is higher in recognition accuracy compared with single data monitoring. Abnormal states such as hesitation and lagging can be accurately captured; when the confusion degree reaches a preset threshold value, three similar operation case videos with the highest matching degree are automatically called and pushed according to a standard operation priority sequence, manual intervention is not needed, the response speed is greatly improved compared with traditional manual guidance, meanwhile, case matching is combined with an SOP step number, the operation scene similarity and the difficulty coefficient, and the operation efficiency is improved. The matching accuracy is guaranteed through a cosine similarity algorithm, key corresponding nodes are marked at the same time, operators are helped to quickly understand standard operation key points, the learning cost is reduced, and the working efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer systems, and more specifically, to a visual SOP management system. Background Technology

[0002] In fields such as industrial production, medical operations, and logistics warehousing, where strict adherence to Standard Operating Procedures (SOPs) is essential, the standardized execution of SOPs is a core prerequisite for ensuring product quality, operational safety, and process consistency. With the development of intelligent technologies, while traditional SOP management methods relying on paper documents and static videos offer various advantages, they also have the following shortcomings, making it difficult to meet the demands of modern production scenarios for real-time monitoring, dynamic optimization, and precise control of operational processes.

[0003] For example, existing SOP management systems mostly focus on displaying and recording operational steps, but lack the function of capturing the operator's emotional state and operational fluency in real time. This makes it difficult to accurately identify abnormal states such as confusion and hesitation regarding SOP steps, leading to an increased risk of operational errors. When operators experience operational delays or errors due to unfamiliarity with or misunderstanding of SOP procedures, the existing system lacks a mechanism to automatically identify and proactively provide assistance. It relies on manual discovery and guidance, resulting in low response efficiency and a tendency to amplify operational deviations. Existing SOP management relies heavily on post-event verification or monitoring of single operational data. It is difficult to combine multi-dimensional data such as operators' facial expressions, vocal emotions, and operational behaviors for real-time analysis. This is not conducive to predicting operational risks in advance and also has an adverse impact on the precise control of operational standards at each position and in each process.

[0004] Therefore, there is an urgent need for a multimodal perception technology that integrates facial expression recognition, voice emotion analysis, and operational fluency analysis. This technology would accurately assess the confusion state of operators during SOP execution, automatically push targeted auxiliary cases, and strengthen real-time control of operational procedures. This would address issues in existing SOP management systems such as incomplete state perception, untimely auxiliary responses, and imprecise control of procedures, thereby improving the standardization and efficiency of SOP execution. Therefore, we propose a visualized SOP management system to solve the aforementioned problems. Summary of the Invention

[0005] 1. Technical problems to be solved While traditional SOP management methods offer several advantages, they lack the ability to capture operators' emotional state and operational fluency in real time. This makes it difficult to accurately identify abnormal states such as confusion or hesitation regarding SOP steps, increasing the risk of operational errors. Furthermore, when operators experience operational pauses or mistakes due to unfamiliarity with or misunderstanding of SOP steps, existing systems lack mechanisms for automatic identification and proactive support. This necessitates manual intervention, resulting in low response efficiency and a tendency to amplify operational deviations. Moreover, SOP control largely relies on post-event verification or monitoring of single operational data, making it difficult to combine multi-dimensional data such as operators' facial expressions, emotional states, and operational behaviors for real-time analysis. This hinders the early prediction of operational risks and negatively impacts the precise control of operational standards across different positions and stages.

[0006] 2. Technical Solution To solve the above problems, the present invention adopts the following technical solution.

[0007] A visual SOP management system, comprising: The facial expression acquisition module is used to acquire facial image data of operators in real time during the execution of SOP steps. The facial image data includes at least eye movement parameters and mouth corner curvature data. The voice data acquisition module is used to synchronously acquire the voice information of the operator during the execution of the SOP. The voice information includes at least the voice frequency and sensitive operation words. The operation behavior acquisition module is used to acquire the operation behavior data of the operator when performing SOP steps. The operation behavior data includes at least the operation action sequence, the execution time of each action, the action interval time, and operation accuracy data. The operation accuracy data is judged based on the deviation threshold between the operation action and the SOP standard action. The emotion and operation state analysis module is bidirectionally connected to the facial expression acquisition module, the voice data acquisition module, and the operation behavior acquisition module, respectively, and is used to receive the facial image data, voice information, and operation behavior data. The emotion and operation state analysis module has a built-in multimodal fusion analysis model, which includes a facial expression feature extraction unit, a voice emotion feature extraction unit, an operation fluency analysis unit, and a deep learning classifier. The facial expression feature extraction unit is used to extract facial features such as eye closure frequency and mouth corner curvature change data from the facial image data. The voice emotion feature extraction unit is used to extract audio standard deviation and sensitive operation words voice features from the voice information. The operation fluency analysis unit is used to calculate the action coherence, standard duration deviation rate and number of operation errors based on the operation behavior data. The operation error is the behavior of the operation action exceeding the SOP standard action deviation threshold. The deep learning classifier is used to normalize the feature data extracted by the facial expression feature extraction unit, the voice emotion feature extraction unit, and the operation fluency analysis unit, and outputs the operator's confusion level score for the current SOP step. The score range is 0-10 points, and the preset confusion level threshold is 6 points. The visualization interaction and case push module is bidirectionally connected to the emotion and operation status analysis module, and is also bidirectionally connected to the SOP standardization storage and retrieval module. The visualization interaction and case push module receives the confusion level score output by the emotion and operation status analysis module. When the confusion level score reaches or exceeds a preset threshold, it sends a case retrieval request to the SOP standardization storage and retrieval module, receives the three most relevant operation case videos with the highest matching degree, sorts them according to standard operation priority, and pushes them to the operator's interactive terminal. The module also marks the key corresponding nodes of the case videos and the current step on the SOP visualization interface. At the same time, the visualization interaction and case push module receives the standardized SOP file output by the SOP standardization storage and retrieval module and visualizes the current SOP step in the form of a dynamic flowchart. The unified standard control module is bidirectionally connected to the operation behavior acquisition module and the SOP standardization storage and retrieval module. The unified standard control module receives the operation behavior data output by the operation behavior acquisition module and the SOP standard operation parameters output by the SOP standardization storage and retrieval module, and performs real-time comparison. When the operation deviation between the two exceeds a preset threshold, a preset form of early warning signal is issued, and the deviation data and early warning timestamp are recorded to form an operation deviation log.

[0008] Furthermore, the multimodal fusion analysis model adopts a CNN-BiLSTM hybrid network structure with attention mechanism enhancement. The facial expression feature extraction unit adopts a lightweight convolutional neural network MobileNetV3, the speech emotion feature extraction unit adopts a bidirectional long short-term memory network BiLSTM, and the operation fluency analysis unit adopts a gradient boosting tree model. The multimodal fusion analysis model is trained based on no less than 10,000 sets of multimodal labeled data.

[0009] Furthermore, the operation behavior acquisition module includes a device sensor interface unit and a motion capture unit. The device sensor interface unit is used to interact with the industrial robot and intelligent operating console in real time to obtain device operation instructions and execution feedback data. The motion capture unit uses a binocular vision camera and extracts the upper limb movement coordinate data of the operator through a skeletal key point detection algorithm.

[0010] Furthermore, the SOP standardization storage and retrieval module has a built-in version management unit, which records the creation time, modification history and approval process of the SOP file, retains at least 10 historical versions, and uses a hash algorithm to uniquely identify each version of the SOP file.

[0011] Furthermore, the visualization interaction and case push module supports multi-terminal adaptation, and the interactive terminals include industrial tablets and operator workstation displays.

[0012] Furthermore, the unified standard control module has a built-in deviation level classification unit, which divides operational deviations into three levels: minor deviation, general deviation, and severe deviation. Different deviation levels correspond to different warning intensities and handling procedures. The unified standard control module is also used to periodically collect deviation data for each position and each SOP step and generate deviation analysis reports.

[0013] Furthermore, it also includes an edge computing processing unit, which is bidirectionally connected to the emotion and operation state analysis module. It receives the multimodal raw data collected by the emotion and operation state analysis module, performs localization processing, and feeds back the feature extraction results to the emotion and operation state analysis module. The edge computing processing unit uses a data desensitization algorithm to perform privacy protection processing on the operator's facial image and voice raw data, retaining only the feature data for analysis.

[0014] Furthermore, the similar operation case videos are stored in a pre-stored form. The matching rules for the similar operation case videos include SOP step number matching, operation scene similarity matching, and operation difficulty coefficient matching. The operation scene similarity is compared by extracting the lighting parameters, device type parameters, and spatial layout parameters of the operation environment. The matching degree is calculated using the cosine similarity algorithm, and the preset matching threshold is manually set.

[0015] 3. Beneficial Effects Compared with the prior art, the advantages of this invention are: (1) This solution integrates three dimensions: facial expression recognition, voice emotion analysis and operation fluency detection. It uses a multimodal fusion analysis model to quantify the degree of confusion of operators. Compared with single data monitoring, the recognition accuracy is higher and can accurately capture abnormal states such as hesitation and lag. (2) In this solution, when the level of confusion reaches the preset threshold, the three most relevant operation case videos with the highest matching degree are automatically retrieved and pushed according to the priority of standard operation. No manual intervention is required, and the response speed is greatly improved compared with traditional manual guidance. At the same time, the case matching combines the SOP step number, operation scenario similarity and difficulty coefficient, and the cosine similarity algorithm ensures the accuracy of matching. At the same time, key corresponding nodes are marked to help operators quickly understand the key points of standard operation, reduce learning costs and improve work efficiency. (3) This scheme unifies and standardizes the real-time comparison of operational behavior data with SOP standard parameters by the control module, triggers graded early warnings according to the deviation level, and combines dynamic early warning signals with operation interception mechanisms to curb the expansion of deviations from the source and promote the unification of operational standards for all positions. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the overall system architecture of the present invention; Figure 2 This is a schematic diagram of the data stream processing and analysis architecture of the present invention; Figure 3 This is a schematic diagram of the data flow in the data acquisition layer of the present invention; Figure 4 This is a schematic diagram of the data flow in the data analysis layer of the present invention; Figure 5 This is a schematic diagram of the data flow for the visualization and case push layer of this invention; Figure 6 This is a schematic diagram of the data flow of the standard control layer of the present invention; Figure 7 This is a schematic diagram of the auxiliary management data flow of the present invention.

[0017] Explanation of the labels in the diagram: 1. Facial expression capture module; 2. Voice data acquisition module; 3. Operation behavior acquisition module; 301. Equipment sensor interface unit; 302. Motion capture unit; 4. Emotion and Operational State Analysis Module; 401. Multimodal Fusion Analysis Model; 4011. Facial Expression Feature Extraction Unit; 4012. Voice Emotion Feature Extraction Unit; 4013. Operational Fluency Analysis Unit; 4014. Deep Learning Classifier; 5. Visual interaction and case study push module; 6. Standard Operating Procedure (SOP) storage and retrieval module; 601. Version management unit; 7. Standardized and unified control module; 701. Deviation level classification unit; 8. Edge computing processing unit. Detailed Implementation

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0019] Example 1: Please see Figures 1-7 A visual SOP management system, comprising: The facial expression acquisition module 1 is used to acquire facial image data of operators in real time during the execution of SOP steps. The facial image data includes at least eye movement parameters and mouth corner curvature data. Voice data acquisition module 2 is used to synchronously acquire the voice information of operators during the execution of SOP. The voice information includes at least voice frequency and sensitive operation words. Operation behavior acquisition module 3 is used to acquire operation behavior data of operators performing SOP steps. The operation behavior data includes at least the sequence of operation actions, the execution time of each action, the action interval time, and operation accuracy data. The operation accuracy data is judged based on the deviation threshold between the operation action and the SOP standard action. The emotion and operation state analysis module 4 is bidirectionally connected to the facial expression acquisition module 1, the voice data acquisition module 2, and the operation behavior acquisition module 3, respectively, and is used to receive facial image data, voice information, and operation behavior data. The emotion and operation state analysis module 4 has a built-in multimodal fusion analysis model 401, which includes a facial expression feature extraction unit 4011, a voice emotion feature extraction unit 4012, an operation fluency analysis unit 4013, and a deep learning classifier 4014. The facial expression feature extraction unit 4011 is used to extract facial features such as eye closing frequency and mouth corner curvature change data from facial image data. The voice emotion feature extraction unit 4012 is used to extract audio standard deviation and sensitive operation words voice features from voice information. The operation fluency analysis unit 4013 is used to calculate the action continuity, standard duration deviation rate and number of operation errors based on operation behavior data. Operation errors are behaviors that exceed the SOP standard action deviation threshold. The deep learning classifier 4014 is used to normalize the feature data extracted by the facial expression feature extraction unit 4011, the voice emotion feature extraction unit 4012 and the operation fluency analysis unit 4013, and output the operator's confusion level score for the current SOP step. The score range is 0-10 points, and the preset confusion level threshold is 6 points. The visualization interaction and case push module 5 is bidirectionally connected to the emotion and operation status analysis module 4, and is also bidirectionally connected to the SOP standardization storage and retrieval module 6. The visualization interaction and case push module 5 receives the confusion level score output by the emotion and operation status analysis module 4. When the confusion level score reaches or exceeds the preset threshold, it sends a case retrieval request to the SOP standardization storage and retrieval module 6, receives the three most relevant operation case videos with the highest matching degree, sorts them according to the standard operation priority, and pushes them to the operator's interactive terminal. It also marks the key corresponding nodes of the case videos and the current step on the SOP visualization interface. At the same time, the visualization interaction and case push module 5 receives the standardized SOP file output by the SOP standardization storage and retrieval module 6 and visualizes the current SOP steps in the form of a dynamic flowchart. The unified standard control module 7 is bidirectionally connected to the operation behavior acquisition module 3 and the SOP standard storage and retrieval module 6. The unified standard control module 7 receives the operation behavior data output by the operation behavior acquisition module 3 and the SOP standard operation parameters output by the SOP standard storage and retrieval module 6, and performs real-time comparison. When the operation deviation between the two exceeds the preset threshold, a preset form of warning signal is issued, and the deviation data and warning timestamp are recorded to form an operation deviation log.

[0020] The multimodal fusion analysis model 401 adopts a CNN-BiLSTM hybrid network structure enhanced by attention mechanism. Among them, the facial expression feature extraction unit 4011 adopts the lightweight convolutional neural network MobileNetV3, the speech emotion feature extraction unit 4012 adopts the bidirectional long short-term memory network BiLSTM, and the operation fluency analysis unit 4013 adopts the gradient boosting tree model. The multimodal fusion analysis model 401 is trained based on no less than 10,000 sets of multimodal labeled data. The operation behavior acquisition module 3 includes a device sensor interface unit 301 and a motion capture unit 302. The device sensor interface unit 301 is used to interact with the industrial robot and the intelligent operating console in real time to obtain the device operation instructions and execution feedback data. The motion capture unit 302 uses a binocular vision camera and extracts the upper limb movement coordinate data of the operator through the skeletal key point detection algorithm. The SOP standardization storage and retrieval module 6 has a built-in version management unit 601, which is used to record the creation time, modification records and approval process of SOP files, retain at least 10 historical versions, and use a hash algorithm to uniquely identify each version of the SOP file. The visualization interaction and case push module 5 supports multi-terminal adaptation, and the interactive terminals include industrial tablets and operator workstation displays. The Unified Standard Control Module 7 has a built-in deviation level classification unit 701, which divides operational deviations into three levels: minor deviation, general deviation, and severe deviation. Different deviation levels correspond to different warning intensities and handling procedures. The Unified Standard Control Module 7 is also used to periodically collect deviation data for each position and each SOP step and generate deviation analysis reports. It also includes an edge computing processing unit 8, which is bidirectionally connected to the emotion and operation state analysis module 4. It receives the multimodal raw data output by the emotion and operation state analysis module 4, performs localization processing, and feeds back the feature extraction results to the emotion and operation state analysis module 4. The edge computing processing unit 8 uses a data desensitization algorithm to perform privacy protection processing on the operator's facial image and voice raw data, retaining only the feature data for analysis. Similar operation case videos are stored in a pre-stored form. The matching rules for similar operation case videos include SOP step number matching, operation scene similarity matching, and operation difficulty coefficient matching. Among them, operation scene similarity is compared by extracting the lighting parameters, equipment type parameters, and spatial layout parameters of the operation environment. The matching degree is calculated using the cosine similarity algorithm, and the preset matching threshold is manually set.

[0021] Example 2: In view of the above embodiment 1, for further description, please refer to Figures 1-7 The motion capture unit 302 of the operation behavior acquisition module 3 uses a binocular vision camera to extract 18 key points of the upper limb bones of the operator (including shoulder, elbow, wrist, finger joints, etc.) through the MediaPipe human posture detection algorithm. The coordinate extraction accuracy is ±0.5mm, and the key point tracking frame rate is synchronized with the camera frame rate. The 401 facial expression feature extraction unit 4011 of the multimodal fusion analysis model adopts the MobileNetV3-Small architecture. The input facial image size is normalized to 224×224 pixels. Features are extracted through 16 convolutional layers. The 8th layer introduces an attention mechanism to enhance the feature weights of the eye and mouth regions. The output feature vector has a dimension of 512. The BiLSTM network of the speech emotion feature extraction unit 4012 contains two hidden layers, each with 256 hidden units. The input speech information is first resampled at a sampling rate of 16kHz and transformed by Mel frequency cepstral coefficients (MFCC) (extracting 13-dimensional MFCC features). Then, it outputs a 256-dimensional speech feature vector through BiLSTM. The audio standard deviation calculation window length is set to 0.5s.

[0022] The gradient boosting tree model of the operation fluency analysis unit 4013 adopts the XGBoost algorithm, with 100 trees, a learning rate of 0.1, and a maximum tree depth of 8. The action coherence is calculated as "actual action interval time / standard action interval time", and the standard duration deviation rate is calculated as "|actual action duration - standard action duration| / standard action duration × 100%". The operation error count period is consistent with the execution time of the current SOP step (not exceeding 10 minutes).

[0023] The 10,000 sets of multimodal labeled data include 3,000 sets of facial expression data, 3,000 sets of voice emotion data, and 4,000 sets of operational behavior data. Each set of data is labeled with a corresponding confusion level score (0-10 points). The labelers consist of 3 SOP domain experts, and the consistency of the scores is verified by the Kappa coefficient (Kappa value ≥ 0.85).

[0024] The deep learning classifier 4014 adopts a fully connected neural network structure. The input layer receives a concatenated feature vector (771 dimensions in total) of facial expression (512 dimensions), voice emotion (256 dimensions), and operation fluency (3 dimensions, including action continuity, standard duration deviation rate, and number of operation errors). After processing by two hidden layers (512 dimensions and 256 dimensions respectively), a 1-dimensional score result is output.

[0025] The normalization process uses the Min-Max normalization algorithm to map each feature data to the [0,1] interval, where facial expression features account for 40% of the weight, voice emotion features account for 30%, and operation fluency features account for 30%. Example of the scoring relationship: 0-2 points (no confusion), 3-5 points (slight confusion), 6-8 points (moderate confusion), 9-10 points (severe confusion).

[0026] The approval process of the version management unit 601 of the SOP standardization storage and retrieval module 6 includes four nodes: "drafting, departmental review, quality department review, and release". The approval time for each node is 24 hours. The system automatically records the approver's account, approval comments, and operation timestamp. The retention period for historical versions is permanent. The hash algorithm used is SHA-256. Each SOP version file generates a unique 64-bit hash value for version integrity verification.

[0027] The standard operating parameters of the SOP are stored in JSON format, which includes fields such as "step number, action name, standard duration (±5s), deviation threshold (action coordinate deviation ≤2mm, duration deviation rate ≤10%), and operating environment requirements". The maximum storage capacity of a single SOP file is 100MB.

[0028] The case matching rules of the Visual Interaction and Case Push Module 5 are weighted as follows: SOP step number matching accounts for 50%, operation scene similarity matching accounts for 30%, and operation difficulty coefficient matching accounts for 20%. The operation scene lighting parameter extraction range is 100-1000 lux, the equipment type parameter is encoded as "equipment model + function type" (e.g., "IRB120-handling"), and the spatial layout parameter adopts coordinate system positioning (X / Y / Z axis accuracy ±1cm). The cosine similarity algorithm calculation process is "scene feature vector dot product / (scene feature vector magnitude × standard scene vector magnitude)", and the preset matching threshold is set to 0.8 (i.e., similarity ≥ 80% is considered a valid match). Key corresponding node annotation method: Use a red flashing border (flashing frequency 2 times / second) to mark the corresponding position of the case video time point and the current SOP step, and at the same time display text description on the right side of the interface (such as "00:35-corresponding step 3: workpiece gripping action, pay attention to wrist joint angle").

[0029] The deviation level thresholds of the unified standard control module 7 are divided into: minor deviation (deviation rate 5%-10%, no safety risk), moderate deviation (deviation rate 10%-20%, may affect product quality), and serious deviation (deviation rate >20%, posing a safety hazard). Corresponding warning levels: Minor deviation (green indicator light stays on), moderate deviation (yellow indicator light flashes + buzzer alert, frequency 1 time / second), severe deviation (red indicator light flashes + buzzer alert, frequency 2 times / second + terminal pop-up blocking operation). Deviation handling process: Minor deviations are automatically recorded and summarized into a report daily; general deviations require operators to submit a rectification explanation within 1 hour; serious deviations trigger an emergency stop order, requiring on-site inspection and resetting by technical personnel; deviation analysis reports include indicators such as "job title, SOP step number, deviation level distribution, monthly deviation trend, and top 5 high-frequency deviations", and the report generation cycle is automatically pushed to the management terminal every Monday.

[0030] The edge computing processing unit 8 uses the NVIDIA Jetson Xavier NX hardware platform and runs on the Ubuntu 20.04 operating system. It supports parallel processing of four 1080P video streams and one audio stream. The data desensitization algorithm uses "facial image blurring (preserving the feature areas of the eyes / corners of the mouth, and pixelating other areas) + voice information voiceprint stripping (preserving frequency and vocabulary features, and removing personal voiceprint features)". The desensitized data complies with the GDPR privacy protection standard. Feature data transmission uses AES-256 encryption and the transmission protocol is MQTT. The feature extraction results are transmitted every 100ms. If the network is interrupted, the local cache capacity can store feature data within 2 hours, and it will be automatically retransmitted after the network is restored.

[0031] Example 3: In view of the above embodiments 1 and 2, for further description, please refer to [link to documentation]. Figures 1-7 A management method for a visual SOP management system includes the following steps: S1, SOP Initialization and Standard Parameter Loading S1-1, SOP File Retrieval and Version Confirmation: The Visual Interaction and Case Push Module 5 sends a request to the SOP retrieval module 6 for the current operation position. The SOP Standardization Storage and Retrieval Module 6 retrieves the latest effective SOP version through the built-in version management unit 601, and verifies the uniqueness and integrity of the SOP file of this version through the SHA-256 hash algorithm to ensure that it has not been tampered with. S1-2, Standard Parameter Distribution: The SOP standardization storage and retrieval module 6 synchronously distributes the SOP standard operation parameters to the operation behavior acquisition module 3, the emotion and operation status analysis module 4, and the unified standard control module 7, providing a benchmark for subsequent data comparison and analysis.

[0032] S2, Multimodal Data Synchronous Acquisition S2-1 Facial Expression Data Acquisition: The facial expression acquisition module 1 starts the image acquisition function to capture facial images of the operator when performing SOP steps in real time. The acquisition frame rate is synchronized with the rhythm of the operation. The module focuses on collecting eye movement parameters (including eye closing frequency, and the statistical period is the duration of the current SOP single step) and mouth corner curvature data (value range 0-1, 0 is completely drooping, 1 is completely upturned). The acquired data is transmitted to the emotion and operation status analysis module 4 in real time. S2-2, Voice Information Acquisition: The voice data acquisition module 2 synchronously starts the recording function, focusing on capturing voice frequencies (the frequency range of the audio waveform is calculated in real time; the normal operation voice frequency range is 100-300Hz, and frequency fluctuations of ±50Hz may occur when confused) and sensitive operation words (the preset vocabulary library contains "how to operate", "wrong", "stuck", etc., which are identified in real time through keyword matching algorithms). The acquired raw voice data is transmitted to the emotion and operation status analysis module 4 and backed up to the edge computing processing unit 8 at the same time. S2-3, Operational Behavior Data Acquisition: Operational behavior acquisition module 3 collects data collaboratively through two sub-units: The device sensor interface unit 301 interacts with the industrial robot and intelligent control panel in real time to obtain device operation instructions (such as "grab the workpiece") and execution feedback data (such as "grab successfully / failed"). Motion capture unit 302: Employs a binocular vision camera and uses the MediaPipe human posture detection algorithm to extract coordinate data of 18 key points of the operator's upper limb skeleton (shoulder, elbow, wrist, finger joints, etc.). It simultaneously records the sequence of operation actions (such as "reaching out → grabbing → moving → placing"), the execution time of each action, the action interval time, and the operation accuracy data (judgment criteria: whether the deviation between the actual action coordinates and the SOP standard action coordinates exceeds the threshold of ≤2mm). The collected operation behavior data is transmitted in real time to the emotion and operation status analysis module 4 and the unified standard control module 7.

[0033] S3. Multimodal Feature Extraction and Perplexity Analysis S3-1. Feature Extraction Based on Multimodal Fusion Analysis Model 401: After receiving multimodal data, the emotion and operational state analysis module 4 performs feature extraction through the built-in multimodal fusion analysis model 401 (which employs a CNN-BiLSTM hybrid network structure enhanced with an attention mechanism). The facial expression feature extraction unit 4011 adopts the MobileNetV3-Small architecture, normalizes the input facial image to 224×224 pixels, extracts features through 16 convolutional layers (the 8th layer introduces an attention mechanism to enhance the feature weights of the eye and mouth regions), and finally outputs a 512-dimensional facial expression feature vector (including the frequency of eye closure and the trend of mouth corner curvature changes). The speech emotion feature extraction unit 4012 uses a BiLSTM network with two hidden layers (256 hidden units per layer). First, the speech data is transformed by Mel-frequency cepstral coefficients (MFCC) (extracting 13-dimensional MFCC features). Then, the BiLSTM network captures the temporal features of speech frequency fluctuations and the frequency of occurrence of sensitive words, and outputs a 256-dimensional speech emotion feature vector (including audio standard deviation, with a calculation window length of 0.5s). Operational fluency analysis unit 4013: It adopts the XGBoost gradient boosting tree model and calculates three core indicators based on operational behavior data: Action continuity = actual action interval time / standard action interval time (the closer the value is to 1, the smoother it is); Standard duration deviation rate = |actual action duration - standard action duration| / standard action duration × 100%; Number of operation errors (the statistical period is consistent with the duration of the current SOP step, and the maximum is no more than 10 minutes. Actions that exceed the deviation threshold are counted as errors), and outputs a 3-dimensional operational fluency feature vector; S3-2, Calculation of Confusion Level Score: The deep learning classifier 4014 receives the above three types of feature vectors, first uses the Min-Max normalization algorithm to map the feature data to the [0,1] interval, and then performs fusion calculation according to the weights of facial expression features (40%), voice emotion features (30%), and operation fluency features (30%), and finally outputs a confusion level score of 0-10 (scoring rules: 0-2 points = no confusion, 3-5 points = slight confusion, 6-8 points = moderate confusion, 9-10 points = severe confusion, and the preset warning threshold is 6 points). The scoring results are transmitted to the visualization interaction and case push module 5 in real time.

[0034] S4, Confusion Response and Intelligent Push of Similar Cases S4-1, Scoring Judgment and Case Retrieval Trigger: The Visual Interaction and Case Push Module 5 monitors the level of confusion in real time. If the score is <6 (no / slight confusion), the Visual Interaction and Case Push Module 5 continuously displays the current SOP step in the form of a dynamic flowchart. The flowchart is generated based on the standard file issued by the SOP Standardization Storage and Retrieval Module 6 and is compatible with two types of terminals: industrial tablets and operator workstation displays. If the score is ≥6 (medium / severe confusion), the Visual Interaction and Case Push Module 5 automatically sends a case retrieval request to the SOP Standardization Storage and Retrieval Module 6. The request parameters include the current SOP step number and operation scenario parameters. S4-2, Case Matching and Sorting: After receiving the request, the SOP standardization storage and retrieval module 6 filters cases according to the preset matching rules: Matching rule weights: SOP step number matching (50%), operation scenario similarity matching (30%), calculated by the cosine similarity algorithm: scene feature vector dot product / (scene feature vector magnitude × standard scene vector magnitude), preset matching threshold 0.8, operation difficulty coefficient matching (20%). The three case videos with the highest matching degree are selected and sorted by "Standard Operating Procedure Compliance Rate" (Compliance Rate = Number of actions in the video that meet the SOP standard / Total number of actions × 100%, with higher compliance rates given priority) and fed back to the Visual Interaction and Case Push Module 5; S4-3, Case Push and Key Node Annotation: The Visual Interaction and Case Push module 5 pushes the three sorted case videos to the operator's interactive terminal, and at the same time annotates key nodes on the SOP dynamic flowchart interface: the corresponding position of the case video time point and the current step is marked with a red flashing border, and text descriptions are displayed on the right side of the interface.

[0035] S5. Real-time monitoring and deviation handling of operating procedures S5-1, Real-time Comparison and Deviation Judgment: The unified standard control module 7 receives two data streams in real time: actual operation behavior data output by the operation behavior acquisition module 3; and SOP standard operation parameters output by the SOP standard storage and recall module 6. The unified standard control module 7 performs real-time comparison through the built-in deviation level classification unit 701, classifying deviations into three levels based on the deviation rate (deviation rate = |actual value - standard value| / standard value × 100%): Minor deviation: deviation rate 5%-10%, no safety risk (e.g., action duration 1 second longer than the standard); General deviation: deviation rate 10%-20%, may affect product quality (e.g., action coordinate deviation 1.5mm, close to the threshold of 2mm); Severe deviation: deviation rate > 20%, posing a safety hazard (e.g., action coordinate deviation 3mm, exceeding the threshold). S5-2, Tiered Early Warning and Operational Intervention: The unified and standardized control module 7 triggers corresponding early warnings based on the deviation level. Minor deviation: The green indicator light at the control station remains constantly lit, with no audible prompt, and deviation data is only recorded in the background; General deviation: The yellow indicator light flashes (frequency 1 time / second) + buzzer prompt (volume 60dB), and a "deviation prompt" pops up on the interactive terminal (containing deviation content, such as "Current action duration deviation 12%, please adjust speed"); Severe deviation: The red indicator light flashes (frequency 2 times / second) + high-decibel buzzer (volume 80dB, lasting 3 seconds), and an "emergency stop command" is sent to the industrial robot / intelligent control console (via the device sensor interface unit 301 of the operation behavior acquisition module 3), intercepting the current operation, requiring on-site confirmation by technical personnel before resetting; S5-3, Deviation Recording and Report Generation: The unified standard control module 7 records all deviations (including minor, general, and serious deviations) and generates operation deviation logs (including deviation timestamp, deviation level, deviation content, operator ID, and job title); at the same time, it regularly compiles deviation data and pushes it to the management terminal to provide data support for SOP optimization.

[0036] S6, Edge Data Processing and Privacy Protection S6-1, Localized processing of raw data: The emotion and operational state analysis module 4 synchronously transmits the raw data collected from the multimodal dataset to the edge computing processing unit 8, and then uses the NVIDIA Jetson Xavier NX hardware platform for localized processing. After extracting key features, the data is fed back to the emotion and operational state analysis module 4, reducing the pressure of cloud transmission. S6-2, Data Desensitization and Secure Transmission: Edge computing processing unit 8 uses privacy-preserving algorithms to desensitize the original data: Facial images: retain the feature areas of the eyes and corners of the mouth (for subsequent analysis), and pixelate other areas to eliminate personal facial recognition information; Voice data: remove personal voiceprint features through a voiceprint stripping algorithm, retaining only features such as voice frequency and vocabulary content for emotion analysis; The desensitized feature data is encrypted using AES-256 and transmitted to the emotion and operation status analysis module 4 via the MQTT protocol; If the network is interrupted, the local cache of edge computing processing unit 8 can store feature data within 2 hours, and automatically retransmit it after the network is restored to ensure that the data is not lost.

[0037] S7 and SOP version maintenance and case library updates S7-1, SOP Version Management: The version management unit 601 of the SOP standardization storage and retrieval module 6 records the entire lifecycle information of the SOP. When the SOP needs to be updated, the version management unit 601 automatically generates a new version and retains the historical version. Each version generates a unique 64-bit identifier through the SHA-256 hash algorithm to ensure traceability. S7-2, Case Library Update: The SOP standardization storage and retrieval module 6 regularly updates the video library of similar operation cases: adding new cases (collecting standard operation videos of excellent operators, which are then reviewed and approved by 3 SOP experts before being added to the library); removing old cases (deleting old cases that do not match the current SOP version, have blurry images, or do not comply with regulations); adding cases to the library is categorized and stored according to "SOP step number + operation scenario + difficulty level" to ensure the synchronization of cases with SOPs.

[0038] The above description is merely a preferred embodiment of the present invention; however, the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and its improved concepts, should be covered within the scope of protection of the present invention.

Claims

1. A visual SOP management system, characterized in that, include: The facial expression acquisition module (1) is used to acquire facial image data of the operator in real time during the execution of SOP steps. The facial image data includes at least eye movement parameters and mouth corner curvature data. The voice data acquisition module (2) is used to synchronously acquire the voice information of the operator during the execution of the SOP. The voice information includes at least the voice frequency and sensitive operation words. The operation behavior acquisition module (3) is used to acquire the operation behavior data of the operator performing the SOP steps. The operation behavior data includes at least the operation action sequence, the execution time of each action, the action interval time and operation accuracy data. The operation accuracy data is judged based on the deviation threshold between the operation action and the SOP standard action. The emotion and operation state analysis module (4) is bidirectionally connected to the facial expression acquisition module (1), the voice data acquisition module (2) and the operation behavior acquisition module (3) respectively, and is used to receive the facial image data, voice information and operation behavior data. The emotion and operation state analysis module (4) has a built-in multimodal fusion analysis model (401). The multimodal fusion analysis model (401) includes a facial expression feature extraction unit (4011), a voice emotion feature extraction unit (4012), an operation fluency analysis unit (4013) and a deep learning classifier (4014). The facial expression feature extraction unit (4011) is used to extract facial features such as eye closing frequency and mouth corner curvature change data from the facial image data. The voice emotion feature extraction unit (4012) is used to extract audio standard deviation and sensitive operation vocabulary voice features from the voice information. The operation fluency analysis unit (4013) is used to calculate the action continuity, standard duration deviation rate and number of operation errors based on the operation behavior data. The operation error is the behavior of the operation action exceeding the SOP standard action deviation threshold. The deep learning classifier (4014) is used to normalize the feature data extracted by the facial expression feature extraction unit (4011), the voice emotion feature extraction unit (4012), and the operation fluency analysis unit (4013), and output the operator's confusion level score for the current SOP step. The score range is 0-10 points, and the preset confusion level threshold is 6 points. The visualization interaction and case push module (5) is bidirectionally connected to the emotion and operation status analysis module (4), and the visualization interaction and case push module (5) is also bidirectionally connected to the SOP standardization storage and retrieval module (6). The visualization interaction and case push module (5) receives the confusion level score output by the emotion and operation status analysis module (4). When the confusion level score reaches or exceeds the preset threshold, it sends a case retrieval request to the SOP standardization storage and retrieval module (6), receives the three most matching operation case videos, sorts them according to the standard operation priority, and pushes them to the operator's interactive terminal. It also marks the key corresponding nodes of the case videos and the current steps on the SOP visualization interface. At the same time, the visualization interaction and case push module (5) receives the standardized SOP file output by the SOP standardization storage and retrieval module (6) and visualizes the current SOP steps in the form of a dynamic flowchart. The unified standard control module (7) is bidirectionally connected to the operation behavior acquisition module (3) and the SOP standard storage and recall module (6). The unified standard control module (7) receives the operation behavior data output by the operation behavior acquisition module (3) and the SOP standard operation parameters output by the SOP standard storage and recall module (6), and performs real-time comparison. When the operation deviation between the two exceeds the preset threshold, a preset form of early warning signal is issued, and the deviation data and early warning timestamp are recorded to form an operation deviation log.

2. The visual SOP management system according to claim 1, characterized in that: The multimodal fusion analysis model (401) adopts a CNN-BiLSTM hybrid network structure with attention mechanism enhancement. The facial expression feature extraction unit (4011) adopts a lightweight convolutional neural network MobileNetV3, the speech emotion feature extraction unit (4012) adopts a bidirectional long short-term memory network BiLSTM, and the operation fluency analysis unit (4013) adopts a gradient boosting tree model. The multimodal fusion analysis model (401) is trained based on no less than 10,000 sets of multimodal labeled data.

3. The visual SOP management system according to claim 1, characterized in that: The operation behavior acquisition module (3) includes a device sensor interface unit (301) and a motion capture unit (302). The device sensor interface unit (301) is used to interact with the industrial robot and the intelligent operating console in real time to obtain the device operation instructions and execution feedback data. The motion capture unit (302) uses a binocular vision camera and extracts the upper limb movement coordinate data of the operator through a skeletal key point detection algorithm.

4. The visual SOP management system according to claim 1, characterized in that: The SOP standardization storage and retrieval module (6) has a built-in version management unit (601) for recording the creation time, modification records and approval process of SOP files, retaining at least 10 historical versions, and using a hash algorithm to uniquely identify each version of the SOP file.

5. A visual SOP management system according to claim 1, characterized in that: The visualization interaction and case push module (5) supports multi-terminal adaptation, and the interactive terminals include industrial tablets and operating station displays.

6. A visual SOP management system according to claim 1, characterized in that: The unified standard control module (7) has a built-in deviation level classification unit (701) that divides operational deviations into three levels: minor deviation, general deviation and serious deviation. Different deviation levels correspond to different warning intensities and handling procedures. The unified standard control module (7) is also used to periodically collect deviation data for each position and each SOP step and generate deviation analysis reports.

7. A visual SOP management system according to claim 1, characterized in that: It also includes an edge computing processing unit (8), which is bidirectionally connected to the emotion and operation state analysis module (4), receives the multimodal raw data output by the emotion and operation state analysis module (4), performs localization processing, and feeds back the feature extraction results to the emotion and operation state analysis module (4). The edge computing processing unit (8) uses a data desensitization algorithm to perform privacy protection processing on the operator's facial image and voice raw data, and only retains the feature data for analysis.

8. A visual SOP management system according to claim 1, characterized in that: The videos of similar operation cases are stored in a pre-stored format. The matching rules for the videos of similar operation cases include SOP step number matching, operation scene similarity matching, and operation difficulty coefficient matching. The operation scene similarity is compared by extracting the lighting parameters, equipment type parameters, and spatial layout parameters of the operation environment. The matching degree is calculated using the cosine similarity algorithm, and the preset matching threshold is manually set.