Operator behavior detection method, apparatus and device, and storage medium

By using deep learning algorithms and natural language processing models to identify operator behavior, an operator behavior detection device is developed. This device generates action rules that are then applied to a standardized operation detection device in the silk processing workshop. This process improves the standardized operation detection device, generates action rule test questions, provides suggestions for operational improvement, corrects operator behavior, increases the operator's action qualification rate, and reduces quality and safety accidents.

CN121096019APending Publication Date: 2025-12-09CHINA TOBACCO GUANGXI IND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511185943.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

The lack of standardization in the behavior of operators in the silk-making workshop leads to a high probability of operational errors, and the existing training and assessment measures have limited effectiveness.

Method used

By using deep learning algorithms and large natural language models, key points of operator behavior are identified, action rule test questions are generated, and suggestions for operation improvement are provided to achieve standardized correction of operational behavior.

Benefits of technology

To improve the pass rate of operators' actions, reduce the probability of quality and safety accidents, and ensure consistency of operating procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121096019A_ABST
    Figure CN121096019A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing, and discloses an operator behavior detection method and device, equipment and a storage medium. The method comprises the following steps: receiving to-be-detected video data, preprocessing the to-be-detected video data, and dividing the to-be-detected video data into training data and test data; training a deep learning algorithm by using the training data to obtain a behavior detection model, and detecting the test data and the standard guidance video data by using the behavior detection model to obtain an included angle number and an included angle degree of a plurality of key points of each actual execution action and the standard execution action; training a natural language large model by using the standard guidance video data, and comparing the number and degree of included angles between the actual execution actions and the plurality of key points of the standard execution actions by using the trained natural language large model to obtain a comparison result; and pushing the standard guidance video data and the action rule test questions to an operator terminal. The action qualification rate of operators is increased, and the occurrence probability of quality and safety accidents is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, device, and storage medium for detecting operator behavior. Background Technology

[0002] The behavior of operators in the tobacco processing workshop directly affects the purity and pass rate of the tobacco. To ensure that operators operate the equipment in a standardized manner, technical training is provided to the operators, and operation manuals are developed. However, due to the subjectivity of human beings, people with different experience levels may have different understandings of the training content and operation manuals, leading to differences in the operators' work behaviors.

[0003] Currently, the common solution adopted by cigarette factories across the country is to increase the frequency of training and formulate strict assessment measures to require operators to reduce operational errors. However, there are still problems such as insufficient standardization of operators' work behavior and a high probability of operational errors. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to overcome the shortcomings of the prior art and provide a method, apparatus, device and storage medium for detecting operator behavior.

[0005] This invention provides the following technical solution: In a first aspect, the present invention provides a method for detecting operator behavior, the method comprising: The system receives video data to be detected sent from the operator's terminal, preprocesses the video data to be detected to obtain preprocessed video data, and divides the preprocessed video data into training data and test data. The deep learning algorithm is trained using the training data to obtain a behavior detection model, and the behavior detection model is used to detect the test data to obtain the number of angles and the angle of multiple key points for each actual action in the test data. Obtain standard instructional video data, use the behavior detection model to detect the standard instructional video data, and obtain the number of angles and the angle of multiple key points for each standard execution action in the standard instructional video data; Obtain the work instruction manual, standard execution action video, and standard execution action requirements. Use the work instruction manual, standard execution action video, and standard execution action requirements as a knowledge base. Use the knowledge base to train a large natural language model to obtain a trained large natural language model. Use the trained large natural language model to compare the number and angle of multiple key points of each actual execution action with the number and angle of multiple key points of each standard execution action to obtain the comparison results. The standard guidance video data is pushed to the operator's terminal. The trained natural language big data model generates action rule test questions based on the comparison results, and the action rule test questions are pushed to the operator's terminal.

[0006] In an optional implementation, the preprocessing of the video data to be detected to obtain preprocessed video data includes: The video data to be detected is labeled to obtain labeled video data. The labeled video data is then subjected to a first preprocessing operation to obtain first preprocessed video data. The first preprocessing operation includes at least one of normalization, cropping, and rotation. The first preprocessed video data is subjected to a second preprocessing operation to obtain the preprocessed video data. The second preprocessing operation includes at least one of flipping and scaling.

[0007] In an optional implementation, the step of training a deep learning algorithm using the training data to obtain a behavior detection model, and then using the behavior detection model to detect the test data to obtain the number and angle of multiple key points for each actual action in the test data, includes: The deep learning algorithm is trained using the training data, and the loss function of the deep learning algorithm is optimized to obtain the behavior detection model. The deep learning algorithm is a human pose recognition algorithm. The behavior detection model is used to detect the test data. The key points of the human body in each actual action in the test data are connected with lines with endpoints. The angle between three connected key points in a single actual action or the number and angle between three connected key points of multiple non-adjacent actions are recorded.

[0008] In an optional implementation, the natural language large model is trained using the knowledge base to obtain a trained natural language large model. The trained natural language large model is then used to compare the number and angles of multiple key points of each actual executed action with the number and angles of multiple key points of each standard executed action to obtain a comparison result, including: The natural language processing model is trained using the knowledge base, and the parameters of the natural language processing model are adjusted using the backpropagation algorithm and the optimization algorithm. The loss function of the natural language processing model is also optimized to obtain the trained natural language processing model. The trained natural language large model is used to determine whether the angle between the three key points of each actual action is within a preset range. If it is within the preset range, a count is performed according to the method of exceeding the preset range after reaching it. The start and end times of the actual action of each of the three key points are recorded to obtain the number of actual actions and the actual execution time interval of each of the three key points. The trained natural language large model is used to determine whether the angle between the three key points of each standard execution action is within the preset degree range. If it is within the preset degree range, a count is performed according to the method of exceeding the preset degree range after reaching the preset degree range. The start time and end time of the standard execution action of each point in the three key points are recorded to obtain the number of standard execution actions and the standard execution time interval of each point in the three key points. Based on the actual number of actions performed, the actual execution time interval, the standard number of actions performed, and the standard execution time interval, the comparison results and suggestions for improving the operation are generated.

[0009] In an optional implementation, generating the comparison result based on the actual number of actions performed, the actual execution time interval, the standard number of actions performed, and the standard execution time interval includes: Determine whether the error between the actual number of actions performed and the standard number of actions performed is within a preset error range, and determine whether the error between the actual execution time interval and the standard execution time interval is within a preset time error range; When the error between the actual number of actions performed and the standard number of actions performed is within the preset error range, and the error between the actual execution time interval and the standard execution time interval is within the preset time error range, the comparison result is determined to meet the requirements; otherwise, the comparison result is determined to not meet the requirements.

[0010] In an optional implementation, the step of pushing the standard guidance video data to the operator's terminal based on the comparison result, generating action rule test questions using the trained natural language processing model based on the comparison result, and pushing the action rule test questions to the operator's terminal includes: When the comparison result does not meet the requirements, the standard guidance video data and the operation action improvement suggestions will be pushed to the operator's terminal; The system receives a confirmation learning signal from the operator's terminal, inputs the operation action improvement suggestions into the trained natural language processing model, uses the trained natural language processing model to generate the action rule test questions, and pushes the action rule test questions to the operator's terminal.

[0011] In an optional implementation, after pushing the action rule test questions to the operator's terminal, the method further includes: The system receives the test answer results returned by the operator and uses the trained natural language model to determine whether the test answer results are correct. If the answer to the question is incorrect, the action rule test question will be pushed to the operator's terminal again. If the test question answer is correct, the correction video data from the operator's end is re-collected, and the behavior detection model is used to detect the correction video data to obtain the number and angle of multiple key points of each correction action in the correction video data. The trained natural language big data model is used to compare the number and angle of multiple key points of each correction action with the number and angle of multiple key points of each standard action to obtain the correction result. This process continues until the correction result meets the requirements; otherwise, the standard guidance video data is pushed back to the operator's end.

[0012] In a second aspect, the present invention provides an operator behavior detection device, the device comprising: The preprocessing module is used to receive the video data to be detected sent by the operator and preprocess the video data to be detected to obtain preprocessed video data, and divide the preprocessed video data into training data and test data. The first detection module is used to train a deep learning algorithm using the training data to obtain a behavior detection model, and to use the behavior detection model to detect the test data to obtain the number of angles and the angle of multiple key points for each actual action in the test data. The second detection module is used to acquire standard guidance video data, and use the behavior detection model to detect the standard guidance video data to obtain the number of included angles and the number of included angles of multiple key points for each standard execution action in the standard guidance video data. The comparison module is used to acquire the work instructions, standard execution action videos, and standard execution action requirements. The work instructions, standard execution action videos, and standard execution action requirements are used as a knowledge base. The natural language processing model is trained using the knowledge base to obtain the trained natural language processing model. The trained natural language processing model is then used to compare the number and angle of multiple key points of each actual execution action with the number and angle of multiple key points of each standard execution action to obtain the comparison result. The push module is used to push the standard guidance video data to the operator's terminal, generate action rule test questions based on the comparison results using the trained natural language big data model, and push the action rule test questions to the operator's terminal.

[0013] Thirdly, this disclosure provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the operator behavior detection method described in the first aspect.

[0014] Fourthly, this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the operator behavior detection method described in the first aspect.

[0015] The beneficial effects of this application are: The operator behavior detection method provided in this application uses computer vision to recognize the work behavior of on-site operators, identifies key behaviors in the operation process, and provides improvement suggestions based on standards. These suggestions are then fed back to the operators to correct their operational behavior, improve the pass rate of operators' actions, and reduce the probability of quality and safety accidents.

[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the various drawings, similar components are numbered similarly.

[0018] Figure 1A flowchart of an operator behavior detection method provided in an embodiment of this application is shown; Figure 2 This illustration shows a schematic diagram of a key point on the human body provided in an embodiment of this application; Figure 3 This illustration shows a schematic diagram of three key points provided in an embodiment of this application; Figure 4 This illustration shows a structural schematic diagram of an operator behavior detection device provided in an embodiment of this application; Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of this application is shown. Detailed Implementation

[0019] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0020] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein in the template description is for the purpose of describing particular embodiments only and is not intended to limit the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0022] Example 1 like Figure 1 The diagram shown is a flowchart of an operator behavior detection method according to an embodiment of this application. The operator behavior detection method provided in this embodiment includes the following steps: Step S110: Receive the video data to be detected sent by the operator's terminal, preprocess the video data to be detected to obtain preprocessed video data, and divide the preprocessed video data into training data and test data.

[0023] In this embodiment, the operator's end is equipped with a video acquisition device aimed at the operator being detected. The video acquisition device can be installed at a fixed point or on a track, such as a fixed camera or a tracking robot. The track-mounted installation method should have a tracking function to ensure that key operation points can be captured. The fixed-point installation method needs to stitch the videos captured by different devices into a complete video stream.

[0024] The video acquisition equipment treats each batch of cigarette production operations as a complete video data set to be inspected and transmits it to the server via Ethernet or other transmission methods. It should be noted that to ensure consistency in data and the field of view of the acquired images, the video acquisition equipment should ideally be of the same model. If inconsistencies occur, standardized data through transcoding, cropping, or other corrective procedures should be used.

[0025] Then, after receiving the video data to be detected sent by the video acquisition device, the server annotates the video data to be detected (for example, by combining manual annotation with semi-automatic tools) to obtain annotated video data. The server then performs a first preprocessing operation on the annotated video data to obtain preprocessed video data, which is adapted to the input requirements of the subsequent model. The first preprocessing operation includes at least one of normalization, cropping, and rotation. The server then performs a second preprocessing operation on the preprocessed video data to increase the robustness of the subsequent model, resulting in preprocessed video data. The second preprocessing operation includes at least one of flipping and scaling.

[0026] The above steps adapt the model input requirements through preprocessing such as normalization, cropping, and rotation, and combine data augmentation techniques such as flipping and scaling to simulate the complex environment of the workshop (such as changes in lighting and differences in operating angles), so that the trained model can better adapt to the actual production scenario and reduce recognition deviations caused by environmental interference.

[0027] Step S120: Train the deep learning algorithm using the training data to obtain a behavior detection model, and use the behavior detection model to detect the test data to obtain the number of angles and the angle of multiple key points for each actual action in the test data.

[0028] Specifically, the deep learning algorithm is trained using training data, and its loss function is optimized to obtain the behavior detection model. In this embodiment, the deep learning algorithm is a human pose recognition algorithm, such as OpenPose or YOLOv8-Pose, and the training methods used include stochastic gradient descent and the Adam optimizer.

[0029] Furthermore, such as Figure 2As shown, a behavior detection model is used to detect test data. Key human body points (e.g., 17 points, including core joints such as head, neck, shoulders, elbows, wrists, hips, knees, and ankles, fully covering the whole-body posture during the operation) in each actual executed action are connected with lines containing endpoints. The number and angle of the included angles between single connected three key points or multiple non-adjacent connected three key points in each actual executed action are recorded. For example, as... Figure 3 As shown, the three key points represent the wrist, elbow, and shoulder, respectively, and the two lines connecting these three key points represent the upper arm and forearm, respectively. Simultaneously, a time label 't' is assigned to this angle, forming a data stream of angle-time pairs.

[0030] The above steps, based on human posture recognition algorithms, accurately detect 17 key points on the human body and quantify the movement by calculating the angles between these key points. This approach is more objective and traceable than manual observation, avoiding subjective judgment errors. By recording the number and angles of single or multiple connected three-point intervals and combining them with time tags to form an "angle-time" data stream, the dynamic characteristics of the movement (such as movement frequency and duration) are fully characterized. This provides refined data support for subsequent comparisons with standard movements, avoiding misjudgments caused by the omission of local movements.

[0031] Step S130: Obtain standard instructional video data, and use the behavior detection model to detect the standard instructional video data to obtain the number of angles and the angle of multiple key points of each standard execution action in the standard instructional video data.

[0032] Understandably, obtaining the work instructions, standard execution action videos, and standard execution action requirements (such as text requirements and frequency requirements) for each position in the cigarette manufacturing workshop can be used in practical applications. After training the operators in the relevant positions based on the work instructions, standard execution action videos, and standard execution action requirements, the operators in the relevant positions can perform standard actions for each operation according to the execution requirements, thereby generating standard guidance video data.

[0033] Then, the aforementioned behavior detection model is used to detect the standard instruction video data, and the number of angles and the angle of the three key points of each standard execution action in the standard instruction video data are obtained.

[0034] Preferably, the acceptable range can be determined by statistically analyzing the angle of the standard execution action multiple times (e.g., 100 times) and taking a certain percentage (e.g., 95%) of the confidence interval, thus avoiding overly strict or lenient standards due to individual differences.

[0035] The above steps are based on the work instructions and the standard actions of experienced operators to create demonstration videos. The same behavior detection model is used to extract the key point angle features of the standard actions to ensure the authority and practicality of the "standard" and to provide a unified and reliable reference for subsequent action comparisons.

[0036] Step S140: Obtain the work instruction manual, standard execution action video, and standard execution action requirements. Use the work instruction manual, standard execution action video, and standard execution action requirements as a knowledge base. Use the knowledge base to train the natural language large model to obtain the trained natural language large model. Use the trained natural language large model to compare the number and angle of multiple key points of each actual execution action with the number and angle of multiple key points of each standard execution action to obtain the comparison result.

[0037] Understandably, the work instructions, standard execution action videos, and standard execution action requirements are used as a knowledge base to train large-scale natural language processing models such as Llama, ChatGLM, qwen, and deepseek. Backpropagation and optimization algorithms (such as Adam and SGD) are used to adjust the parameters of the large-scale natural language processing model, and the model performance is optimized by minimizing the loss function (such as the cross-entropy loss function), thus obtaining the trained large-scale natural language processing model. Preferably, hyperparameters such as learning rate, batch size, and training epochs can be adjusted during training to improve the model's training effect. The main function of this large-scale natural language processing model is to detect the difference between the standard execution actions and the actual execution actions of the operator, and to provide reasonable suggestions and action guidance based on the differences.

[0038] Specifically, the trained natural language large model is used to determine whether the angle between the three key points of each actual action is within a preset range. If it is within the preset range, a count is performed according to the method of exceeding the preset range after reaching it. The start and end times of the actual action at each of the three key points are recorded to obtain the actual execution time interval at each of the three key points and the number of actual actions within each actual execution time interval. For example, the actual execution time interval is determined based on the start and end times of each actual action at each point. Then, it is determined whether the angle between each point is within the preset range. If it is within the preset range, a count is performed according to the method of exceeding the preset range after reaching it, thus obtaining the total number of actual actions.

[0039] Similarly, the trained natural language large model is used to determine whether the angle between the three key points of each standard execution action is within the preset degree range. If it is within the preset degree range, a count is performed according to the method of exceeding the preset degree range after reaching the preset degree range. The start time and end time of the standard execution action of each point in the three key points are recorded to obtain the standard execution time interval of each point in the three key points and the number of standard execution actions within each standard execution time interval.

[0040] Furthermore, after obtaining the actual number of actions performed, the actual execution time interval, the standard number of actions performed, and the standard execution time interval, it is determined whether the error between the actual number of actions performed and the standard number of actions performed is within the preset error range, and whether the error between the actual execution time interval and the standard execution time interval is within the preset time error range.

[0041] When the error between the actual number of actions performed and the standard number of actions performed is within the preset error range, and the error between the actual execution time interval and the standard execution time interval is within the preset time error range, the comparison result is determined to meet the requirements; otherwise, the comparison result is determined to not meet the requirements, and corresponding suggestions for improving the operation are generated.

[0042] The above steps are based on large models such as Llama and ChatGLM combined with RAG technology. They link knowledge bases such as work instructions, standard execution action videos and standard execution action requirements with real-time action data to achieve automated judgment of the number of actions and time intervals. This is more efficient than manual assessment and reduces judgment bias caused by differences in experience.

[0043] Step S150: Based on the comparison results, push the standard guidance video data to the operator's terminal, use the trained natural language big data model to generate action rule test questions based on the comparison results, and push the action rule test questions to the operator's terminal.

[0044] Specifically, when the comparison result meets the requirements, it is recorded in the operator's personal performance. When the comparison result does not meet the requirements, standard guidance video data and operational action improvement suggestions are pushed to the operator's terminal. After learning, the operator sends a confirmation learning signal to the server. After receiving the confirmation learning signal from the operator's terminal, the server inputs the operational action improvement suggestions into the trained natural language processing model. Using the trained natural language processing model, it generates action rule test questions, such as multiple-choice questions and open-ended questions, outlining the relevant action execution requirements and how to execute them. These action rule test questions are then pushed to the operator's terminal.

[0045] In one optional implementation, after the operator answers the action rule test questions, the answer is sent to the server. Upon receiving the answer from the operator, the server uses a trained natural language processing model to determine if the answer is correct. If the answer is incorrect, the action rule test questions are re-pushed to the operator; if the answer is correct, the operator is required to re-execute the learned correction action, and the correction video data from the operator's end is re-collected. The correction video data is then re-evaluated to ensure it meets requirements. Specifically, a behavior detection model is used to re-detect the correction video data, obtaining the number and angle of multiple key points for each correction action. The trained natural language processing model then compares the number and angle of multiple key points for each correction action with those for each standard action to obtain the correction result.

[0046] If the correction result of a certain point meets the requirements, the correction result of the next point is judged, until the correction results of all points after re-execution are judged to meet the requirements. Otherwise, the standard guidance video data is pushed to the operator's terminal again, and the operator is required to perform the action in front of the corresponding camera again, so as to correct the operator's wrong actions and failure to perform according to the standard requirements.

[0047] Preferably, this embodiment can also set a maximum number of corrections (e.g., 3 times). If the standard is still not met after exceeding the maximum number of corrections, manual intervention will be automatically triggered (e.g., notifying the team leader to provide on-site guidance) to avoid production stoppage caused by mechanical correction.

[0048] The above steps, through the delivery of standard videos, improvement suggestions, and targeted tests (theoretical Q&A + practical reproduction), achieve a closed loop of "identifying deviations - learning standards - verifying mastery," ensuring that operators not only understand the specifications but also can actually implement them, fundamentally correcting erroneous actions. By continuously correcting non-standard operations, the number of defective tobacco products caused by operational errors is directly reduced, while safety accidents caused by violations are also reduced, ultimately improving the tobacco qualification rate and production safety. Moreover, compared to the "experience-based" guidance of traditional manual training, this step, through model-based standards and automated correction, ensures that all operators have consistent correction criteria, facilitating standardized workshop management and reducing the labor costs of repetitive training.

[0049] The operator behavior detection method provided in this application uses computer vision to recognize the work behavior of on-site operators, identifies key behaviors in the operation process, and provides improvement suggestions based on standards. These suggestions are then fed back to the operators to correct their operational behavior, improve the pass rate of operators' actions, and reduce the probability of quality and safety accidents.

[0050] Example 2 like Figure 4 The diagram shown is a structural schematic of an operator behavior detection device 400 according to an embodiment of this application. The device includes: The preprocessing module 410 is used to receive the video data to be detected sent by the operator and preprocess the video data to be detected to obtain preprocessed video data, and divide the preprocessed video data into training data and test data. The first detection module 420 is used to train a deep learning algorithm using the training data to obtain a behavior detection model, and to use the behavior detection model to detect the test data to obtain the number of angles and the angle of multiple key points of each actual action in the test data. The second detection module 430 is used to acquire standard guidance video data, and use the behavior detection model to detect the standard guidance video data to obtain the number of angles and the angle of multiple key points of each standard execution action in the standard guidance video data. The comparison module 440 is used to acquire the work instruction, standard execution action video, and standard execution action requirements. The work instruction, standard execution action video, and standard execution action requirements are used as a knowledge base. The natural language large model is trained using the knowledge base to obtain the trained natural language large model. The trained natural language large model is then used to compare the number and angle of multiple key points of each actual execution action with the number and angle of multiple key points of each standard execution action to obtain the comparison result. The push module 450 is used to push the standard guidance video data to the operator's terminal, generate action rule test questions based on the comparison results using the trained natural language large model, and push the action rule test questions to the operator's terminal.

[0051] The operator behavior detection device provided in this application embodiment can realize each process of the operator behavior detection method corresponding to Embodiment 1, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0052] The operator behavior detection device provided in this application embodiment uses computer vision to recognize the work behavior of on-site operators, identifies key behaviors in the operation process, and provides improvement suggestions based on standards. These suggestions are then fed back to the operators to correct their operational behavior, improve the pass rate of operators' actions, and reduce the probability of quality and safety accidents.

[0053] Example 3 This application also provides a computer device. Please refer to the following for details. Figure 5 , Figure 5 This is a basic structural block diagram of the computer device in this embodiment.

[0054] The computer device 5 includes a memory 51, a processor 52, and a network interface 53 that are interconnected via a system bus. It should be noted that only a computer device 5 with a memory 51, a processor 52, and a network interface 53 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0055] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0056] The memory 51 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or D slot compatibility test memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 51 may be an internal storage unit of the computer device 5, such as the hard disk or memory of the computer device 5. In other embodiments, the memory 51 may also be an external storage device of the computer device 5, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 5. Of course, the memory 51 may also include both the internal storage unit and its external storage device of the computer device 5. In this embodiment, the memory 51 is typically used to store the operating system and various application software installed on the computer device 5, such as computer-readable instructions for slot compatibility testing methods. In addition, the memory 51 can also be used to temporarily store various types of data that have been output or will be output.

[0057] In some embodiments, the processor 52 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other operator behavior detection chip. The processor 52 is typically used to control the overall operation of the computer device 5. In this embodiment, the processor 52 is used to execute computer-readable instructions stored in the memory 51 or to process data, such as executing computer-readable instructions for the slot compatibility testing method.

[0058] The network interface 53 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 5 and other electronic devices.

[0059] The computer device provided in this embodiment can execute the above-described operator behavior detection method. The operator behavior detection method here can be any of the operator behavior detection methods described in the various embodiments above.

[0060] Example 4 This embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the operator behavior detection method in this embodiment.

[0061] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the computer-readable storage medium can be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device. Of course, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is typically used to store the operating system and various application software installed on the computer device. In addition, the computer-readable storage medium can also be used to temporarily store various types of data that have been output or will be output.

[0062] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, as an alternative implementation, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0063] In addition, the functional modules or units in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0064] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium can be a non-volatile storage medium or a volatile storage medium. For example, the storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other medium capable of storing program code.

[0065] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for detecting operator behavior, characterized in that, The method includes: The system receives video data to be detected sent from the operator's terminal, preprocesses the video data to be detected to obtain preprocessed video data, and divides the preprocessed video data into training data and test data. The deep learning algorithm is trained using the training data to obtain a behavior detection model, and the behavior detection model is used to detect the test data to obtain the number of angles and the angle of multiple key points for each actual action in the test data. Obtain standard instructional video data, use the behavior detection model to detect the standard instructional video data, and obtain the number of angles and the angle of multiple key points for each standard execution action in the standard instructional video data; Obtain the work instruction manual, standard execution action video, and standard execution action requirements. Use the work instruction manual, standard execution action video, and standard execution action requirements as a knowledge base. Use the knowledge base to train a large natural language model to obtain a trained large natural language model. Use the trained large natural language model to compare the number and angle of multiple key points of each actual execution action with the number and angle of multiple key points of each standard execution action to obtain the comparison results. The standard guidance video data is pushed to the operator's terminal. The trained natural language big data model generates action rule test questions based on the comparison results, and the action rule test questions are pushed to the operator's terminal.

2. The operator behavior detection method according to claim 1, characterized in that, The preprocessing of the video data to be detected to obtain preprocessed video data includes: The video data to be detected is labeled to obtain labeled video data. The labeled video data is then subjected to a first preprocessing operation to obtain first preprocessed video data. The first preprocessing operation includes at least one of normalization, cropping, and rotation. The first preprocessed video data is subjected to a second preprocessing operation to obtain the preprocessed video data. The second preprocessing operation includes at least one of flipping and scaling.

3. The operator behavior detection method according to claim 2, characterized in that, The process involves training a deep learning algorithm using the training data to obtain a behavior detection model, and then using the behavior detection model to detect the test data to obtain the number and angle of multiple key points for each actual action in the test data, including: The deep learning algorithm is trained using the training data, and the loss function of the deep learning algorithm is optimized to obtain the behavior detection model. The deep learning algorithm is a human pose recognition algorithm. The behavior detection model is used to detect the test data. The key points of the human body in each actual action in the test data are connected with lines with endpoints. The angle between three connected key points in a single actual action or the number and angle between three connected key points of multiple non-adjacent actions are recorded.

4. The operator behavior detection method according to claim 3, characterized in that, The method involves training a large-scale natural language model using the knowledge base to obtain a trained large-scale natural language model. Then, the trained large-scale natural language model is used to compare the number and angles of multiple key points for each actual executed action with the number and angles of multiple key points for each standard executed action, yielding comparison results, including: The natural language processing model is trained using the knowledge base, and the parameters of the natural language processing model are adjusted using the backpropagation algorithm and the optimization algorithm. The loss function of the natural language processing model is also optimized to obtain the trained natural language processing model. The trained natural language large model is used to determine whether the angle between the three key points of each actual action is within a preset range. If it is within the preset range, a count is performed according to the method of exceeding the preset range after reaching it. The start and end times of the actual action of each of the three key points are recorded to obtain the number of actual actions and the actual execution time interval of each of the three key points. The trained natural language large model is used to determine whether the angle between the three key points of each standard execution action is within the preset degree range. If it is within the preset degree range, a count is performed according to the method of exceeding the preset degree range after reaching the preset degree range. The start time and end time of the standard execution action of each point in the three key points are recorded to obtain the number of standard execution actions and the standard execution time interval of each point in the three key points. Based on the actual number of actions performed, the actual execution time interval, the standard number of actions performed, and the standard execution time interval, the comparison results and suggestions for improving the operation are generated.

5. The operator behavior detection method according to claim 4, characterized in that, The step of generating the comparison result based on the actual number of actions performed, the actual execution time interval, the standard number of actions performed, and the standard execution time interval includes: Determine whether the error between the actual number of actions performed and the standard number of actions performed is within a preset error range, and determine whether the error between the actual execution time interval and the standard execution time interval is within a preset time error range; When the error between the actual number of actions performed and the standard number of actions performed is within the preset error range, and the error between the actual execution time interval and the standard execution time interval is within the preset time error range, the comparison result is determined to meet the requirements; otherwise, the comparison result is determined to not meet the requirements.

6. The operator behavior detection method according to claim 5, characterized in that, The step of pushing the standard guidance video data to the operator's terminal based on the comparison results, generating action rule test questions using the trained natural language processing model based on the comparison results, and pushing the action rule test questions to the operator's terminal includes: When the comparison result does not meet the requirements, the standard guidance video data and the operation action improvement suggestions will be pushed to the operator's terminal; The system receives a confirmation learning signal from the operator's terminal, inputs the operation action improvement suggestions into the trained natural language processing model, uses the trained natural language processing model to generate the action rule test questions, and pushes the action rule test questions to the operator's terminal.

7. The operator behavior detection method according to claim 6, characterized in that, After pushing the action rule test questions to the operator's terminal, the method further includes: The system receives the test answer results returned by the operator and uses the trained natural language model to determine whether the test answer results are correct. If the answer to the question is incorrect, the action rule test question will be pushed to the operator's terminal again. If the test question answer is correct, the correction video data from the operator's end is re-collected, and the behavior detection model is used to detect the correction video data to obtain the number and angle of multiple key points of each correction action in the correction video data. The trained natural language big data model is used to compare the number and angle of multiple key points of each correction action with the number and angle of multiple key points of each standard action to obtain the correction result. This process continues until the correction result meets the requirements; otherwise, the standard guidance video data is pushed back to the operator's end.

8. An operator behavior detection device, characterized in that, The device includes: The preprocessing module is used to receive the video data to be detected sent by the operator and preprocess the video data to be detected to obtain preprocessed video data, and divide the preprocessed video data into training data and test data. The first detection module is used to train a deep learning algorithm using the training data to obtain a behavior detection model, and to use the behavior detection model to detect the test data to obtain the number of angles and the angle of multiple key points for each actual action in the test data. The second detection module is used to acquire standard guidance video data, and use the behavior detection model to detect the standard guidance video data to obtain the number of included angles and the number of included angles of multiple key points for each standard execution action in the standard guidance video data. The comparison module is used to acquire the work instructions, standard execution action videos, and standard execution action requirements. The work instructions, standard execution action videos, and standard execution action requirements are used as a knowledge base. The natural language processing model is trained using the knowledge base to obtain the trained natural language processing model. The trained natural language processing model is then used to compare the number and angle of multiple key points of each actual execution action with the number and angle of multiple key points of each standard execution action to obtain the comparison result. The push module is used to push the standard guidance video data to the operator's terminal, generate action rule test questions based on the comparison results using the trained natural language large model, and push the action rule test questions to the operator's terminal.

9. A computer device, characterized in that, The device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the operator behavior detection method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the operator behavior detection method according to any one of claims 1-7.