User behavior identification method and related device

By analyzing and identifying the actions and identities of staff through surveillance images, and generating warning messages, this solves the problem of ineffective enforcement of existing mobile phone confiscation methods, and achieves safety standards and protection against technology leaks within the workshop.

CN121600588APending Publication Date: 2026-03-03BEIJING CHEHEJIA AUTOMOBILE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411179443.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

In existing technologies, the method of preventing technology leaks by confiscating mobile phones is poorly implemented and cannot reliably regulate the behavior of workers in the workshop, resulting in a high risk of information leakage.

Method used

By analyzing the surveillance images of the monitored area, the system identifies the actions and behaviors of the detected objects and performs identity recognition, generates warning information and saves the image information of violations, uses preset posture algorithms and key bone joint coordinates to judge the violations, and combines the feature data of work clothes, safety helmets and armbands to identify identity and permissions.

Benefits of technology

It enables the timely detection and regulation of violations by workshop staff without infringing on personal privacy, forming a strong confidentiality and security network and effectively preventing technology leaks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600588A_ABST
    Figure CN121600588A_ABST
Patent Text Reader

Abstract

The invention discloses a user behavior recognition method and a related device, and relates to the field of image processing, and the method comprises the steps: carrying out the detection object recognition of a monitoring image of a monitoring region, carrying out the analysis to obtain the motion behavior of a detection object in the monitoring image, and when the motion behavior belongs to a preset behavior set, carrying out the recognition of the detection object; the method comprises the following steps: acquiring a detection object in a monitoring area, performing identity recognition on the detection object, and when the identity recognition result is that the detection object does not have the authority of the action behavior, generating warning information and storing image information of detecting that the detection object executes the action behavior, so that whether the detection object in the monitoring area has an illegal behavior or not can be found in time. The traditional scheme is matched with the user behavior identification method disclosed by the invention to form a powerful workshop confidentiality security network, so that the limitation of manual inspection is effectively made up, and the behaviors of the workers in the workshop are reliably and safely standardized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a user behavior recognition method and related apparatus. Background Technology

[0002] In the existing scheme, in order to ensure safe production in the factory workshop and keep the company's technology confidential, it is necessary to regulate the behavior of the workers in the workshop and prohibit them from engaging in certain behaviors that are likely to cause safety accidents or technology leaks.

[0003] The most common form of technology leakage is taking photos with a mobile phone. To prevent this, current workshop security management primarily involves temporarily confiscating workers' mobile phones at the workshop entrance. However, relying solely on confiscation is insufficient for maintaining workshop security. This is because the subjective judgment or negligence of security personnel can lead to inadequate security controls, allowing undetected phones to be brought into the workshop. For example, during peak hours, security personnel may miss phones due to pressure or fatigue when faced with a large influx of employees. This means that even with explicit regulations prohibiting mobile phones, the actual enforcement may be compromised, leading to information leaks within the factory. Therefore, reliably regulating the behavior of workers within the workshop has become a pressing technical problem for those skilled in the art. Summary of the Invention

[0004] In view of the above problems, this application provides a user behavior recognition method and related apparatus to achieve the purpose of safely regulating the behavior of workers in the workshop. The specific solution is as follows:

[0005] The first aspect of this application provides a user behavior recognition method, including:

[0006] A user behavior recognition method, comprising:

[0007] Acquire surveillance images of the monitored area;

[0008] When a detection object exists, acquire the action behavior of the detection object;

[0009] If the action of the detected object belongs to a preset set of actions, the detected object is identified. The preset set of actions contains at least one action that is not allowed.

[0010] When the identity recognition result indicates that the detected object does not have the permission to perform the action, a warning message is generated and the image information of the detected object performing the action is saved.

[0011] Optionally, in the above user behavior recognition method, obtaining the action behavior of the detected object includes:

[0012] A preset pose algorithm is used to analyze the detected objects in the monitoring image to obtain the action behavior of the detected objects;

[0013] Alternatively, the key bone joint coordinates of the detected object can be obtained, and the action behavior of the detected object can be identified based on the key bone joint coordinates.

[0014] Optionally, in the above user behavior recognition method, the identification of the detected object includes:

[0015] Feature extraction is performed on the detected object in the monitoring image to obtain feature data that characterizes the identity information of the detected object;

[0016] The identity information of the detected object is identified based on the feature data.

[0017] Optionally, in the above user behavior recognition method, feature extraction of the detected object in the monitoring image includes:

[0018] The system identifies the feature data of the work clothes of the object being tested, the feature data of the safety helmet of the object being tested, and / or the feature data of the armband of the object being tested;

[0019] Identifying the identity information of the detected object based on the feature data includes:

[0020] The feature data is compared with preset feature data, which is the feature data of the work clothes, safety helmet or armband of a first identity user who has the permission to perform the action, or the feature data of the work clothes, safety helmet or armband of a second identity user who does not have the permission to perform the action.

[0021] Based on the comparison results, it is determined whether the detected object is a first-identity user with the permission to perform the action or a second-identity user without the permission to perform the action.

[0022] Optionally, in the above user behavior recognition methods,

[0023] After acquiring the monitoring image of the monitored area and before acquiring the action behavior of the detected object, the process also includes:

[0024] Obtain feature information from the collected images of the target object to be monitored;

[0025] The presence of the target object in the monitoring image is determined based on the feature information of the acquired image of the target object.

[0026] Optionally, in the above user behavior recognition method, after determining that a detection object exists in the monitoring image and before acquiring the action behavior of the detection object, the method further includes:

[0027] Determine whether the number of detected objects in the monitored image is greater than 1;

[0028] When the number of detected objects is greater than 1, the detected objects are bound by ID;

[0029] Track each detected object based on the bound ID.

[0030] Optionally, in the above user behavior recognition method, a preset pose algorithm is used to analyze the detected object in the monitoring image to obtain the action behavior of the detected object, including:

[0031] A preset posture algorithm is used to analyze the detected objects in the monitoring image to determine whether the detected objects are using a mobile phone and whether they are using a mobile phone to take pictures or record videos.

[0032] Obtain the key bone joint coordinates of the detected object, and identify the action behavior of the detected object based on the key bone joint coordinates, including:

[0033] The key bone joint coordinates of the detected object and the position coordinates of the image acquisition device carried by the detected object are obtained. Based on the key bone joint coordinates and the position coordinates of the image acquisition device, it is determined whether the detected object is using the image acquisition device to take pictures or record videos.

[0034] A user behavior recognition device, comprising:

[0035] The image acquisition unit is used to acquire monitoring images of the monitored area;

[0036] The behavior recognition unit is used to acquire the action behavior of the detected object when there is a detected object in the monitoring image;

[0037] An object identification unit is used to identify the detected object if the action behavior of the detected object belongs to a preset set of behaviors, wherein the preset set of behaviors includes at least one disallowed action behavior.

[0038] The warning unit is used to generate warning information and save image information of the detected object performing the action when the identity recognition result indicates that the detected object does not have the authority to perform the action.

[0039] A computer program product includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the user behavior recognition method as described in any of the preceding claims.

[0040] An electronic device includes at least one processor and a memory connected to the processor, wherein:

[0041] The memory is used to store computer programs;

[0042] The processor is used to execute the computer program to enable the electronic device to implement the user behavior recognition method as described in any of the above.

[0043] A computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the user behavior recognition method as described in any of the preceding claims.

[0044] By employing the above technical solution, the solution provided in this application identifies the detected objects in the monitoring images of the monitored area and analyzes the actions of the detected objects in the monitoring images. When the actions belong to a preset set of actions, the detected object is identified. When the identification result indicates that the detected object does not have the authority to perform the action, a warning message is generated and the image information of the detected object performing the action is saved. This allows for timely detection of any violations by the detected objects in the monitored area, ensuring safe production and preventing technology leaks. By combining traditional solutions with the user behavior identification method disclosed in this application, a robust workshop security network can be formed without infringing on personal privacy, effectively compensating for the limitations of manual inspection and reliably regulating the behavior of workers in the workshop. Attached Figure Description

[0045] The above and other features, advantages, and aspects of the embodiments disclosed in this application will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0046] Figure 1 A schematic diagram of an implementation system architecture for the user behavior recognition method provided in this application embodiment;

[0047] Figure 2 A schematic diagram of a terminal structure provided in an embodiment of this application;

[0048] Figure 3 A schematic diagram of a server structure provided in an embodiment of this application;

[0049] Figure 4 This is a development flow diagram of the server disclosed in the embodiments of this application;

[0050] Figure 5 This is a flowchart illustrating a user behavior recognition method disclosed in an embodiment of this application;

[0051] Figure 6 This is a schematic diagram of the structure of a user behavior recognition device disclosed in an embodiment of this application;

[0052] Figure 7 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. Detailed Implementation

[0053] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0054] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0055] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0056] This application can be applied to the field of image processing. The following section will introduce several application scenarios that have been implemented in products, taking user behavior recognition as an example.

[0057] First, let's introduce the application scenarios of this application.

[0058] This application can be applied, but is not limited to, to applications with image processing capabilities or cloud services provided by cloud-side servers, which will be described in detail below:

[0059] See Figure 1 , Figure 1A schematic diagram of a system architecture is shown. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers (…). Figure 1 (The example includes a server), and the server 200 can provide the method provided in the embodiments of this application to one or more terminals.

[0060] The terminal 100 may be equipped with an application for implementing user behavior recognition methods. The application and webpage can provide an interface. The terminal 100 can receive relevant parameters input by the user on the interactive interface and send the parameters to the server 200. The server 200 can obtain the processing result based on the received parameters and return the processing result to the terminal 100. At this time, the image acquisition device can send the acquired image to the server 200 through the terminal 100 or directly.

[0061] It should be understood that in some optional implementations, the terminal 100 can also complete the action of obtaining the processing result based on the received parameters on its own, without the need for the server to cooperate. This application embodiment is not limited to this. In this case, the image acquisition device can directly send the acquired image to the terminal 100.

[0062] The following description Figure 1 The product form of the mid-terminal 100;

[0063] The terminal 100 in this application embodiment can be a mobile phone, tablet computer, wearable device, vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc., and this application embodiment does not impose any restrictions on it.

[0064] Figure 2 A schematic diagram of an optional hardware structure for terminal 100 is shown.

[0065] refer to Figure 2 As shown, the terminal 100 may include a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), a headphone jack 163 (optional), a processor 170, an external interface 180, a power supply 190, and other components. Those skilled in the art will understand that... Figure 2These are merely examples of terminals or multi-functional devices and do not constitute a limitation on terminals or multi-functional devices. They may include more or fewer components than shown, or combine certain components, or different components.

[0066] The input unit 130 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the portable multi-functional device. Specifically, the input unit 130 may include a touchscreen 131 (optional) and / or other input devices 132. The touchscreen 131 can collect touch operations performed by the user on or near it (such as operations performed by the user using fingers, knuckles, styluses, or any suitable object on or near the touchscreen), and drive the corresponding connection devices according to a pre-set program. The touchscreen can detect the user's touch actions, convert the touch actions into touch signals and send them to the processor 170, and can receive and execute commands sent by the processor 170; the touch signal includes at least touch point coordinate information. The touchscreen 131 can provide an input interface and an output interface between the terminal 100 and the user. In addition, various types of touchscreens, such as resistive, capacitive, infrared, and surface acoustic wave, can be used to implement the touchscreen. Besides the touchscreen 131, the input unit 130 may also include other input devices. Specifically, other input devices 132 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0067] Among them, the input device 132 can receive input data, etc.

[0068] The display unit 140 can be used to display information input by the user or information provided to the user, various menus of the terminal 100, interactive interfaces, file display, and / or playback of any multimedia file. In this embodiment, the display unit 140 can be used to display acquired video, processing results, etc.

[0069] The memory 120 can be used to store instructions and data. The memory 120 may primarily include an instruction storage area and a data storage area. The data storage area can store various types of data, such as multimedia files and text. The instruction storage area can store software units such as operating systems, applications, and instructions required for at least one function, or subsets or extended sets thereof. It may also include non-volatile random access memory. It provides the processor 170 with hardware, software, and data resources for managing the computing device, supporting control software and applications. It is also used for storing multimedia files, as well as storing running programs and applications.

[0070] The processor 170 is the control center of the terminal 100. It connects various parts of the terminal 100 via various interfaces and lines. By running or executing instructions stored in the memory 120 and calling data stored in the memory 120, it performs various functions of the terminal 100 and processes data, thereby controlling the terminal device as a whole. Optionally, the processor 170 may include one or more processing units; preferably, the processor 170 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 170. In some embodiments, the processor and memory can be implemented on a single chip; in some embodiments, they can also be implemented separately on independent chips. The processor 170 can also be used to generate corresponding operation control signals, send them to the corresponding components of the computing processing device, read and process data in the software, especially read and process data and programs in the memory 120, so that the various functional modules therein perform corresponding functions, thereby controlling the corresponding components to act according to the instructions.

[0071] The memory 120 can be used to store software code related to the user behavior recognition method, and the processor 170 can execute the steps of the user behavior recognition method and can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to achieve the corresponding functions.

[0072] The radio frequency unit 110 (optional) can be used for receiving and transmitting signals during information transmission or calls. For example, it can receive downlink information from the base station and process it for the processor 170; additionally, it can transmit uplink data to the base station. Typically, the RF circuit includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, the radio frequency unit 110 can also communicate wirelessly with network devices and other devices. This wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0073] In this embodiment of the application, the radio frequency unit 110 can send data to the server 200 and receive the processing results sent by the server 200.

[0074] It should be understood that the radio frequency unit 110 is optional and can be replaced with other communication interfaces, such as a network port.

[0075] The terminal 100 also includes a power supply 190 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 170 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.

[0076] Terminal 100 also includes an external interface 180, which can be a standard Micro USB interface or a multi-pin connector, which can be used to connect terminal 100 to other devices for communication or to connect a charger to charge terminal 100.

[0077] Although not shown, terminal 100 may also include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with various functions, etc., which will not be described in detail here. Some or all of the methods described below can be applied to, for example... Figure 2 In the terminal 100 shown.

[0078] The following description Figure 1 The product form of the mid-range server 200;

[0079] Figure 3 A structural diagram of a server 200 is provided, as follows: Figure 3 As shown, server 200 includes bus 201, processor 202, communication interface 203, and memory 204. Processor 202, memory 204, and communication interface 203 communicate with each other via bus 201.

[0080] Bus 201 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0081] The processor 202 can be any one or more of the following processors: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).

[0082] Memory 204 may include volatile memory, such as random access memory (RAM). Memory 204 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0083] The memory 204 can be used to store software code related to the user behavior recognition method, and the processor 202 can execute the steps of the user behavior recognition method of the chip, and can also schedule other units to achieve the corresponding functions.

[0084] It should be understood that the aforementioned terminal 100 and server 200 can be centralized or distributed devices. The processors (e.g., processor 170 and processor 202) in the aforementioned terminal 100 and server 200 can be hardware circuits (such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), general-purpose processors, digital signal processors (DSPs), microprocessors or microcontrollers, etc.) or combinations of these hardware circuits. For example, the processor can be a hardware system with instruction execution capabilities, such as a CPU or DSP, or a hardware system without instruction execution capabilities, such as an ASIC or FPGA, or a combination of the aforementioned hardware systems without instruction execution capabilities and hardware systems with instruction execution capabilities.

[0085] In one specific embodiment disclosed in this application, the specific process of the user behavior recognition method can be implemented in a server. The server sends the processing result to the client, and the client displays the monitoring image and the processor's recognition result through a browser. In this case, the development process of the server includes:

[0086] Step S1: Acquire video data stream;

[0087] The acquisition of video data streams can be achieved using the DEEPSTREAM framework or other frameworks. DEEPSTREAM is a library based on the open-source video streaming framework GStreamer, used for building artificial intelligence applications. It supports acquiring video data streams from various data sources (such as USB / CSI cameras, RTSP streams, etc.) and analyzing them using artificial intelligence and computer vision techniques. In the DEEPSTREAM framework, acquiring the video data stream is the first step in building a video analysis workflow. It involves capturing video data from a selected data source and transmitting it to the framework for subsequent processing.

[0088] Step S2: Algorithm model development and deployment;

[0089] The algorithm model refers to the visual algorithm model used in this application to implement the user behavior recognition method. The specific tasks of the model may include object recognition, object pose recognition, object action recognition, and determining whether the current action of the object is a prohibited action. The visual algorithm model mainly involves frameworks such as the POSEC3D motion vision recognition framework, YOLOv8, and other object recognition and pose recognition frameworks. This application can set the specific workflow and content of each framework according to specific needs. During model development, the video data stream obtained in step S1 can be used as a data sample to train the algorithm model. Then, the trained algorithm model is converted and deployed to the server.

[0090] Since this solution is applied to real-time detection scenarios and requires timely display of the processing results of monitored images, it also features specialized optimizations for model algorithm deployment. The code within the system framework (original code can be PYTHON (Guido van Rossum)) is converted to C++, and the backend processing system is developed using C++, including video stream reception, preprocessing, and model computation. To achieve high-performance video stream parsing, this solution uses the DEEPSTREAM framework on the server side to develop video processing and analysis algorithms, calls them through corresponding data interfaces, and integrates TensorRT Engine for hardware acceleration. Ultimately, it enables rapid analysis of the acquired video stream using the aforementioned algorithm model.

[0091] Step S3: Front-end development;

[0092] The aforementioned front-end development refers to the development of the client browser interface, which includes UI design and front-end / back-end interaction development.

[0093] In this solution, the client can also directly obtain the image information corresponding to the video data stream from the video source and display it through the browser. In this case, the user can view the video stream, the server's analysis results, and receive alarms in real time through the browser. Since the algorithm model uses the POSEC3D motion vision recognition framework, which requires a relatively long time to process alarm results, in order to obtain alarm results instantly, the client can use an independent thread to run and display the algorithm results of the alarm frames to display the video alarm results.

[0094] To address the aforementioned problems, this application provides a user behavior recognition method. The user behavior recognition method of this application embodiment will be described in detail below with reference to the accompanying drawings.

[0095] Reference Figure 5 , Figure 5This is a flowchart illustrating a user behavior recognition method provided in an embodiment of this application. The algorithm model described above is used to implement this user behavior recognition method, such as... Figure 5 As shown in the embodiment of this application, a data processing method may include steps 501 to 405, which are described in detail below.

[0096] Step S501: Obtain monitoring images of the monitored area.

[0097] The monitored area refers to the coverage area of ​​the image acquisition device (e.g., a camera), and the monitored image can refer to the video data stream acquired by the image acquisition device. In this step, the video data stream acquired by the image acquisition device can be obtained through a wireless network or a data cable.

[0098] Step S502: When a detection object exists, obtain the action behavior of the detection object.

[0099] The detection target can refer to the staff within the monitored area. In this step, after acquiring the monitoring image, it can be preprocessed before determining whether the detection target exists in the preprocessed image. The preprocessing can refer to video frame extraction, which involves converting the video stream (monitoring image) into a multi-dimensional spatiotemporal image array to be identified. That is, each frame of the video (i.e., spatial information) is arranged chronologically to form an image sequence containing multiple time points. This image sequence can be considered a multi-dimensional spatiotemporal image array.

[0100] Then, the preprocessed surveillance images are used to identify potential targets. For example, the preprocessed surveillance images are analyzed to determine if there are any workers present; these workers are the potential targets. When a potential target is identified, pose recognition technology is used to analyze the pose of the target in the surveillance image, extracting key points and skeleton information. This allows for the modeling and recognition of the target's pose, and based on the recognition results, the target's actions are determined.

[0101] Step S503: When the action behavior of the detected object belongs to the preset behavior set, the detected object is identified.

[0102] In this step, a preset set of behaviors can be pre-configured. This set includes at least one prohibited action, such as taking a photo or other unauthorized operation of mechanical equipment. After identifying the action of the detected object, it is determined whether the action belongs to the prohibited actions in the preset set. If the action is not included in the preset set, it indicates that the detected object has not performed an unauthorized action, and no further steps are required. If the action is included in the preset set, it indicates that the detected object is performing an unauthorized action, and further steps are required.

[0103] The permissions for actions vary depending on the job type of the staff. For example, if the job type of the detected object is A, it is allowed to take photos, but if the job type of the detected object is B, it is not allowed to take photos. Therefore, in this step, when the action of the detected object belongs to the preset action set, it is necessary to identify the detected object and then determine the job type of the detected object based on the identification result, and determine whether the detected object has the permission for the action.

[0104] Step S504: When the identity recognition result indicates that the detected object does not have the permission to perform the action, generate a warning message and save the image information of the detected object performing the action.

[0105] In this step, a set of prohibited behaviors that match the identity of the detected object can be pre-established. When the detected behavior belongs to the set of behaviors corresponding to the detected object, the detected behavior is a violation for the detected object, and the detected object does not have the permission to use the behavior. When the detected behavior does not belong to the set of behaviors corresponding to the detected object, the detected behavior is a permitted behavior for the detected object, and the detected object has the permission to use the behavior.

[0106] When it is detected that the monitored object does not have the permission to perform the action (e.g., the monitored object does not have the permission to take a picture), it indicates that the monitored object has committed a violation. At this time, in order to ensure safe production or prevent technology leakage, it is necessary to generate an alert message and save the image information of the monitored object performing the action. This image information can be a video clip. The alert message is used to notify security monitoring personnel that a staff member is performing a violation, so that security monitoring personnel can correct the violation of the monitored object in a timely manner.

[0107] As can be seen from the above scheme, this application identifies the detected objects in the monitoring images of the monitored area and analyzes the actions of the detected objects in the monitoring images. When the actions belong to a preset set of actions, the detected object is identified. When the identification result shows that the detected object does not have the authority to perform the action, a warning message is generated and the image information of the detected object performing the action is saved. This allows for timely detection of any violations by the detected objects in the monitored area, ensuring safe production and preventing technology leaks. By combining traditional solutions with the user behavior identification method disclosed in this application, a strong workshop security network can be formed without infringing on personal privacy, effectively compensating for the limitations of manual inspection and reliably regulating the behavior of workers in the workshop.

[0108] In the technical solutions disclosed in this application, when acquiring the action behavior of the detected object, a suitable algorithm can be selected to identify the action behavior of the detected object according to the design requirements. For example, in the technical solutions disclosed in this embodiment, a preset pose algorithm can be used to analyze the detected object in the monitoring image to obtain the action behavior of the detected object. The preset pose algorithm can be the POSEC3D (Pose Convolution 3D) algorithm. The POSEC3D algorithm is a skeletal behavior recognition framework based on 3D convolutional neural network (3D-CNN). The POSEC3D algorithm aims to efficiently extract the spatiotemporal features in the skeletal sequence through 3D-CNN, thereby achieving accurate recognition of human actions. The skeletal sequence is obtained by skeletal annotation of each frame image in the multidimensional spatiotemporal image array. For example, a preset pose algorithm can be used to analyze the detected object in the monitoring image to determine whether the detected object is using a mobile phone and whether it is using a mobile phone to take pictures or record videos. When it is determined that the detected object is taking pictures or recording videos, the taking pictures or recording videos is the action behavior of the user detected this time. Furthermore, since the positional relationships of the user's joints differ depending on the user's actions, this application can also obtain the key bone joint coordinates of the detected object in the monitoring image and identify the detected object's actions based on these key bone joint coordinates. When monitoring the detected object's actions of taking photos or recording videos, to ensure more accurate and reliable judgment results, this application can also incorporate the coordinates of the image acquisition device (which can be a mobile phone, camera, camcorder, or other device) into the user's action recognition process. In this case, the coordinates of the image acquisition device carried by the detected object and the key bone joint coordinates of the detected object are identified from the monitoring image. Then, based on the key bone joint coordinates and the position coordinates of the image acquisition device, it is determined whether the detected object is using the image acquisition device and whether it is taking photos or recording videos. When it is determined that the detected object is taking photos or recording videos, the taking photos or recording videos is the user's action detected this time.

[0109] Considering the differences in clothing worn by workers in different jobs, the identity information of the identified object can be determined based on their attire. Here, the identity information refers to the job type of the identified object. For example, the colors of workers' clothes, safety helmets, or armbands may differ depending on their job type. Therefore, when identifying the detected object, this application can extract features from the monitored image to obtain feature data characterizing the identity information of the detected object, and then identify the identity information of the detected object based on the feature data. The feature data can be feature data of work clothes, safety helmets, and / or armbands, more specifically, the color of the work clothes, the color of the safety helmet, and / or the color of the armband. These feature data allow for rapid identification of the detected object's identity information.

[0110] After obtaining the feature data (feature data of work clothes, safety helmet, or armband) of the detected object, the detected feature data is compared with preset feature data. Based on the comparison result, it is determined whether the detected object is a first identity user with the permission to perform the action or a second identity user without the permission to perform the action. The preset feature data is the feature data of the work clothes, safety helmet, or armband of the first identity user with the permission to perform the action, or the feature data of the work clothes, safety helmet, or armband of the second identity user without the permission to perform the action.

[0111] When the preset feature data is the feature data of the work clothes, safety helmet, or armband of a first-identified user with the pre-marked permission to perform the action, if the obtained feature data is the feature data of the work clothes, then the feature data is compared with the feature data of the work clothes in the preset feature data. If the two are consistent, it indicates that the detected object is a first-identified user with the permission to perform the action. If the obtained feature data is the feature data of the safety helmet, then the feature data is compared with the feature data of the safety helmet in the preset feature data. If the two are consistent, it indicates that the detected object is a first-identified user with the permission to perform the action. If the obtained feature data is the feature data of the armband, then the feature data is compared with the feature data of the armband in the preset feature data. If the two are consistent, it indicates that the detected object is a first-identified user with the permission to perform the action.

[0112] When the preset feature data is the feature data of the work clothes, safety helmet, or armband of a pre-marked second-identity user who does not have the permission to perform the action, if the obtained feature data is the feature data of the work clothes, then the feature data is compared with the feature data of the work clothes in the preset feature data. If the two are consistent, it indicates that the detected object is a second-identity user who does not have the permission to perform the action. If the obtained feature data is the feature data of the safety helmet, then the feature data is compared with the feature data of the safety helmet in the preset feature data. If the two are consistent, it indicates that the detected object is a second-identity user who does not have the permission to perform the action. If the obtained feature data is the feature data of the armband, then the feature data is compared with the feature data of the armband in the preset feature data. If the two are consistent, it indicates that the detected object is a second-identity user who does not have the permission to perform the action.

[0113] In some scenarios, the monitored population may only be certain specific groups, while other users do not need to be monitored. Users can configure the characteristic information of the specific groups to be monitored based on their monitoring needs. For example, the specific groups to be monitored may be people wearing yellow or white safety helmets. In this case, after acquiring the monitoring image of the monitored area, before acquiring the action behavior of the target, it is necessary to determine whether the target exists in the monitoring image. This process can be specifically as follows: based on the characteristic information of the acquired image of the target to be monitored, determine whether the target exists in the monitoring image, that is, determine whether there is a target carrying the characteristic information in the monitoring image. If there is a target carrying the characteristic information, that target is taken as the target to be monitored. For example, when the characteristic information is a yellow safety helmet, and there are multiple workers in the monitoring image, determine whether there are workers wearing yellow safety helmets in the monitoring image, and take the workers wearing yellow safety helmets as the target.

[0114] In the technical solution disclosed in this embodiment, when multiple objects to be monitored are detected in the detection image, in order to stably track and recognize the objects, this solution further includes, after determining that there are objects in the monitoring image and before obtaining the action behavior of the objects, determining whether the number of objects in the monitoring image is greater than 1. When the number of objects is greater than 1, matching IDs to each object and binding IDs to each object, so that each object is bound to an ID, and then tracking each object based on the bound ID, thereby maintaining the stability of the tracking of the objects.

[0115] The above describes a user behavior recognition method provided by the embodiments of this application. The following will describe the apparatus for performing the above user behavior recognition method.

[0116] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a user behavior recognition device provided in an embodiment of this application. Figure 6 As shown, the user behavior recognition device includes:

[0117] Image acquisition unit 10 is used to acquire monitoring images of the monitored area;

[0118] The behavior recognition unit 20 is used to acquire the action behavior of the detected object when there is a detected object in the monitoring image;

[0119] The object identification unit 30 is used to identify the object if the action of the detected object belongs to a preset set of actions, wherein the preset set of actions contains at least one disallowed action.

[0120] The warning unit 40 is used to generate warning information and save image information of the detected object performing the action when the identity recognition result indicates that the detected object does not have the authority to perform the action.

[0121] In one possible implementation, corresponding to the above method, the above device may further include an ID binding unit, which is used to determine whether the number of detected objects in the monitoring image is greater than 1; when the number of detected objects is greater than 1, the detected objects are ID-bound; and each detected object is tracked based on the bound ID.

[0122] The specific functions of the image acquisition unit 10, the object recognition unit 20, the behavior recognition unit 30, the monitoring action recognition unit 40, the object identity recognition unit 50, the permission detection and acquisition unit 60, and the warning unit 70 are described in the above method and will not be repeated here.

[0123] This application also provides an electronic device in its embodiments. (See reference...) Figure 7 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 7 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0124] like Figure 7As shown, the electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can execute various steps of the user behavior recognition methods disclosed in the embodiments of this application according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, the RAM 603 also stores various programs and data required for the operation of the electronic device. The processor 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0125] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, memory cards, hard drives, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0126] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the user behavior recognition methods provided in this application.

[0127] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the user behavior recognition methods provided in this application.

[0128] The user information (including but not limited to user image information) and data (including but not limited to data used for analysis, data stored, data displayed) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0129] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0130] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0131] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0132] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A user behavior recognition method, characterized in that, include: Acquire surveillance images of the monitored area; When a detection object exists, acquire the action behavior of the detection object; If the action of the detected object belongs to a preset set of actions, the detected object is identified. The preset set of actions contains at least one action that is not allowed. When the identity recognition result indicates that the detected object does not have the permission to perform the action, a warning message is generated and the image information of the detected object performing the action is saved.

2. The user behavior recognition method according to claim 1, characterized in that, Acquiring the action behavior of the detected object includes: A preset pose algorithm is used to analyze the detected objects in the monitoring image to obtain the action behavior of the detected objects; Alternatively, the key bone joint coordinates of the detected object can be obtained, and the action behavior of the detected object can be identified based on the key bone joint coordinates.

3. The user behavior recognition method according to claim 1, characterized in that, The identification of the detected object includes: Feature extraction is performed on the detected object in the monitoring image to obtain feature data that characterizes the identity information of the detected object; The identity information of the detected object is identified based on the feature data.

4. The user behavior recognition method according to claim 3, characterized in that, Feature extraction of the detected object in the monitoring image includes: The system identifies the feature data of the work clothes of the object being tested, the feature data of the safety helmet of the object being tested, and / or the feature data of the armband of the object being tested; Identifying the identity information of the detected object based on the feature data includes: The feature data is compared with preset feature data, which is the feature data of the work clothes, safety helmet or armband of a first identity user who has the permission to perform the action, or the feature data of the work clothes, safety helmet or armband of a second identity user who does not have the permission to perform the action. Based on the comparison results, it is determined whether the detected object is a first-identity user with the permission to perform the action or a second-identity user without the permission to perform the action.

5. The user behavior recognition method according to claim 1, characterized in that, After acquiring the monitoring image of the monitored area and before acquiring the action behavior of the detected object, the process also includes: Obtain feature information from the collected images of the target object to be monitored; The presence of the target object in the monitoring image is determined based on the feature information of the acquired image of the target object.

6. The user behavior recognition method according to claim 1, characterized in that, After determining that a detection object exists in the surveillance image, but before acquiring the action behavior of the detection object, the process includes: Determine whether the number of detected objects in the monitored image is greater than 1; When the number of detected objects is greater than 1, the detected objects are bound by ID; Track each detected object based on the bound ID.

7. The user behavior recognition method according to claim 2, characterized in that, The detection objects in the monitoring image are analyzed using a preset pose algorithm to obtain the action behavior of the detection objects, including: A preset posture algorithm is used to analyze the detected objects in the monitoring image to determine whether the detected objects are using a mobile phone and whether they are using a mobile phone to take pictures or record videos. Obtain the key bone joint coordinates of the detected object, and identify the action behavior of the detected object based on the key bone joint coordinates, including: The key bone joint coordinates of the detected object and the position coordinates of the image acquisition device carried by the detected object are obtained. Based on the key bone joint coordinates and the position coordinates of the image acquisition device, it is determined whether the detected object is using the image acquisition device to take pictures or record videos.

8. A user behavior recognition device, characterized in that, include: The image acquisition unit is used to acquire monitoring images of the monitored area; The behavior recognition unit is used to acquire the action behavior of the detected object when there is a detected object in the monitoring image; An object identification unit is used to identify the detected object if the action behavior of the detected object belongs to a preset set of behaviors, wherein the preset set of behaviors includes at least one disallowed action behavior. The warning unit is used to generate warning information and save image information of the detected object performing the action when the identity recognition result indicates that the detected object does not have the authority to perform the action.

9. A computer program product, characterized in that, It includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the user behavior recognition method as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the user behavior recognition method as described in any one of claims 1 to 7.

11. A computer storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the user behavior recognition method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Action type recognition method and device, storage medium and computer equipment

    CN110321852A

  • Identity recognition method for illegal object based on illegal behavior recognition

    CN113255542A

  • Video monitoring method and device, computer equipment and storage medium

    CN114173094A

  • Multi-person identity and action association recognition method and device and readable medium

    CN114419480A

  • Method and device for monitoring personnel in dangerous area based on machine vision

    CN116206255A