A target behavior detection method and device, and a target behavior detection system
By combining multiple information sources such as image, audio, and temperature field data, the confidence level of target behavior is determined, which solves the problems of high false negative rate and data leakage in existing technologies, and achieves more accurate target behavior detection and higher security.
Patent Information
- Application Number
- CN202511054092.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Existing target behavior detection methods have high false negative and false positive rates in adverse scenarios such as occlusion, backlighting, or dense crowds, and also pose security risks of data leakage and privacy breaches.
By combining image, audio, and temperature field data, the confidence level of target behavior is determined using a combination of multiple data sources, including human keypoint detection, audio feature analysis, temperature field gradient feature extraction, and behavior modeling. This avoids detection using single data sources and reduces data transmission to the cloud.
It improves the accuracy of target behavior detection, reduces the false negative rate, reduces the risk of data leakage, and enhances the security and privacy protection of the system.
Smart Images

Figure CN120564271B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of security monitoring, in particular to a target behavior detection method and device and a target behavior detection system. BACKGROUND
[0002] With the development of science and technology and the surge in public safety needs, public place safety monitoring systems play an increasingly important role in maintaining social order and protecting personal safety. Among them, using safety monitoring systems for target behavior detection is also becoming more and more important.
[0003] However, the existing monitoring system relies on single camera visual analysis when detecting target behavior, and has a high miss detection rate for target behavior in adverse scenes such as occlusion, backlight or crowd density. In addition, the existing target behavior detection method often uses a fixed threshold rule engine, which causes some misleading actions to be misjudged as target behavior. For example, some actions during basketball playing will be misjudged as fighting actions, affecting the accuracy of target behavior recognition. In addition, the existing target behavior detection method often needs to perform centralized data analysis and calculation in the cloud, and needs to return a large amount of high-definition video stream, which not only increases the complexity and computing cost of the system, but also increases the security risks of data leakage and privacy leakage.
[0004] For the existing target behavior detection method, there is no effective solution to the problems of high miss detection rate, low target recognition accuracy, and security risks of data leakage and privacy leakage. SUMMARY
[0005] Therefore, it is necessary to provide a target behavior detection method, device and system to solve the above technical problems.
[0006] In a first aspect, the present application provides a target behavior detection method. The method comprises:
[0007] obtaining a to-be-recognized image, a sequence frame corresponding to the to-be-recognized image, audio data corresponding to the to-be-recognized image, and temperature field data corresponding to the to-be-recognized image; the sequence frame corresponding to the to-be-recognized image is a preset number of images adjacent to the to-be-recognized image, which together with the to-be-recognized image constitutes a time sequence frame sequence;
[0008] determining a confidence of a target behavior of a detection target in the to-be-recognized image based on the obtained to-be-recognized image, sequence frame corresponding to the to-be-recognized image, audio data corresponding to the to-be-recognized image, and temperature field data corresponding to the to-be-recognized image;
[0009] Determine the confidence of the target behavior of the detection target in the to-be-identified image based on the confidence of the target behavior of the detection target in the to-be-identified image.
[0010] In one of the embodiments, the confidence of the target behavior of the detection target in the to-be-identified image is determined based on the to-be-identified image, the sequence frame corresponding to the to-be-identified image, the audio data corresponding to the to-be-identified image, and the temperature field data corresponding to the to-be-identified image, including:
[0011] The confidence of the target action of the detection target in the to-be-identified image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target are determined based on the to-be-identified image, the sequence frame corresponding to the to-be-identified image, the audio data corresponding to the to-be-identified image, and the temperature field data corresponding to the to-be-identified image.
[0012] The confidence of the target behavior of the detection target in the to-be-identified image is determined based on the confidence of the target action of the detection target in the to-be-identified image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target.
[0013] In one of the embodiments, the confidence of the target action of the detection target in the to-be-identified image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target are determined based on the to-be-identified image, the sequence frame corresponding to the to-be-identified image, the audio data corresponding to the to-be-identified image, and the temperature field data corresponding to the to-be-identified image, including:
[0014] The detection target in the to-be-identified image is determined based on the to-be-identified image.
[0015] The human key point detection is performed on the detection target in the sequence frame corresponding to the to-be-identified image to obtain the confidence of the target action of the detection target in the to-be-identified image.
[0016] The sound abnormality degree of the audio data corresponding to the to-be-identified image is determined based on the audio data corresponding to the to-be-identified image.
[0017] determine a thermal intensity of the limb contact area of the detection target in the to-be-identified image based on temperature field data corresponding to the to-be-identified image;
[0018] determine an entropy value corresponding to a behavior of the detection target in the to-be-identified image, a contact force field matching degree between a motion of the detection target and a preset target motion, and a panic index of a non-detection target in the to-be-identified image based on a frame sequence of the to-be-identified image.
[0019] In one of the embodiments, the human key point detection on the detection target in the sequence frame corresponding to the to-be-identified image is performed to obtain a confidence degree of a target motion of the detection target in the to-be-identified image, which includes:
[0020] The human key point detection is respectively performed on the detection target in the sequence frame corresponding to the to-be-identified image to obtain coordinates of human skeleton joints of the detection target in each image in the sequence frame.
[0021] Based on the coordinates of the human skeleton joints of the detection target in each image in the sequence frame, a fitting degree of a preset curve of the detection target, a joint acceleration mutation number, and a trunk inclination angle variance are determined; the fitting degree of the preset curve is a fitting degree between a motion curve of the detection target and a curve of a preset target motion; based on the fitting degree of the preset curve of the detection target, the joint acceleration mutation number, and the trunk inclination angle variance, a confidence degree of a target motion of the detection target in the to-be-identified image is determined.
[0022] In one of the embodiments, the determination of the sound abnormality degree of the audio data corresponding to the to-be-identified image based on the audio data corresponding to the to-be-identified image includes:
[0023] Based on the audio data corresponding to the to-be-identified image, a sound feature of the audio data is determined.
[0024] Based on the sound feature of the audio data, the sound abnormality degree of the audio data is determined.
[0025] In one of the embodiments, the determination of the thermal intensity of the limb contact area of the detection target in the to-be-identified image based on the temperature field data corresponding to the to-be-identified image includes:
[0026] Gradient features of the temperature field data corresponding to the to-be-identified image are extracted;
[0027] Based on the gradient features, a thermal intensity of the temperature field data corresponding to the to-be-identified image is determined.
[0028] determine a thermal intensity of a limb contact area of the detection target in the to-be-identified image based on the thermal intensity of the temperature field data corresponding to the to-be-identified image.
[0029] In one of the embodiments, the determining, based on the frame sequence of the to-be-identified image, of an entropy value corresponding to the behavior of the detection target in the to-be-identified image, a contact force field matching degree of the action of the detection target and a preset target action, and a panic index of a non-detection target in the to-be-identified image, comprises:
[0030] modeling, based on the frame sequence of the to-be-identified image, the behavior of the detection target in the to-be-identified image in the frame sequence to obtain a state transition matrix of the detection target in the to-be-identified image in the frame sequence;
[0031] determining, based on the state transition matrix, the entropy value corresponding to the behavior of the detection target in the to-be-identified image;
[0032] determining a joint point pressure distribution spatiotemporal feature of the detection target in the to-be-identified image;
[0033] determining, based on the joint point pressure distribution spatiotemporal feature of the detection target in the to-be-identified image and a joint point pressure distribution spatiotemporal feature corresponding to a preset target action template, the contact force field matching degree of the action of the detection target in the to-be-identified image and the preset target action;
[0034] calculating, based on an expression feature of a face of the non-detection target in the to-be-identified image, the panic index of the non-detection target in the to-be-identified image; the expression feature comprises a pupil dilation rate and a limb withdrawal speed.
[0035] In one of the embodiments, the determining, based on the confidence of the target action of the detection target in the to-be-identified image, a sound abnormality degree of audio data, a thermal intensity of a limb contact area of the detection target, an entropy value corresponding to the behavior of the detection target, a contact force field matching degree of the action of the detection target and a preset target action, and a panic index of a non-detection target, of a confidence of a target behavior of the detection target in the to-be-identified image, comprises:
[0036] determining, based on the confidence of the target action of the detection target in the to-be-identified image, the panic index of the non-detection target, the thermal intensity of the limb contact area of the detection target, and the sound abnormality degree of the audio data, whether the detection target in the to-be-identified image satisfies a preset first judgment condition; the preset first judgment condition is that the action satisfies a preset first action preset condition and the sound satisfies a preset sound preset condition;
[0037] when the detection target in the to-be-identified image meets the preset first judgment condition, judging whether the detection target in the to-be-identified image meets a preset second judgment condition based on an entropy value corresponding to a behavior of the detection target in the to-be-identified image, a contact force field matching degree of the action of the detection target and the preset target action, and a panic index of the non-detection target; the preset second judgment condition is at least one of a preset second action preset condition or a preset collision condition;
[0038] when the detection target in the to-be-identified image meets the preset second judgment condition, determining the confidence degree of the target behavior of the detection target in the to-be-identified image based on a confidence degree of a target action of the detection target in the to-be-identified image, the entropy value corresponding to the behavior of the detection target, the panic index of the non-detection target, and the contact force field matching degree of the action of the detection target and the preset target action.
[0039] In a second aspect, the present application further provides a target behavior detection device. The device comprises:
[0040] a data acquisition module configured to acquire a to-be-identified image, a sequence frame corresponding to the to-be-identified image, audio data corresponding to the to-be-identified image, and temperature field data corresponding to the to-be-identified image; the sequence frame corresponding to the to-be-identified image is a preset number of image frames adjacent to the to-be-identified image, and the to-be-identified image together constitutes a time sequence frame sequence;
[0041] a confidence degree determination module configured to determine a confidence degree of a target behavior of a detection target in the to-be-identified image based on the acquired to-be-identified image, the sequence frame corresponding to the to-be-identified image, the audio data corresponding to the to-be-identified image, and the temperature field data corresponding to the to-be-identified image;
[0042] and a detection module configured to determine a detection result of the target behavior in the to-be-identified image based on the confidence degree of the target behavior of the detection target in the to-be-identified image.
[0043] In a third aspect, the present application further provides a target behavior detection system, which comprises a data acquisition module, a feature extraction module, a classification decision module, and an optimization deployment module.
[0044] The data acquisition module comprises an infrared imaging sensing unit, a skeletal key point trajectory acquisition unit and a high-frequency voiceprint acquisition unit; the infrared imaging sensing unit is configured to acquire instantaneous temperature changes of a limb contact area in a to-be-identified image, and obtain temperature field data corresponding to the to-be-identified image; the skeletal key point trajectory acquisition unit is configured to perform human key point detection on a detection target in a sequence frame corresponding to the to-be-identified image, and obtain coordinates of human skeletal joint points of the detection target in each image in the sequence frame; and the high-frequency voiceprint acquisition unit is configured to acquire audio data corresponding to the to-be-identified image.
[0045] The feature extraction module is configured to extract gradient features of the temperature field data corresponding to the to-be-identified image; determine thermal intensity of the temperature field data corresponding to the to-be-identified image based on the gradient features of the temperature field data corresponding to the to-be-identified image; determine a fitting degree of a preset curve of the detection target in the to-be-identified image, a number of joint acceleration mutations and a trunk inclination angle variance of the detection target based on the coordinates of the human skeletal joint points of the detection target in each image in the sequence frame; the fitting degree of the preset curve is a fitting degree of a motion curve of the detection target and a curve of a preset target motion; determine a confidence degree of the target motion of the detection target in the to-be-identified image based on the fitting degree of the preset curve of the detection target in the to-be-identified image, the number of joint acceleration mutations and the trunk inclination angle variance; determine a sound feature of the audio data based on the audio data corresponding to the to-be-identified image; and determine a sound abnormality degree of the audio data based on the sound feature of the audio data.
[0046] The feature extraction module is further configured to calculate a KL divergence of a state transition matrix of the detection target in the sequence frame by using a hidden Markov model, and obtain an entropy value corresponding to a behavior of the detection target; determine a joint point pressure distribution space-time feature of the detection target in the to-be-identified image by using a contact force field model; determine a contact force field matching degree of a motion of the detection target in the to-be-identified image and a preset target motion based on the joint point pressure distribution space-time feature of the detection target in the to-be-identified image and a joint point pressure distribution space-time feature corresponding to a preset target motion template; extract an expression feature of a face of a non-detection target in the to-be-identified image by using a micro-expression extraction model; the expression feature comprises a pupil dilation rate and a limb withdrawal speed; and determine a panic index of the non-detection target in the to-be-identified image based on the expression feature of the face of the non-detection target in the to-be-identified image.
[0047] The classification decision module is configured to determine whether the detection target in the to-be-identified image meets a preset first judgment condition based on the confidence of the target action of the detection target in the to-be-identified image, the panic index of the non-detection target, the thermal intensity of the limb contact area of the detection target, and the sound abnormality of the audio data; the preset first judgment condition is that the action meets a preset first action preset condition and the sound meets a preset sound preset condition; when the detection target in the to-be-identified image meets the preset first judgment condition, it is determined whether the detection target in the to-be-identified image meets a preset second judgment condition based on the entropy value corresponding to the behavior of the detection target in the to-be-identified image, the contact force field matching degree between the action of the detection target and the preset target action, and the panic index of the non-detection target in the to-be-identified image; the preset second judgment condition is that the action meets at least one of a preset second action preset condition or a preset collision condition; when the detection target in the to-be-identified image meets the preset second judgment condition, the confidence of the target behavior of the detection target in the to-be-identified image is determined based on the confidence of the target action of the detection target in the to-be-identified image, the entropy value corresponding to the behavior of the detection target, the panic index of the non-detection target, and the contact force field matching degree between the action of the detection target and the preset target action.
[0048] The optimization deployment module is configured to determine a detection result of the target behavior of the detection target in the to-be-identified image based on the confidence of the target behavior of the detection target in the to-be-identified image, generate an adversarial sample of the target behavior based on the detection result of the target behavior of the detection target in the to-be-identified image and an artificial detection result, and update the model in the feature extraction module by using the adversarial sample of the target behavior.
[0049] The target behavior detection method, device and system determine the confidence of the target behavior of the detection target in the to-be-identified image by acquiring the to-be-identified image, the sequence frame corresponding to the to-be-identified image, the audio data corresponding to the to-be-identified image, and the temperature field data corresponding to the to-be-identified image, and then determine the detection result of the target behavior in the to-be-identified image based on the confidence of the target behavior of the detection target in the to-be-identified image. The confidence of the target behavior of the detection target in the to-be-identified image is determined by combining the to-be-identified image, the sequence frame corresponding to the to-be-identified image, the audio data and the temperature field data, so that the detection result of the target behavior is more accurate, the missed detection caused by single data detection is avoided, and when the target behavior is detected, the data to be detected does not need to be transmitted to the cloud, thereby solving the problems of the existing target behavior detection method, such as high missed detection rate, low target recognition accuracy, and security risks of data leakage and privacy leakage.
[0050] The details of one or more embodiments of the application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the application will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF DRAWINGS
[0051] The accompanying drawings are included to provide a further understanding of the application, and are incorporated in and constitute a part of this application, illustrate embodiments of the application, and explain them, which do not limit the present application. In the drawings:
[0052] Figure 1 The hardware structure block diagram of the terminal of the target behavior detection method provided by an embodiment of the application is shown in the figure;
[0053] Figure 2 The flow chart of the target behavior detection method provided by an embodiment of the application is shown in the figure;
[0054] Figure 3 The flow chart of the target behavior detection method provided by a preferred embodiment of the application is shown in the figure;
[0055] Figure 4 The structure block diagram of the target behavior detection device provided by an embodiment of the application is shown in the figure. DETAILED DESCRIPTION
[0056] In order to more clearly understand the object, technical scheme and advantages of the application, the application is described and explained in the following with reference to the drawings and embodiments.
[0057] Unless otherwise defined, technical terms or scientific terms used in the present application shall have the same meaning as those commonly understood by a person of ordinary skill in the art to which the present application belongs. The terms "one", "a", "an", "the", "these", and similar terms in the present application do not mean "only one" or "exactly one", but can mean "one or more". The terms "include", "contain", "have", and any variant thereof in the present application are intended to cover the non-exclusive inclusion; for example, a process, method, and system, product or device containing a series of steps or modules (units) are not limited to the listed steps or modules (units), but can include steps or modules (units) not listed, or can include other steps or modules (units) inherent to the process, method, product or device. The terms "connect", "connect", "couple" and the like in the present application are not limited to physical or mechanical connection, but can include electrical connection, whether direct or indirect. The term "multiple" in the present application means two or more. The term "and / or" describes the association between the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that A exists alone, A and B exist together, and B exists alone. Generally, the character " / " represents the relationship between the objects before and after it as "or". The terms "first", "second", "third" and the like in the present application are only used to distinguish similar objects, and do not represent a specific order of the objects.
[0058] The method embodiments provided in the present embodiment can be executed in a terminal, a computer or a similar computing device. For example, the method embodiments are executed on a terminal, Figure 1 is a hardware structure diagram of the terminal of the target behavior detection method of the present embodiment. As shown in Figure 1 , the terminal can include one or more (only one is shown in Figure 1 ) processor 102 and memory 104 for storing data, wherein the processor 102 can include but not limited to processing devices such as microprocessor MCU or programmable logic device FPGA. The above terminal can also include transmission device 106 for communication function and input / output device 108. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above terminal. For example, the terminal can also include more or less components than those shown in Figure 1 , or have a different configuration from that shown in Figure 1 .
[0059] The memory 104 can be used to store computer programs, such as software programs of application software and modules, such as a computer program corresponding to the target behavior detection method in the embodiment. The processor 102 can execute various functional applications and data processing, i.e., implement the method described above, by running the computer programs stored in the memory 104. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include memories remotely arranged with respect to the processor 102, which can be connected to the terminal through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0060] The transmission device 106 is configured to receive or send data via a network. The network includes a wireless network provided by a communication provider of the terminal. In an example, the transmission device 106 includes a network interface controller (NIC) which can be connected to other network devices through a base station so as to communicate with the Internet. In an example, the transmission device 106 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.
[0061] In the embodiment, a target behavior detection method is provided, Figure 2 is a flowchart of the target behavior detection method of the embodiment, as shown in Figure 2 The flowchart includes the following steps:
[0062] In step S210, an image to be recognized, a sequence frame corresponding to the image to be recognized, audio data corresponding to the image to be recognized, and temperature field data corresponding to the image to be recognized are obtained. The sequence frame corresponding to the image to be recognized is a preset number of images adjacent to the image to be recognized, which constitute a time sequence frame sequence together with the image to be recognized.
[0063] The above obtaining the to-be-identified image can be acquiring an image frame of the target region by using an image acquisition device. The above image acquisition device can be one of a visible light imaging device (such as a single-lens reflex or micro-single camera), an infrared or thermal imaging device (such as an infrared camera), a hyperspectral camera, or the like. The above obtaining the to-be-identified image can also be acquiring a key frame picture in a video stream in real time (acquiring according to a setting parameter of a resolution of 1080P and a frame rate of 25 fps) from a streaming media collected by a camera. Specifically, the key frame picture in the video stream can be extracted according to a preset time interval. The above preset time interval can be specifically set based on specific requirements, which is not specifically limited in this embodiment. For example, the above preset time interval can be 200 ms.
[0064] In addition, the above target region can be specifically set according to specific scenarios, for example, can be a teaching building corridor or a classroom region in a school scenario. The above preset number of images can be specifically set according to specific requirements, which is not specifically limited in this embodiment, as long as the to-be-identified image and the preset number of images adjacent to the to-be-identified image can form a sequence frame that can determine the confidence of the target behavior of the detection target in the to-be-identified image. The above obtaining the audio data corresponding to the to-be-identified image can be acquiring an audio stream in the streaming media collected by the camera by using a high-frequency voiceprint acquisition module, or can be acquiring the audio data corresponding to the to-be-identified image by using an independently installed audio acquisition device. The above audio acquisition device can be a MEMS (Micro-Electro-Mechanical Systems, micro-electromechanical system) ultrasonic microphone. It should be noted that the above audio acquisition device can also be other microphones or other audio acquisition devices, as long as the audio data corresponding to the to-be-identified image can be acquired, which is not specifically limited in this embodiment. The above obtaining the temperature field data corresponding to the to-be-identified image can be acquiring the temperature field data corresponding to the to-be-identified image by using an infrared thermal imaging sensor. The above temperature field data can be complete distribution information of temperature values of each point in space at the moment corresponding to the to-be-identified image. The resolution of the above temperature field data can be specifically set based on specific requirements, for example, the resolution of the temperature field data can be set to 320x240. The above infrared thermal imaging sensor can be one or more of a non-cooled infrared microbolometer, a cooled infrared focal plane array, an infrared thermal imager module, an infrared thermal imaging chip, or the like.
[0065] Step S220, determining the confidence of the target behavior of the detection target in the to-be-identified image based on the obtained to-be-identified image, the sequence frame corresponding to the to-be-identified image, the audio data corresponding to the to-be-identified image, and the temperature field data corresponding to the to-be-identified image.
[0066] In this step, the confidence of the target behavior of the detection target can be used to represent the credibility of the behavior of the detection target being the target behavior. The higher the confidence, the greater the possibility that the behavior of the detection target is the target behavior, and vice versa. The lower the confidence, the smaller the possibility that the behavior of the detection target is the target behavior. The target behavior can be specifically set according to specific requirements or application scenarios, which are not specifically limited in this embodiment. For example, the target behavior is fighting behavior. The confidence of the target behavior of the detection target in the to-be-recognized image based on the acquired to-be-recognized image, the sequence frame corresponding to the to-be-recognized image, the audio data corresponding to the to-be-recognized image, and the temperature field data corresponding to the to-be-recognized image can be the confidence of the target action of the detection target in the to-be-recognized image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target. Furthermore, the confidence of the target behavior of the detection target in the to-be-recognized image is determined based on the confidence of the target action of the detection target in the to-be-recognized image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target.
[0067] In step S230, the detection result of the target behavior in the to-be-recognized image is determined based on the confidence of the target behavior of the detection target in the to-be-recognized image.
[0068] The detection result includes the presence of the target behavior, the absence of the target behavior, and the suspicious behavior of whether the target behavior is present or not. The detection result of the target behavior in the to-be-recognized image is determined based on the confidence of the target behavior of the detection target in the to-be-recognized image. When the confidence of the target behavior of the detection target in the to-be-recognized image is greater than a preset first confidence threshold, it is determined that the target behavior is present in the to-be-recognized image. When the confidence of the target behavior of the detection target in the to-be-recognized image is less than a preset second confidence threshold, it is determined that the target behavior is not present in the to-be-recognized image. When the confidence of the target behavior of the detection target in the to-be-recognized image is between the preset second confidence threshold and the preset first confidence threshold, the behavior is suspicious and needs to be manually reviewed. The preset second confidence threshold is less than the preset first confidence threshold. The preset second confidence threshold and the preset first confidence threshold can be specifically set according to specific conditions, which are not specifically limited in this embodiment. For example, the preset second confidence threshold is 0.6, and the preset first confidence threshold is 0.8.
[0069] It should be noted that after determining the detection result of the target behavior in the image to be recognized, if the system outputs a high confidence alarm for 3 consecutive times but the manual review is a false alarm, the multi-modal features of the misjudgment samples will be automatically collected, the detection threshold and feature weight will be adjusted after comparative analysis, the misjudgment features will be added to the non-target behavior template library, the sample will be enhanced through the generation of an adversarial network and incremental training will be performed on the edge device, and after the backtest verifies that the false alarm rate is reduced by more than 80%, the feature library will be formally updated and synchronized to all nodes.
[0070] The steps S210 to S230 determine the confidence of the target behavior of the detection target in the image to be recognized by acquiring the image to be recognized, the sequence frame corresponding to the image to be recognized, the audio data corresponding to the image to be recognized, and the temperature field data corresponding to the image to be recognized, and then determine the detection result of the target behavior in the image to be recognized based on the confidence of the target behavior of the detection target in the image to be recognized. By combining the image to be recognized, the sequence frame corresponding to the image to be recognized, the audio data, and the temperature field data, the confidence of the target behavior of the detection target in the image to be recognized is determined, so that the detection result of the target behavior is more accurate, the missed detection caused by single data detection is avoided, and when the target behavior is detected, the data to be detected does not need to be transmitted to the cloud, thereby solving the problems of high missed detection rate, low target recognition accuracy, and security risks of data leakage and privacy leakage in the existing target behavior detection method.
[0071] In one embodiment, step S220 determines the confidence of the target behavior of the detection target in the image to be recognized based on the acquired image to be recognized, the sequence frame corresponding to the image to be recognized, the audio data corresponding to the image to be recognized, and the temperature field data corresponding to the image to be recognized, including:
[0072] Step S222 determines the confidence of the target action of the detection target in the image to be recognized, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target based on the acquired image to be recognized, the sequence frame corresponding to the image to be recognized, the audio data corresponding to the image to be recognized, and the temperature field data corresponding to the image to be recognized.
[0073] The confidence of the target action can be used to represent the reliability of the action of the detected target as the target action. The higher the confidence, the greater the possibility that the action of the detected target is the target action, and vice versa. The target action can be specifically set according to specific requirements or application scenarios, which are not specifically limited in the embodiment. For example, the target action is a punching action in a fighting behavior. The confidence of the target action of the detected target in the to-be-recognized image can be determined based on the obtained to-be-recognized image, the detected target in the to-be-recognized image is determined, and then human key point detection is performed on the detected target in the sequence frame corresponding to the to-be-recognized image to obtain the confidence of the target action of the detected target in the to-be-recognized image.
[0074] The sound abnormality degree can be used to represent the deviation of the audio data from the standard audio data of the current scene. The standard audio data of the current scene can be a pre-stored acoustic model corresponding to the current scene. The acoustic model can be a sound model of normal audio in the current scene. The sound abnormality degree can be a scalar value, which can be a value in the range of [0, 1]. The greater the value, the higher the probability of sound abnormality, and the smaller the value, the lower the probability of sound abnormality, indicating that the sound in the audio data is more normal.
[0075] Further, the thermal intensity of the limb contact area can be the intensity of the thermal signal of the limb contact area. The intensity of the thermal signal can be the instantaneous heat power or the instantaneous temperature increment per unit area, which can be calculated by the product of the temperature rise value of the limb contact area and the contact heat transfer coefficient. The temperature rise value of the limb contact area can be the average temperature rise value or the maximum temperature rise value. The thermal intensity of the limb contact area can be a value in the range of [0, 1].
[0076] The entropy value corresponding to the behavior can be the predictability or regularity of the current behavior pattern of the detected target, that is, used to represent the severity of the current behavior anomaly of the detected target. The greater the entropy value corresponding to the behavior, the greater the difference between the current action sequence and the "normal habit pattern", and the more unpredictable and abnormal the behavior. The smaller the entropy value corresponding to the behavior, the higher the compliance of the action sequence to the normal regularity, and the more stable and predictable the behavior. The entropy value corresponding to the behavior can be a value in the range of [0, 1].
[0077] The contact force field matching degree can be used to measure the similarity in the time-space form between the joint-pressure distribution sequence of the detected target's action and the contact force template of the preset target action (such as fighting, falling, or carrying). The interval range of the numerical value of the contact force field matching degree can be set according to specific requirements, which is not specifically limited in the embodiment. For example, the contact force field matching degree can be set as a numerical value between 0 and 1. When the contact force field matching degree of the detected target's action and the preset target action is closer to 1, it indicates that the joint-pressure distribution sequence of the detected target's action is more similar to the contact force template of the preset target action in the time-space form, and the detected target's action is closer to the preset target action. When the contact force field matching degree of the detected target's action and the preset target action is closer to 0, it indicates that the joint-pressure distribution sequence of the detected target's action is less similar to the contact force template of the preset target action in the time-space form, and the detected target's action is less close to the preset target action.
[0078] In this step, the panic index of the non-detected target can be the panic index of the onlookers or other people (except the detected target) in the environment, which is used to represent the degree of panic or surprise of the crowd in the current scene due to the sudden event. The greater the numerical value of the panic index of the non-detected target, the more obvious the panic of the crowd, and the more intense the reaction of the crowd in the environment caused by the abnormal event; the smaller the numerical value of the panic index of the non-detected target, the more stable the emotion of the crowd. The behavior of the detected target can be determined by the panic index of the non-detected target.
[0079] In step S224, the confidence of the target behavior of the detected target in the to-be-recognized image is determined based on the confidence of the target action of the detected target in the to-be-recognized image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detected target, the entropy value corresponding to the behavior of the detected target, the contact force field matching degree of the detected target's action and the preset target action, and the panic index of the non-detected target.
[0080] In steps S222 to S224, the confidence of the target behavior of the detected target in the to-be-recognized image is determined based on the to-be-recognized image, the sequence frame corresponding to the to-be-recognized image, the audio data corresponding to the to-be-recognized image, and the temperature field data corresponding to the to-be-recognized image. The confidence of the target behavior of the detected target in the to-be-recognized image is determined, so that the detection result of the target behavior in the to-be-recognized image can be determined according to the confidence of the target behavior of the detected target in the to-be-recognized image.
[0081] Specifically, in one embodiment, step S222, based on the obtained to-be-recognized image, the sequence frame corresponding to the to-be-recognized image, the audio data corresponding to the to-be-recognized image, and the temperature field data corresponding to the to-be-recognized image, determine the confidence of the target action of the detection target in the to-be-recognized image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target, comprising:
[0082] Step S2221, based on the obtained to-be-recognized image, determine the detection target in the to-be-recognized image.
[0083] In this step, the above-mentioned detection target can be a target that needs to be detected in the current scene. For example, the above-mentioned detection target can be a person or an animal that needs to be detected. The above-mentioned detection target in the to-be-recognized image can be detected by using a preset detection method. The above-mentioned preset detection method can be specifically set according to specific needs, and the present embodiment is not specifically limited herein. For example, when detecting the target behavior of a person, a person with a distance less than a preset distance threshold can be regarded as a detection target. At this time, the preset detection method can be to use a trained neural network model to recognize the to-be-recognized image, and to determine a person with a distance less than a preset distance threshold as a detection target in the to-be-recognized image.
[0084] Step S2222, perform human key point detection on the detection target in the sequence frame corresponding to the to-be-recognized image to obtain the confidence of the target action of the detection target in the to-be-recognized image.
[0085] The confidence of the target action of the detected target in the to-be-identified image can be obtained by respectively performing human key point detection on the detected target in the sequence frames corresponding to the to-be-identified image. In this way, the coordinates of the human skeleton joints of the detected target in each image in the sequence frames are obtained. Then, based on the coordinates of the human skeleton joints of the detected target in each image in the sequence frames, the fitting degree of the preset curve, the number of joint acceleration mutations, and the trunk inclination angle variance of the detected target in the to-be-identified image are determined. The fitting degree of the preset curve is the fitting degree of the action curve of the detected target and the curve of the preset target action. Based on the fitting degree of the preset curve, the number of joint acceleration mutations, and the trunk inclination angle variance of the detected target in the to-be-identified image, the confidence of the target action of the detected target in the to-be-identified image is determined. The action curve of the detected target can be an action curve determined based on the position changes of one or more skeleton joints of the detected target in each image in the sequence frames. The preset target action can be a target action corresponding to a target behavior. For example, when the target behavior is fighting, the target action can be a punching or slapping action. The curve of the preset target action can be a position change curve of one or more skeleton joints corresponding to the preset target action. For example, when the preset target action is a punching action, the curve of the preset target action can be a position change curve of a hand skeleton joint during the punching action.
[0086] In step S2223, based on the audio data corresponding to the to-be-identified image, the sound abnormality degree of the audio data corresponding to the to-be-identified image is determined.
[0087] The determination of the sound abnormality degree of the audio data corresponding to the to-be-identified image based on the audio data corresponding to the to-be-identified image can include determining the sound features of the audio data based on the audio data corresponding to the to-be-identified image, and then determining the sound abnormality degree of the audio data based on the sound features of the audio data.
[0088] In step S2224, based on the temperature field data corresponding to the to-be-identified image, the thermal intensity of the limb contact area of the detected target in the to-be-identified image is determined.
[0089] The determination of the thermal intensity of the limb contact area of the detected target in the to-be-identified image based on the temperature field data corresponding to the to-be-identified image can include extracting gradient features of the temperature field data corresponding to the to-be-identified image, and then determining the thermal intensity of the temperature field data corresponding to the to-be-identified image based on the gradient features.
[0090] Step S2225, based on the frame sequence of the to-be-identified image, determine the entropy value corresponding to the behavior of the detection target in the to-be-identified image, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target in the to-be-identified image.
[0091] In this step, the entropy value corresponding to the behavior of the detection target in the to-be-identified image, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target in the to-be-identified image based on the frame sequence of the to-be-identified image can be based on the frame sequence of the to-be-identified image, modeling the behavior of the detection target in the to-be-identified image in the image frame sequence, obtaining the state transition matrix of the detection target in the to-be-identified image in the frame sequence, and then determining the entropy value corresponding to the behavior of the detection target in the to-be-identified image based on the state transition matrix. Then, determine the joint node pressure distribution spatiotemporal feature of the detection target in the to-be-identified image, and determine the contact force field matching degree of the action of the detection target in the to-be-identified image and the preset target action based on the joint node pressure distribution spatiotemporal feature of the detection target in the to-be-identified image and the joint node pressure distribution spatiotemporal feature corresponding to the preset target action template. Based on the expression features (pupil dilation rate and body shrinkage speed) of the face of the non-detection target in the to-be-identified image, calculate the panic index of the non-detection target in the to-be-identified image.
[0092] The steps S2221 to S2225 above determine the confidence of the target action of the detection target in the to-be-identified image, the sound abnormality degree of the audio data, the thermal intensity of the body contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target based on the to-be-identified image, the sequence frame corresponding to the to-be-identified image, the audio data corresponding to the to-be-identified image, and the temperature field data corresponding to the to-be-identified image. Through the determination of the confidence of the target action of the detection target in the to-be-identified image, the sound abnormality degree of the audio data, the thermal intensity of the body contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target, it is convenient for subsequent determination of the confidence of the target behavior of the detection target in the to-be-identified image.
[0093] In addition, in one embodiment, step S2222, human key point detection is performed on the detection target in the sequence frame corresponding to the to-be-identified image to obtain the confidence of the target action of the detection target in the to-be-identified image, comprising:
[0094] Step S11, respectively performing human key point detection on the detection target in the sequence frame corresponding to the to-be-identified image to obtain the coordinates of the human skeleton joints of the detection target in each image in the sequence frame.
[0095] In this step, the detection target in each image in the sequence frame is subjected to human key point detection, and the coordinates of the human skeleton joints of the detection target in each image in the sequence frame are obtained. The edge computing device can be used to perform grayscale normalization, noise reduction and other preprocessing on the received image to be identified after receiving the sequence frame corresponding to the image to be identified transmitted by the video monitoring platform, and then use the skeleton joint trajectory acquisition module to perform human key point detection on each image of the sequence frame, and real-time extract the two-dimensional coordinates of a preset number of joints in each image (the error of the extracted coordinates is required to be less than or equal to a pixel), and obtain the coordinates of the human skeleton joints of the detection target in each image in the sequence frame. The above-mentioned preset number can be set according to specific requirements, and the positions of the above-mentioned joints can also be set according to specific requirements, which are not limited in the embodiment. For example, the preset number can be set to 17, and the positions of the joints are respectively nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle.
[0096] In step S12, based on the coordinates of the human skeleton joints of the detection target in each image in the sequence frame, the fitting degree of the preset curve of the detection target in the image to be identified, the number of joint acceleration mutations, and the trunk inclination angle variance are determined; the fitting degree of the preset curve is the fitting degree of the action curve of the detection target and the curve of the preset target action; based on the fitting degree of the preset curve of the detection target in the image to be identified, the number of joint acceleration mutations, and the trunk inclination angle variance, the confidence of the target action of the detection target in the image to be identified is determined.
[0097] The determination of the fitting degree of the preset curve of the detected target in the to-be-identified image can comprise the following steps: firstly, based on the coordinates of the human body skeleton joints of the detected target in each image in the sequence frames, the coordinates sequence of the key points of the detected target (human or object) in the sequence frames is determined; the mean square error of the coordinates sequence of the key points of the detected target (human or object) in the sequence frames and each corresponding point on the preset target motion curve is calculated; and based on the mean square error, the fitting degree of the motion curve of the detected target and the preset target motion curve is determined. Specifically, the greater the mean square error, the smaller the fitting degree of the motion curve of the detected target and the preset target motion curve; and the smaller the mean square error, the greater the fitting degree of the motion curve of the detected target and the preset target motion curve. The fitting degree of the motion curve of the detected target and the preset target motion curve can also be determined based on the relationship between the mean square error and a preset mean square error threshold value. When the mean square error is greater than or equal to the preset mean square error threshold value, it is determined that the motion of the detected target does not belong to the preset target motion; and when the mean square error is less than the preset mean square error threshold value, it is determined that the motion of the detected target belongs to the preset target motion. The above-mentioned curve can be a parabola. If the target motion is punching, the above-mentioned curve can be a parabola corresponding to the punching motion. It should be noted that whether the motion trajectory of the detected target is an abnormal curve can also be determined according to the coordinates of the human body skeleton joints of the detected target in each image in the sequence frames, that is, the curvature of the motion curve of the detected target is determined according to the coordinates of the human body skeleton joints of the detected target in each image in the sequence frames. When the curvature of the motion curve is greater than a preset threshold value, it is determined that the motion curve of the detected target is an abnormal parabola, and the motion of the detected target can be a target motion such as punching.
[0098] Further, the number of joint acceleration mutations can be the cumulative number of significant and discontinuous step changes in joint angular acceleration (or linear acceleration) within a short period of time within a continuous period of time (a preset time length within the acquisition time length of the sequence frames, for example, a time length of 100 ms). The determination of the number of joint acceleration mutations can be based on the coordinates of the human body skeleton joints of the detected target in each image in the sequence frames to determine the speed change rate (acceleration) of the target joint within the preset time length. When the acceleration is greater than a preset mutation threshold value, it is recorded as one mutation, the number of mutations within the preset time is counted, and the number of joint acceleration mutations is obtained. The mutation threshold value can be a preset value that can measure whether the acceleration of the joint has mutated. When the acceleration is greater than the mutation threshold value, it is determined that a mutation occurs at this time, and when the acceleration is less than or equal to the mutation threshold value, it is determined that the acceleration is normal. The mutation threshold value can be specifically set according to specific situations, which is not limited in the present embodiment. For example, the mutation threshold value can be set to .
[0099] In addition, the trunk inclination angle variance can be a statistical quantity for measuring the fluctuation of the trunk inclination angle in a continuous time period, and can be used to quantify the body stability or the motion intensity. The determination of the trunk inclination angle variance can be based on the coordinates of the detected human skeleton joints in each image in the sequence frames, determining the position changes of the scapula and hip joint coordinates of the detected target in the sequence frames, determining the trunk axis offset of the detected target during the sequence frames based on the position changes of the scapula and hip joint coordinates of the detected target in the sequence frames, and determining the trunk inclination angle variance of the detected target in the sequence frames (or a predetermined time length) in the to-be-identified image based on the trunk axis offset of the detected target during the sequence frames. If the trunk inclination angle variance of the detected target in the to-be-identified image in the sequence frames (or a predetermined time length) is greater than a predetermined threshold, it is determined that the detected target is out of balance and needs to trigger the imbalance warning.
[0100] In this step, the determination of the confidence of the target motion of the detected target in the to-be-identified image based on the fitting degree of the preset curve, the number of joint acceleration mutations, and the trunk inclination angle variance can be that the fitting degree of the preset curve, the number of joint acceleration mutations, and the trunk inclination angle variance are set with corresponding weights according to specific application scenarios, and the confidence S of the target motion of the detected target in the to-be-identified image is determined based on the weights of the fitting degree of the preset curve, the number of joint acceleration mutations, and the trunk inclination angle variance, and the values of the fitting degree of the preset curve, the number of joint acceleration mutations, and the trunk inclination angle variance. The specific calculation process is as follows:
[0101] ;
[0102] In the above, A is the weight of the fitting degree of the preset curve, is the fitting degree of the preset curve of the detected target, B is the weight of the number of joint acceleration mutations, is the number of joint acceleration mutations of the detected target, C is the weight of the trunk inclination angle variance, is the trunk inclination angle variance.
[0103] It should be noted that the weights of the fitting degree of the preset curve, the number of joint acceleration mutations, and the trunk inclination angle variance can be set according to specific conditions, and the present embodiment does not make specific limitations here. For example, A can be set to 0.2, B can be set to 0.3, and C can be set to 0.5.
[0104] The steps S11 to S12 above obtain the coordinates of the human skeleton joints of the detection target in each image in the sequence frame by performing human key point detection on the detection target in the sequence frame corresponding to the to-be-identified image, and further obtain the confidence of the target action of the detection target in the to-be-identified image. Through determination of the confidence of the target action of the detection target in the to-be-identified image, the confidence of the target behavior of the detection target in the to-be-identified image is determined through the confidence of the target action of the detection target.
[0105] In one embodiment, step S2223, based on the audio data corresponding to the to-be-identified image, determining the sound abnormality degree of the audio data corresponding to the to-be-identified image, comprising:
[0106] Step S21, based on the audio data corresponding to the to-be-identified image, determining the sound feature of the audio data.
[0107] The above determining the sound feature of the audio data based on the audio data corresponding to the to-be-identified image can be extracting a 13-dimensional feature vector of the audio data using the Mel frequency cepstrum coefficient on the audio data.
[0108] Step S22, based on the sound feature of the audio data, determining the sound abnormality degree of the audio data.
[0109] The above determining the sound abnormality degree of the audio data based on the sound feature of the audio data can be analyzing the energy ratio of the target frequency band in the sound feature of the audio data based on the sound feature of the audio data, and if the energy ratio of the target frequency band exceeds a preset energy ratio threshold of the total energy, calculating the sound abnormality degree of the audio data by a Gaussian mixture model. The above target frequency band can be specifically set based on specific situations, which is not specifically limited in the embodiment, for example, the above target frequency band can be The above energy ratio threshold can be specifically set based on specific situations, which is not specifically limited in the embodiment, for example, the above energy ratio threshold can be 35%. The above Gaussian mixture model can be a Gaussian mixture model trained by a large number of normal audio, which can be used to judge the sound abnormality degree of the audio data.
[0110] The above steps S21 to S22 determine the sound abnormality degree of the audio data corresponding to the to-be-identified image based on the audio data corresponding to the to-be-identified image, which facilitates subsequent determination of the confidence of the target behavior of the detection target in the to-be-identified image based on the sound abnormality degree of the audio data.
[0111] In addition, in one embodiment, step S2224, based on the temperature field data corresponding to the to-be-identified image, determining the thermal intensity of the limb contact area of the detection target in the to-be-identified image, comprising:
[0112] Step S31, extract gradient features of the temperature field data corresponding to the to-be-identified image.
[0113] The above-mentioned extraction of the gradient features of the temperature field data corresponding to the to-be-identified image can be the extraction of the gradient features of the temperature field data by using a 5x5 convolution kernel.
[0114] Step S32, determine the thermal intensity of the temperature field data corresponding to the to-be-identified image based on the gradient features.
[0115] In this step, the above-mentioned determination of the thermal intensity of the temperature field data corresponding to the to-be-identified image based on the gradient features can be the generation of the thermal intensity of the temperature field data corresponding to the to-be-identified image by using a preset convolutional neural network based on the gradient features. The above-mentioned preset convolutional neural network can be a network that can be used to generate the thermal intensity of the temperature field data and is pre-trained.
[0116] Step S33, determine the thermal intensity of the limb contact area of the detection target in the to-be-identified image based on the thermal intensity of the temperature field data corresponding to the to-be-identified image.
[0117] The above-mentioned determination of the thermal intensity of the limb contact area of the detection target in the to-be-identified image based on the thermal intensity of the temperature field data corresponding to the to-be-identified image can be the determination of the contact area based on the thermal intensity of the temperature field data corresponding to the to-be-identified image and a preset contact area temperature mutation threshold, and the determination of the thermal intensity of the limb contact area of the detection target in the to-be-identified image based on the determined contact area. The above-mentioned preset contact area temperature mutation threshold can be specifically set based on specific situations, which is not specifically limited in this embodiment. For example, the above-mentioned preset contact area temperature mutation threshold can be that is, when the temperature mutation is greater than or equal to the preset contact area temperature mutation threshold, the area is determined as a contact area.
[0118] The above-mentioned steps S31 to S33 determine the thermal intensity of the limb contact area of the detection target in the to-be-identified image based on the temperature field data corresponding to the to-be-identified image, which facilitates the subsequent determination of the confidence of the target behavior of the detection target in the to-be-identified image based on the thermal intensity of the limb contact area of the detection target in the to-be-identified image.
[0119] Further, in one embodiment, step S2225, based on the frame sequence of the to-be-identified image, determines the entropy value corresponding to the behavior of the detection target of the to-be-identified image, the contact force field matching degree between the action of the detection target and the preset target action, and the panic index of the non-detection target in the to-be-identified image, comprising:
[0120] Step S41, based on the frame sequence of the to-be-identified image, modeling the behavior of the detected target in the to-be-identified image in the frame sequence to obtain a state transition matrix of the detected target in the to-be-identified image in the frame sequence.
[0121] The state transition matrix can be an N*N probability table in a hidden Markov model, used to represent the probability of jumping to a hidden state j at the next time if the system is currently in a hidden state i. The above modeling the behavior of the detected target in the to-be-identified image in the frame sequence to obtain a state transition matrix of the detected target in the to-be-identified image in the frame sequence can be modeling the behavior of the detected target in the image frame sequence by using a hidden Markov model to obtain a state transition matrix of the detected target in the to-be-identified image in the frame sequence. It should be noted that a preset number of images in the sequence frame can be selected, and the selected images can be modeled by using a hidden Markov model to obtain a state transition matrix of the detected target in the to-be-identified image in the frame sequence. The above preset number can be specifically set according to specific requirements, which is not specifically limited in the embodiment. For example, the above preset number can be 10.
[0122] Step S42, based on the state transition matrix, determining the entropy value corresponding to the behavior of the detected target in the to-be-identified image.
[0123] The above determining the entropy value corresponding to the behavior of the detected target in the to-be-identified image based on the state transition matrix can be calculating the KL divergence (Kullback-Leibler divergence) of the state transition matrix and the normal behavior model based on the state transition matrix, and determining the entropy value corresponding to the behavior of the detected target in the to-be-identified image based on the KL divergence of the state transition matrix and the normal behavior model. The above determining the entropy value corresponding to the behavior of the detected target in the to-be-identified image based on the KL divergence of the state transition matrix and the normal behavior model can be when the KL divergence of the state transition matrix and the normal behavior model is greater than a preset divergence threshold, converting the difference greater than the preset divergence threshold into the entropy value corresponding to the behavior of the detected target in the to-be-identified image. The above preset divergence threshold can be specifically set according to specific requirements, which is not specifically limited in the embodiment. For example, the preset divergence threshold can be set to 1.2. When the KL divergence of the state transition matrix and the normal behavior model is greater than the preset divergence threshold, it is determined that the behavior of the detected target at this time is an abnormal behavior, and the entropy value is used to represent the abnormal degree of this abnormal behavior.
[0124] The calculation process of converting the difference greater than the preset divergence threshold into the entropy value E corresponding to the behavior of the detected target in the to-be-identified image is as follows:
[0125] ;
[0126] Where D is the KL divergence between the state transition matrix and the normal behavior model, and μ is the mean KL divergence of the normal behavior model. The mean KL divergence of the normal behavior model can be obtained empirically.
[0127] Step S43: Determine the spatiotemporal characteristics of the joint pressure distribution of the target in the image to be identified.
[0128] The aforementioned determination of the spatiotemporal characteristics of the joint pressure distribution of the target in the image to be identified can be achieved by using a pre-set model.
[0129] Step S44: Based on the spatiotemporal characteristics of the joint pressure distribution of the detected target in the image to be identified, and the spatiotemporal characteristics of the joint pressure distribution corresponding to the preset target action template, determine the contact force field matching degree between the action of the detected target in the image to be identified and the preset target action.
[0130] In this step, the determination of the contact force field matching degree between the action of the detected target in the image to be identified and the preset target action, based on the spatiotemporal features of the joint pressure distribution of the detected target in the image to be identified and the spatiotemporal features of the joint pressure distribution corresponding to the preset target action template, can be achieved by using an improved dynamic time warping algorithm to match the spatiotemporal features of the joint pressure distribution of the detected target in the image to be identified with the spatiotemporal features of the joint pressure distribution corresponding to the preset target action template, thereby obtaining the contact force field matching degree between the action of the detected target in the image to be identified and the preset target action. The preset target action can be a standard fighting action (e.g., pushing, punching, kicking, etc.).
[0131] The formula for calculating the contact force field matching degree M between the action of the detected target in the image to be identified and the preset target action is as follows:
[0132] ;
[0133] Among them, the above The spatiotemporal characteristics of joint pressure distribution corresponding to the preset target motion template. The spatiotemporal characteristics of joint pressure distribution of the target in the image to be identified.
[0134] When M is greater than the preset contact threshold C1, contact verification is triggered. The preset contact threshold can be set according to specific needs, and this embodiment does not impose a specific limitation. For example, the preset contact threshold can be 0.5.
[0135] Step S45: Based on the facial expression features of non-detected targets in the image to be identified; the facial expression features include pupil dilation rate and limb withdrawal speed; calculate the panic index of non-detected targets in the image to be identified.
[0136] The aforementioned non-detected targets can be any targets in the image to be identified, excluding the detected targets. Specifically, they can be people in the image to be identified, excluding the detected targets. The calculation of the panic index of the non-detected targets in the image to be identified, based on their facial expression features, can be achieved by using a Vision Transformer to extract the facial expression features (pupil dilation rate and limb withdrawal speed) of the non-detected targets. The panic index D of the non-detected targets in the image to be identified is then calculated by weighting these facial expression features and their weights (the weights of pupil dilation rate and limb withdrawal speed).
[0137] When pupil dilation rate When the limb withdrawal speed is >0.8 m / s, the D value exceeds the warning threshold of 0.7. At this point, a fight or other targeted behavior may have occurred. The weights of the pupil dilation rate and the limb withdrawal speed mentioned above can be set according to specific circumstances, and this embodiment does not impose specific limitations. For example, the weight of the pupil dilation rate can be 0.4, and the weight of the limb withdrawal speed can be 0.6.
[0138] Steps S41 to S45 above determine the entropy value corresponding to the behavior of the detected target in the image to be identified, the contact force field matching degree between the action of the detected target and the preset target action, and the panic index of the non-detected target in the image to be identified, based on the frame sequence of the image to be identified. By determining the entropy value corresponding to the behavior of the detected target in the image to be identified, the contact force field matching degree between the action of the detected target and the preset target action, and the panic index of the non-detected target in the image to be identified, it is convenient to subsequently determine the confidence level of the target behavior of the detected target in the image to be identified.
[0139] In one embodiment, step S224, based on the confidence level of the target action of the detected target in the image to be identified, the sound anomaly of the audio data, the thermal intensity of the limb contact area of the detected target, the entropy value corresponding to the behavior of the detected target, the matching degree of the contact force field between the action of the detected target and the preset target action, and the panic index of the non-detected target, determines the confidence level of the target behavior of the detected target in the image to be identified, including:
[0140] Step S2242: Based on the confidence level of the target action of the detected target in the image to be identified, the panic index of the non-detected target, the thermal intensity of the limb contact area of the detected target, and the sound abnormality of the audio data, determine whether the detected target in the image to be identified meets the preset first judgment condition; the preset first judgment condition is that the action meets the preset first action preset condition and the sound meets the preset sound preset condition.
[0141] The preset first action preset condition can be specifically set according to specific situations, and embodiments are not specifically limited herein.
[0142] For example, the preset first action preset condition can be that:
[0143] ;
[0144] S is a confidence degree of a target action of a detection target in the to-be-identified image, and D is a KL divergence of a state transition matrix and a normal behavior model.
[0145] The preset sound preset condition can be specifically set according to specific situations, and embodiments are not specifically limited herein. D is a panic index of a non-detection target, which is used for environmental interference compensation. When the panic degree of the surrounding crowd is higher, the triggering threshold of S is lower.
[0146] For example, the preset sound preset condition can be that:
[0147] ;
[0148] A1 is a sound abnormality degree of audio data, and H is a heat intensity of a limb contact area of the detection target.
[0149] It should be noted that when the thermal imaging detection device detects limb contact, the weight of the sound abnormality degree of the audio data is increased. When the detection target in the to-be-identified image satisfies the preset first judgment condition, the behavior of the detection target can be a target behavior or a high-energy limb interaction (such as violent movement). When the detection target in the to-be-identified image does not satisfy the preset first judgment condition, the behavior of the detection target is a normal behavior (such as walking, talking, etc.).
[0150] In step S2244, when the detection target in the to-be-identified image satisfies the preset first judgment condition, the detection target is judged whether to satisfy a preset second judgment condition based on an entropy value corresponding to the behavior of the detection target in the to-be-identified image, a contact force field matching degree of the action of the detection target and a preset target action, and a panic index of a non-detection target. The preset second judgment condition is that the action satisfies at least one of a preset second action preset condition or a preset collision condition.
[0151] The preset second action preset condition can be specifically set according to specific situations, and embodiments are not specifically limited herein.
[0152] For example, the preset second action preset condition can be that:
[0153] ;
[0154] Wherein, C1 is a preset contact threshold, and E is an entropy value corresponding to the behavior of the detected target in the image to be recognized.
[0155] The above-mentioned preset collision condition can be specifically set according to specific situations, and the embodiment is not specifically limited here.
[0156] For example, the above-mentioned preset collision condition can be:
[0157] ;
[0158] Wherein, C1 is a preset contact threshold, and D is the KL divergence of the state transition matrix and the normal behavior model.
[0159] When the detected target in the image to be recognized meets the preset second judgment condition, it can be determined that the current detected target is highly suspected of target behavior, and a conflict or an accidental violent collision may be occurring. If the preset second judgment condition is not met, it is determined as non-target behavior.
[0160] Step S2246, when the detected target in the image to be recognized meets the preset second judgment condition, the confidence of the target action of the detected target in the image to be recognized, the entropy value corresponding to the behavior of the detected target, the panic index of the non-detected target, and the contact force field matching degree of the action of the detected target and the preset target action are determined. The confidence of the target behavior of the detected target in the image to be recognized.
[0161] The above-mentioned confidence of the target action of the detected target in the image to be recognized, the entropy value corresponding to the behavior of the detected target, the panic index of the non-detected target, and the contact force field matching degree of the action of the detected target and the preset target action are determined. The calculation process of the confidence P of the target behavior of the detected target in the image to be recognized is:
[0162] P=0.4S'+0.3E'+0.2D'+0.1C';
[0163] Wherein, S' is a false alarm prevention mechanism of the cross thermal environment, which prevents false alarms caused by too fast movement trajectory in part of the limbless collision such as dance and basketball, S'=S÷(1+0.2|1-H|), through cross verification of thermal imaging and action mode, when H value is high (with contact), S' is amplified, and the weight of the confidence of the target action of the detected target is enhanced; It is a false alarm prevention mechanism of dynamic trajectory and thermal feeling, which prevents false alarms caused by behaviors such as hugging, helping and other contact but low thermal induction, non-violent trajectory, ;D' is a false alarm prevention mechanism of panic degree and movement trajectory cross verification, which prevents false alarms caused by violent behavior, E' is a multi-dimensional confidence mechanism, which introduces a hybrid behavior entropy value of force field matching degree and acoustic effect, and is used to reflect multiple dimensions affecting the confidence, .
[0164] The weights of the above parameters can be set according to the actual scene, which is not specifically limited in the embodiment.
[0165] The steps S2242 to S2246 determine the confidence of the target behavior of the detection target in the to-be-identified image based on the confidence of the target action of the detection target in the to-be-identified image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target. Through the determination of the confidence of the target behavior of the detection target in the to-be-identified image, it is convenient to determine the detection result of the target behavior in the to-be-identified image according to the confidence of the target behavior of the detection target in the to-be-identified image.
[0166] The embodiment will be described and illustrated below through preferred embodiments.
[0167] Figure 3 is a flowchart of a target behavior detection method provided by a preferred embodiment of the application. As shown in Figure 3 , the target behavior detection method comprises the following steps:
[0168] Step S301, acquiring a to-be-identified image, a sequence frame corresponding to the to-be-identified image, audio data corresponding to the to-be-identified image, and temperature field data corresponding to the to-be-identified image; the sequence frame corresponding to the to-be-identified image is a preset number of images adjacent to the to-be-identified image, which constitutes a time sequence frame sequence together with the to-be-identified image;
[0169] Step S302, determining a detection target in the to-be-identified image based on the acquired to-be-identified image;
[0170] Step S303, performing human key point detection on the detection target in the sequence frame corresponding to the to-be-identified image to obtain a confidence of a target action of the detection target in the to-be-identified image;
[0171] Step S304, determining a sound abnormality degree of the audio data corresponding to the to-be-identified image based on the audio data corresponding to the to-be-identified image;
[0172] Step S305, determining a thermal intensity of a limb contact area of the detection target in the to-be-identified image based on the temperature field data corresponding to the to-be-identified image;
[0173] Step S306, based on the frame sequence of the to-be-recognized image, determine the entropy value corresponding to the behavior of the detection target in the to-be-recognized image, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target in the to-be-recognized image.
[0174] Step S307, based on the confidence of the target action of the detection target in the to-be-recognized image, the panic index of the non-detection target, the thermal intensity of the limb contact area of the detection target, and the sound abnormality degree of the audio data, determine whether the detection target in the to-be-recognized image satisfies the preset first judgment condition; the preset first judgment condition is that the action satisfies the preset first action preset condition, and the sound satisfies the preset sound preset condition.
[0175] Step S308, when the detection target in the to-be-recognized image satisfies the preset first judgment condition, based on the entropy value corresponding to the behavior of the detection target in the to-be-recognized image, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target, determine whether the detection target in the to-be-recognized image satisfies the preset second judgment condition; the preset second judgment condition is that the action satisfies at least one of the preset second action preset condition or the preset collision condition.
[0176] Step S309, when the detection target in the to-be-recognized image satisfies the preset second judgment condition, based on the confidence of the target action of the detection target in the to-be-recognized image, the entropy value corresponding to the behavior of the detection target, the panic index of the non-detection target, and the contact force field matching degree of the action of the detection target and the preset target action, determine the confidence of the target behavior of the detection target in the to-be-recognized image.
[0177] Step S310, based on the confidence of the target behavior of the detection target in the to-be-recognized image, determine the detection result of the target behavior in the to-be-recognized image.
[0178] The steps S301 to S310 described above determine the confidence of the target behavior of the detection target in the to-be-identified image by acquiring the to-be-identified image, the sequence frame corresponding to the to-be-identified image, the audio data corresponding to the to-be-identified image, and the temperature field data corresponding to the to-be-identified image, and then determining the detection result of the target behavior in the to-be-identified image based on the confidence of the target behavior of the detection target in the to-be-identified image. The confidence of the target behavior of the detection target in the to-be-identified image is determined by combining the to-be-identified image, the sequence frame corresponding to the to-be-identified image, the audio data, and the temperature field data, so that the detection result of the target behavior is more accurate, the missed detection caused by single data detection is avoided, and when the target behavior is detected, the to-be-detected data does not need to be transmitted to the cloud, thereby solving the problems of the existing target behavior detection method, such as high missed detection rate, low target recognition accuracy, and security risks of data leakage and privacy leakage.
[0179] It should be understood that, although each step in the flowchart involved in each embodiment as described above is shown in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps.
[0180] Based on the same inventive concept, in the present embodiment, a target behavior detection device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and will not be described again. The terms "module", "unit", "sub-unit" and the like used below can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware or a combination of software and hardware is also possible and contemplated.
[0181] In one embodiment, Figure 4 is a structural block diagram of a target behavior detection device provided by an embodiment of the present application, as Figure 4 shown, the target behavior detection device comprises:
[0182] The data acquisition module 42 is configured to acquire the to-be-identified image, sequence frames corresponding to the to-be-identified image, audio data corresponding to the to-be-identified image, and temperature field data corresponding to the to-be-identified image. The sequence frames corresponding to the to-be-identified image are a preset number of image frames adjacent to the to-be-identified image, and the to-be-identified image together form a time sequence frame sequence.
[0183] The confidence determination module 44 is configured to determine the confidence of the target behavior of the detection target in the to-be-identified image based on the acquired to-be-identified image, sequence frames corresponding to the to-be-identified image, audio data corresponding to the to-be-identified image, and temperature field data corresponding to the to-be-identified image.
[0184] The detection module 46 is configured to determine the detection result of the target behavior in the to-be-identified image based on the confidence of the target behavior of the detection target in the to-be-identified image.
[0185] The target behavior detection device described above determines the confidence of the target behavior of the detection target in the to-be-identified image by acquiring the to-be-identified image, sequence frames corresponding to the to-be-identified image, audio data corresponding to the to-be-identified image, and temperature field data corresponding to the to-be-identified image, and then determines the detection result of the target behavior in the to-be-identified image based on the confidence of the target behavior of the detection target in the to-be-identified image. The confidence of the target behavior of the detection target in the to-be-identified image is determined by combining the to-be-identified image, sequence frames corresponding to the to-be-identified image, audio data, and temperature field data, so that the detection result of the target behavior is more accurate, and the problem of high omission rate, low target recognition accuracy, and security risks such as data leakage and privacy leakage in the existing target behavior detection method is solved.
[0186] It should be noted that each of the above modules can be a functional module or a program module, and can be implemented by software or hardware. For the modules implemented by hardware, each of the above modules can be located in the same processor; or each of the above modules can be located in different processors in any combination.
[0187] In one embodiment, a target behavior detection system is provided, which includes a data acquisition module, a feature extraction module, a classification decision module, and an optimization deployment module.
[0188] The data acquisition module comprises an infrared imaging sensing unit, a skeleton key point trajectory acquisition unit and a high-frequency voiceprint acquisition unit; the infrared imaging sensing unit is configured to acquire the instantaneous temperature change of the limb contact area in the to-be-identified image, and obtain the temperature field data corresponding to the to-be-identified image; the skeleton key point trajectory acquisition unit is configured to perform human key point detection on the detection target in the sequence frames corresponding to the to-be-identified image, and obtain the coordinates of the human skeleton joints of the detection target in each image in the sequence frames; and the high-frequency voiceprint acquisition unit is configured to acquire the audio data corresponding to the to-be-identified image.
[0189] The skeleton key point trajectory acquisition unit can further comprise a plurality of cameras configured to acquire the coordinate information of the human skeleton key points in real time. The high-frequency voiceprint acquisition unit is configured to identify and acquire sound features such as high-frequency screaming, heavy object impact sound and glass breaking sound.
[0190] The feature extraction module is configured to extract the gradient feature of the temperature field data corresponding to the to-be-identified image; determine the thermal intensity of the temperature field data corresponding to the to-be-identified image based on the gradient feature of the temperature field data corresponding to the to-be-identified image; determine the fitting degree of the preset curve, the number of joint acceleration mutations and the trunk inclination angle variance of the detection target in the to-be-identified image based on the coordinates of the human skeleton joints of the detection target in each image in the sequence frames; the fitting degree of the preset curve is the fitting degree of the action curve of the detection target and the curve of the preset target action; determine the confidence of the target action of the detection target in the to-be-identified image based on the fitting degree of the preset curve, the number of joint acceleration mutations and the trunk inclination angle variance of the detection target in the to-be-identified image; determine the sound feature of the audio data based on the audio data corresponding to the to-be-identified image; and determine the sound abnormality degree of the audio data based on the sound feature of the audio data.
[0191] The feature extraction module is configured to pre-process and standardize the collected data. In the feature extraction module, the system pre-processes and standardizes the collected data, and extracts the key features of the target behaviors such as fighting and brawling. The feature extraction module particularly adopts a multi-head self-attention mechanism and a Transformer network, realizes cross-modal correlation analysis of multi-modal data, and enhances the feature representation capability.
[0192] In the feature extraction module, the system uses a weighted fusion algorithm to convert the skeleton trajectory features into action confidence S, wherein the fitting degree of the preset curve can be assigned a weight of 0.2 to reflect the action normativity, the number of joint acceleration mutations can be assigned a weight of 0.3 to capture the violent features, and the trunk inclination angle variance can be assigned a weight of 0.5 to quantify the degree of body imbalance. The voiceprint feature extracts the abnormal frequency band energy ratio through the mel frequency cepstral coefficient to construct the sound abnormality degree of the audio data; the thermal imaging data uses a convolutional neural network to extract the temperature gradient features of the contact area to generate the thermal intensity H of the temperature field data.
[0193] The feature extraction module is further configured to calculate the KL divergence of the state transition matrix of the detection target in the sequence frame by using a hidden Markov model, to obtain an entropy value corresponding to the behavior of the detection target; determine the joint node pressure distribution spatiotemporal feature of the detection target in the to-be-identified image by using a contact force field model; determine the action of the detection target in the to-be-identified image based on the joint node pressure distribution spatiotemporal feature of the detection target in the to-be-identified image; determine the contact force field matching degree between the action of the detection target in the to-be-identified image and the preset target action based on the action of the detection target in the to-be-identified image and the preset target action template; extract the expression feature of the face of a person other than the detection target in the to-be-identified image by using a micro-expression extraction model; the expression feature includes the pupil dilation rate and the limb retraction speed; and determine the panic index of the person other than the detection target in the to-be-identified image based on the expression feature of the face of the person other than the detection target in the to-be-identified image.
[0194] The module uses a dynamic threshold classifier to generate a multi-scene noise data training model based on adversarial learning, and maps the original features to a space suitable for classification through a generative adversarial network. The classifier can dynamically adjust the detection sensitivity according to environmental parameters and crowd density, set appropriate threshold rules for different scenes, and effectively reduce the false positive rate.
[0195] The classification decision module is used for determining whether the detected target in the to-be-identified image meets a preset first judgment condition based on a confidence degree of a target action of the detected target in the to-be-identified image, a panic index of a non-detected target, a thermal intensity of a limb contact area of the detected target, and a sound abnormality degree of audio data. The preset first judgment condition is that the action meets a preset first action preset condition and the sound meets a preset sound preset condition. When the detected target in the to-be-identified image meets the preset first judgment condition, it is judged whether the detected target in the to-be-identified image meets a preset second judgment condition based on an entropy value corresponding to a behavior of the detected target in the to-be-identified image, a contact force field matching degree of the action of the detected target and a preset target action, and the panic index of the non-detected target in the to-be-identified image. The preset second judgment condition is that the action meets at least one of a preset second action preset condition or a preset collision condition. When the detected target in the to-be-identified image meets the preset second judgment condition, a confidence degree of a target behavior of the detected target in the to-be-identified image is determined based on the confidence degree of the target action of the detected target in the to-be-identified image, the entropy value corresponding to the behavior of the detected target, the panic index of the non-detected target, and the contact force field matching degree of the action of the detected target and the preset target action.
[0196] The classification decision module is the core of the system, which introduces a multi-dimensional dynamic decision mechanism. First, the system constructs a basic feature layer (feature extraction module) to quantify the skeleton trajectory, voiceprint feature and thermal imaging data into action confidence, acoustic abnormality and thermal intensity, respectively. The calculation of action confidence S integrates the parabolic fitting degree of punching trajectory, the number of joint acceleration mutations and the variance of trunk inclination angle to form a composite index. This module also introduces an advanced feature layer, which calculates the KL divergence E of the behavior pattern deviating from the normal state by analyzing the Markov transition probability matrix of the action sequence, and outputs the contact matching degree C1 by constructing a real-time contact force field model. In addition, the system also integrates a scene semantic analysis module to identify the panic index D of the surrounding crowd (non-detected target).
[0197] The final decision adopts a three-level fusion strategy, considering multiple indicators of the basic feature layer and the advanced feature layer. The primary threshold screening requires the action confidence and acoustic abnormality to reach a certain threshold; the intermediate verification requires the behavior entropy evaluation or the contact matching degree to reach a certain standard; and the advanced confirmation corrects the final confidence by considering environmental factors. Experiments show that this scheme significantly improves the recognition accuracy of fighting behavior in a specific scene and greatly reduces the false positive rate.
[0198] The decision mechanism adopts a three-level cascade verification architecture: the primary screening sets dynamic thresholds (environmental interference compensation), (thermal enhancement); the intermediate verification introduces composite conditions ; the confidence of the advanced decision wherein, The modal cross-validation mechanism is implemented. Experimental data show that the mechanism still maintains an accuracy of 89.7% when the crowd density is greater than 3 people / ㎡, which is 23.6% higher than that of the traditional method. The system integrates an online adversarial training module. When the manual judgment is inconsistent with the system for three consecutive times, the adversarial sample generator automatically updates the feature library, and the false positive rate decreases by 8.2% per month.
[0199] The optimization deployment layer adopts an end-to-end lightweight scheme. The fighting features of a large 3D CNN teacher model are migrated to a lightweight student model through hierarchical knowledge distillation. The layer also includes a model compression and acceleration module. An 8-bit integer quantization scheme is customized for edge devices to significantly improve the inference speed. The system is based on a domestic Hisilicon chip to build a hundred-yuan-level hardware. Only the coordinates of the skeleton points and the voiceprint spectrum graph are transmitted. The data security and privacy protection are realized through ISO 31700-2024 privacy certification.
[0200] The optimization deployment module is configured to determine a detection result of a target behavior of a detection target in the to-be-recognized image based on a confidence of the target behavior of the detection target in the to-be-recognized image, generate an adversarial sample of the target behavior based on the detection result of the target behavior of the detection target in the to-be-recognized image and the manual detection result, and update a model in the feature extraction module by using the adversarial sample of the target behavior.
[0201] The optimization deployment module adopts an end-to-end lightweight scheme. The fighting features of a large 3D CNN (Three-Dimensional Convolutional Neural Network) teacher model are migrated to a lightweight student model through hierarchical knowledge distillation. The layer also includes a model compression and acceleration module. An 8-bit integer quantization scheme is customized for edge devices to significantly improve the inference speed. The system is based on a traditional chip to build a hundred-yuan-level hardware. Only the coordinates of the skeleton points and the voiceprint spectrum graph are transmitted. The data security and privacy protection are realized through ISO 31700-2024 privacy certification.
[0202] In addition, the system also designs a dynamic learning mechanism. When a high confidence is detected for multiple times but the manual review is a false positive, the adversarial sample generator automatically updates the feature library, further improving the adaptive ability and robustness of the system. This multi-modal edge intelligent detection system not only realizes the accurate and real-time detection of fighting and fighting behavior in complex scenes, but also effectively reduces the missed detection rate and false positive rate, providing strong support for safety monitoring in public places.
[0203] In one embodiment, the target behavior detection system provided by the embodiments of the present application can be applied in a school security monitoring scene, taking a teaching building corridor as a specific monitoring area. After accessing various sensors, the system collects multi-modal data and performs analysis and processing, thereby realizing accurate and real-time detection of fighting behavior.
[0204] In one example, an infrared thermal imaging sensor, a skeleton joint trajectory acquisition module, and a high-frequency voiceprint acquisition module can be deployed at the corner and intersection of the teaching building corridor, and integrated and installed with the existing monitoring stand on the hardware. On the software level, the system is connected to the school video monitoring platform through the interface to realize protocol docking. The monitoring platform collects the camera stream media of the corridor area in real time (resolution 1080P, frame rate 25fps), and transmits the key frame pictures (extract one frame every 200ms) in the video stream to the edge computing device through the RTSP protocol, and synchronously outputs the audio stream to the high-frequency voiceprint acquisition module, thereby forming a real-time access link of multi-modal data.
[0205] Further, after receiving the key frame pictures transmitted by the video monitoring platform, the infrared thermal imaging sensor synchronously collects the temperature field data of the same scene (resolution 320x240), the skeleton joint trajectory acquisition module detects the human key points of the pictures, and real-time extracts the two-dimensional coordinates of 17 joints (error ≤2 pixels). The collected picture data is subjected to gray scale normalization and noise reduction processing by the preprocessing module, and the temperature field data and joint coordinate information are packaged into multi-modal data frames (format is JSON, including timestamp, coordinate matrix, and temperature matrix). The multi-modal data frames are transmitted to the feature extraction layer through the gigabit network port of the edge node in the UDP (User Datagram Protocol) protocol, and the transmission delay is controlled within 50ms, thereby ensuring the real-time performance of the data.
[0206] The feature extraction module is deployed in the high-performance computing unit of the edge computing device, and adopts a heterogeneous computing architecture to realize parallel processing of multi-modal data. After receiving the multi-modal data frames transmitted by the data acquisition layer, the basic feature extraction layer performs skeleton feature extraction, voiceprint feature extraction, and thermal imaging feature extraction. The advanced feature extraction layer constructs a space-time dynamic analysis model based on the basic features, and the specific processing logic is as follows: behavior entropy E calculation, contact force field matching degree C1 calculation, and panic index D calculation.
[0207] In addition, the system will automatically adjust the detection sensitivity and resource allocation according to the environmental characteristics and personnel activity rules at different times: during the daytime class period (8:00-12:00, 14:00-18:00), each sensor maintains the standard collection frequency, the feature extraction layer preferentially processes high-frequency voiceprint and skeleton track features, and at the same time, the primary threshold of action confidence S is set to 0.7, balancing detection accuracy and power consumption; during the class break period (10 minutes), the crowd mobility is enhanced, the system automatically reduces the sensor collection frequency by 20%, and raises the S threshold to 0.85, while reducing the weight of the trunk tilt angle variance to reduce false positives caused by normal crowd pushing, and at this time the edge device power resource priority ensures multi-target detection; during the night period (22:00 to 6:00 the next day), the infrared thermal imaging sensor switches to night vision mode, the voiceprint module starts the noise reduction algorithm, and at the same time, the warning threshold of thermal intensity H is reduced from 0.6 to 0.45, the contact detection sensitivity in dim environments is improved, and at the same time, the model power consumption is reduced by 40% through dynamic quantization technology, achieving a balance between energy saving and high sensitivity.
[0208] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0209] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0210] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method of target behavior detection, the method comprising: The method comprises: acquiring a to-be-recognized image, sequence frames corresponding to the to-be-recognized image, audio data corresponding to the to-be-recognized image, and temperature field data corresponding to the to-be-recognized image; the sequence frames corresponding to the to-be-recognized image are a preset number of images adjacent to the to-be-recognized image, which together constitute a time sequence frame sequence with the to-be-recognized image; based on the to-be-recognized image, the sequence frames corresponding to the to-be-recognized image, the audio data corresponding to the to-be-recognized image, and the temperature field data corresponding to the to-be-recognized image, determining a confidence degree of a target action of a detection target in the to-be-recognized image, a sound abnormality degree of the audio data, a thermal intensity of a limb contact area of the detection target, an entropy value corresponding to a behavior of the detection target, a contact force field matching degree of an action of the detection target and a preset target action, and a panic index of a non-detection target; based on the confidence degree of the target action of the detection target in the to-be-recognized image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target, determining a confidence degree of a target behavior of the detection target in the to-be-recognized image; based on the confidence degree of the target behavior of the detection target in the to-be-recognized image, determining a detection result of the target behavior in the to-be-recognized image.
2. The target behavior detection method of claim 1, wherein, The method comprises: based on the to-be-recognized image, the sequence frames corresponding to the to-be-recognized image, the audio data corresponding to the to-be-recognized image, and the temperature field data corresponding to the to-be-recognized image, determining a confidence degree of a target action of a detection target in the to-be-recognized image, a sound abnormality degree of the audio data, a thermal intensity of a limb contact area of the detection target, an entropy value corresponding to a behavior of the detection target, a contact force field matching degree of an action of the detection target and a preset target action, and a panic index of a non-detection target; based on the confidence degree of the target action of the detection target in the to-be-recognized image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target, determining a confidence degree of a target behavior of the detection target in the to-be-recognized image; based on the to-be-recognized image, the sequence frames corresponding to the to-be-recognized image, the audio data corresponding to the to-be-recognized image, and the temperature field data corresponding to the to-be-recognized image, determining a confidence degree of a target action of a detection target in the to-be-recognized image, a sound abnormality degree of the audio data, a thermal intensity of a limb contact area of the detection target, an entropy value corresponding to a behavior of the detection target, a contact force field matching degree of an action of the detection target and a preset target action, and a panic index of a non-detection target; based on the confidence degree of the target action of the detection target in the to-be-recognized image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target, determining a confidence degree of a target behavior of the detection target in the to-be-recognized image; based on the to-be-recognized image, the sequence frames corresponding to the to-be-recognized image, the audio data corresponding to the to-be-recognized image, and the temperature field data corresponding to the to-be-recognized image, determining a confidence degree of a target action of a detection target in the to-be-recognized image, a sound abnormality degree of the audio data, a thermal intensity of a limb contact area of the detection target, an entropy value corresponding to a behavior of the detection target, a contact force field matching degree of an action of the detection target and a preset target action, and a panic index of a non-detection target; based on the confidence degree of the target action of the detection target in the to-be-recognized image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target, determining a confidence degree of a target behavior of the detection target in the to-be-recognized image; the human body key point detection on the detection target in the sequence frames corresponding to the to-be-recognized image to obtain the confidence degree of the target action of the detection target in the to-be-recognized image, comprises: 3. The target behavior detection method of claim 2, wherein, perform human key point detection on the detection target in each of the sequence frames corresponding to the to-be-identified image, to obtain coordinates of human skeleton joints of the detection target in each of the sequence frames; determine a fitting degree of a preset curve of the detection target in the to-be-identified image, a number of joint acceleration mutations, and a trunk inclination angle variance based on the coordinates of the human skeleton joints of the detection target in each of the sequence frames; the fitting degree of the preset curve is a fitting degree of a motion curve of the detection target and a curve of a preset target motion; determine a confidence degree of a target motion of the detection target in the to-be-identified image based on the fitting degree of the preset curve of the detection target in the to-be-identified image, the number of joint acceleration mutations, and the trunk inclination angle variance.
4. The target behavior detection method of claim 2, wherein, The determining, based on the audio data corresponding to the to-be-identified image, of the sound abnormality degree of the audio data corresponding to the to-be-identified image includes: determining a sound feature of the audio data based on the audio data corresponding to the to-be-identified image; determining the sound abnormality degree of the audio data based on the sound feature of the audio data.
5. The method of claim 2, wherein, The determining, based on the temperature field data corresponding to the to-be-identified image, of the thermal intensity of the limb contact area of the detection target in the to-be-identified image includes: extracting gradient features of the temperature field data corresponding to the to-be-identified image; determining the thermal intensity of the temperature field data corresponding to the to-be-identified image based on the gradient features; determining the thermal intensity of the limb contact area of the detection target in the to-be-identified image based on the thermal intensity of the temperature field data corresponding to the to-be-identified image.
6. The method of claim 2, wherein, The determining, based on the frame sequence of the to-be-identified image, of an entropy value corresponding to a behavior of the detection target in the to-be-identified image, a contact force field matching degree of a motion of the detection target and a preset target motion, and a panic index of a non-detection target in the to-be-identified image includes: modeling a behavior of the detection target in the to-be-identified image in the frame sequence based on the frame sequence of the to-be-identified image, to obtain a state transition matrix of the detection target in the to-be-identified image in the frame sequence; determining the entropy value corresponding to the behavior of the detection target in the to-be-identified image based on the state transition matrix; determining a joint node pressure distribution spatiotemporal feature of the detection target in the to-be-identified image; determining the contact force field matching degree of the motion of the detection target and the preset target motion based on the joint node pressure distribution spatiotemporal feature of the detection target in the to-be-identified image and a joint node pressure distribution spatiotemporal feature corresponding to a preset target motion template; calculating the panic index of the non-detection target in the to-be-identified image based on an expression feature of a face of the non-detection target in the to-be-identified image; the expression feature includes a pupil dilation rate and a limb withdrawal speed.
7. The target behavior detection method of claim 1, wherein, The confidence degree of the target action of the detection target in the to-be-identified image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target are used to determine the confidence degree of the target behavior of the detection target in the to-be-identified image, and the confidence degree of the target behavior of the detection target in the to-be-identified image is determined by the following steps: The confidence degree of the target action of the detection target in the to-be-identified image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target are used to determine the confidence degree of the target behavior of the detection target in the to-be-identified image, and the confidence degree of the target behavior of the detection target in the to-be-identified image is determined by the following steps: When the detection target in the to-be-identified image meets the preset first judgment condition, the entropy value corresponding to the behavior of the detection target in the to-be-identified image, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target are used to determine whether the detection target in the to-be-identified image meets the preset second judgment condition; the preset second judgment condition is that at least one of the preset second action preset condition or the preset collision condition is met. When the detection target in the to-be-identified image meets the preset second judgment condition, the confidence degree of the target action of the detection target in the to-be-identified image, the entropy value corresponding to the behavior of the detection target, the panic index of the non-detection target, and the contact force field matching degree of the action of the detection target and the preset target action are used to determine the confidence degree of the target behavior of the detection target in the to-be-identified image.
8. A target behavior detection apparatus characterized by comprising: The device comprises: A data acquisition module is configured to acquire a to-be-identified image, a sequence frame corresponding to the to-be-identified image, audio data corresponding to the to-be-identified image, and temperature field data corresponding to the to-be-identified image; the sequence frame corresponding to the to-be-identified image is a preset number of image frames adjacent to the to-be-identified image, and the to-be-identified image forms a time sequence frame sequence together; The confidence determination module is configured to determine a confidence of a target action of a detection target in the to-be-identified image, a sound abnormality degree of the audio data, a thermal intensity of a limb contact area of the detection target, an entropy value corresponding to a behavior of the detection target, a contact force field matching degree of an action of the detection target and a preset target action, and a panic index of a non-detection target based on the to-be-identified image, the sequence frame corresponding to the to-be-identified image, the audio data corresponding to the to-be-identified image, and the temperature field data corresponding to the to-be-identified image; and determine a confidence of a target behavior of the detection target in the to-be-identified image based on the confidence of the target action of the detection target in the to-be-identified image, the sound abnormality degree of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree of the action of the detection target and the preset target action, and the panic index of the non-detection target. The detection module is configured to determine a detection result of the target behavior in the to-be-identified image based on the confidence of the target behavior of the detection target in the to-be-identified image.
9. A target behavior detection system characterized by, The system comprises a data acquisition module, a feature extraction module, a classification decision module, and an optimization deployment module. The data acquisition module comprises an infrared imaging sensing unit, a skeletal key point trajectory acquisition unit, and a high-frequency voiceprint acquisition unit; the infrared imaging sensing unit is configured to acquire instantaneous temperature changes of a limb contact area in a to-be-identified image to obtain temperature field data corresponding to the to-be-identified image; the skeletal key point trajectory acquisition unit is configured to perform human key point detection on a detection target in a sequence frame corresponding to the to-be-identified image to obtain coordinates of human skeletal joint points of the detection target in each image in the sequence frame; and the high-frequency voiceprint acquisition unit is configured to acquire audio data corresponding to the to-be-identified image. The feature extraction module is configured to extract gradient features of the temperature field data corresponding to the to-be-identified image; determine a thermal intensity of the temperature field data corresponding to the to-be-identified image based on the gradient features of the temperature field data corresponding to the to-be-identified image; determine a fitting degree of a preset curve of a detection target in the to-be-identified image, a joint acceleration mutation number, and a trunk inclination angle variance based on coordinates of human skeletal joint points of the detection target in each image in the sequence frame; the fitting degree of the preset curve is a fitting degree of an action curve of the detection target and a curve of a preset target action; determine a confidence of a target action of the detection target in the to-be-identified image based on the fitting degree of the preset curve of the detection target in the to-be-identified image, the joint acceleration mutation number, and the trunk inclination angle variance; determine a sound feature of the audio data based on the audio data corresponding to the to-be-identified image; and determine a sound abnormality degree of the audio data based on the sound feature of the audio data. The feature extraction module is further configured to calculate a KL divergence of a state transition matrix of the detected target in a sequence frame by using a hidden Markov model to obtain an entropy value corresponding to a behavior of the detected target; determine a joint node pressure distribution space-time feature of the detected target in the to-be-identified image by using a contact force field model; determine a contact force field matching degree of a motion of the detected target in the to-be-identified image and a preset target motion based on the joint node pressure distribution space-time feature of the detected target in the to-be-identified image and a joint node pressure distribution space-time feature corresponding to the preset target motion template; extract an expression feature of a face of a person other than the detected target in the to-be-identified image by using a micro-expression extraction model; the expression feature includes a pupil dilation rate and a body withdrawal speed; and determine a panic index of the person other than the detected target in the to-be-identified image based on the expression feature of the face of the person other than the detected target in the to-be-identified image. The classification decision module is configured to determine whether the detected target in the to-be-identified image meets a preset first judgment condition based on the confidence of the target motion of the detected target in the to-be-identified image, the panic index of the person other than the detected target, the thermal intensity of the body contact area of the detected target, and the sound abnormality degree of the audio data; the preset first judgment condition is that the motion meets a preset first motion preset condition and the sound meets a preset sound preset condition; when the detected target in the to-be-identified image meets the preset first judgment condition, determine whether the detected target in the to-be-identified image meets a preset second judgment condition based on the entropy value corresponding to the behavior of the detected target in the to-be-identified image, the contact force field matching degree of the motion of the detected target and the preset target motion, and the panic index of the person other than the detected target in the to-be-identified image; the preset second judgment condition is that the motion meets at least one of a preset second motion preset condition or a preset collision condition; and when the detected target in the to-be-identified image meets the preset second judgment condition, determine the confidence of the target behavior of the detected target in the to-be-identified image based on the confidence of the target motion of the detected target in the to-be-identified image, the entropy value corresponding to the behavior of the detected target, the panic index of the person other than the detected target, and the contact force field matching degree of the motion of the detected target and the preset target motion. The optimization deployment module is configured to determine a detection result of the target behavior of the detected target in the to-be-identified image based on the confidence of the target behavior of the detected target in the to-be-identified image; generate an adversarial sample of the target behavior based on the detection result of the target behavior of the detected target in the to-be-identified image and an artificial detection result; and update the model in the feature extraction module by using the adversarial sample of the target behavior.
Citation Information
Patent Citations
Behavior detection method, device and system, computer equipment and storage medium
CN119577651A