Target behavior detection method and device and target behavior detection system
By combining multiple feature extraction and pattern matching technologies of image, audio and temperature field data, the existing target behavior detection methods are solved in harsh scenarios, achieving high-accuracy local detection, reducing the risk of data leakage.
Patent Information
- Application Number
- CN202511054092.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-07-30
AI Technical Summary
The existing target behavior detection methods have high missed detection rates and high misjudgment rates in harsh scenarios such as occlusion, backlighting or crowd-intensive scenarios, and there is a security risk of data leakage and privacy leakage.
By combining image, audio and temperature field data, a variety of data feature extraction and pattern matching techniques are used to determine the confidence of target behavior, avoid missed detection caused by single data detection, and perform local analysis to reduce data transmission.
It improves the accuracy of target behavior detection, reduces the missed detection rate, reduces the risk of data leakage, and avoids the complexity and cost of cloud computing.
Smart Images

Figure CN120564271A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of security monitoring technology, and in particular to a target behavior detection method, device, and target behavior detection system. Background Art
[0002] With the development of science and technology and the surge in public safety needs, public security monitoring systems are playing an increasingly important role in maintaining social order and protecting personal safety. Among them, the use of security monitoring systems to detect target behavior is becoming increasingly important.
[0003] However, existing monitoring systems rely on single-camera visual analysis when detecting target behaviors, resulting in a high rate of missed detection of target behaviors in harsh scenarios such as occlusion, backlighting, or dense crowds. In addition, existing target behavior detection methods often use fixed-threshold rule engines, which result in some misleading actions being misjudged as target behaviors. For example, certain basketball actions may be misjudged as fighting, affecting the accuracy of target behavior recognition. In addition, existing target behavior detection methods often require centralized data analysis and calculations in the cloud, requiring the transmission of large amounts of high-definition video streams, which not only increases the complexity and computing cost of the system, but also increases the security risks of data leakage and privacy leakage.
[0004] There is no effective solution to the problems of high missed detection rate, low target recognition accuracy, and security risks of data leakage and privacy leakage in existing target behavior detection methods. Summary of the Invention
[0005] Based on this, it is necessary to provide a target behavior detection method, device and a target behavior detection system to address the above technical problems.
[0006] In a first aspect, the present application provides a method for detecting target behavior. The method comprises:
[0007] Acquire an image to be recognized, a sequence of frames corresponding to the image to be recognized, audio data corresponding to the image to be recognized, and temperature field data corresponding to the image to be recognized; the sequence of frames corresponding to the image to be recognized is a time-sequential frame sequence formed by a preset number of images adjacent to the image to be recognized and the image to be recognized;
[0008] Determining a confidence level of a target behavior of a detection target in the image to be identified based on the acquired image to be identified, a sequence of frames corresponding to the image to be identified, audio data corresponding to the image to be identified, and temperature field data corresponding to the image to be identified;
[0009] A detection result of the target behavior in the image to be recognized is determined based on the confidence level of the target behavior of the detected target in the image to be recognized.
[0010] In one embodiment, determining the confidence level of the target behavior of the detection target in the image to be identified based on the acquired image to be identified, the sequence of frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified includes:
[0011] Based on the acquired image to be identified, the sequence of frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified, determine the confidence level of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the detection target action and the preset target action, and the panic index of the non-detection target;
[0012] The confidence level of the target action of the detection target in the image to be identified is determined based on the confidence level of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the detection target's action and the preset target action, and the panic index of the non-detection target.
[0013] In one embodiment, the method of determining the confidence level of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the detection target action and the preset target action, and the panic index of the non-detection target based on the acquired image to be identified, the sequence frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified includes:
[0014] Determining a detection target in the image to be identified based on the acquired image to be identified;
[0015] Performing human key point detection on the detection target in the sequence frames corresponding to the image to be identified, and obtaining a confidence level of a target action of the detection target in the image to be identified;
[0016] determining, based on the audio data corresponding to the image to be identified, the degree of sound abnormality of the audio data corresponding to the image to be identified;
[0017] Determining the thermal intensity of the limb contact area of the detection target in the image to be identified based on the temperature field data corresponding to the image to be identified;
[0018] Based on the frame sequence of the image to be identified, the entropy value corresponding to the behavior of the detection target in the image to be identified, the contact force field matching degree between the detection target's action and the preset target action, and the panic index of the non-detection target in the image to be identified are determined.
[0019] In one embodiment, the performing of human key point detection on the detection target in the sequence frames corresponding to the image to be identified to obtain the confidence level of the target action of the detection target in the image to be identified includes:
[0020] Performing human body key point detection on the detection target in the sequence frames corresponding to the image to be identified, respectively, to obtain the coordinates of the human skeleton joint points of the detection target in each image in the sequence frames;
[0021] Based on the coordinates of the human skeletal joint points of the detection target in each image in the sequence frame, the fitting degree of the preset curve, the number of joint acceleration mutations, and the variance of the trunk inclination angle of the detection target in the image to be identified are determined; the fitting degree of the preset curve is the degree of fitting between the action curve of the detection target and the curve of the preset target action; based on the fitting degree of the preset curve of the detection target in the image to be identified, the number of joint acceleration mutations, and the variance of the trunk inclination angle, the confidence level of the target action of the detection target in the image to be identified is determined.
[0022] In one embodiment, determining the sound abnormality of the audio data corresponding to the image to be identified based on the audio data corresponding to the image to be identified includes:
[0023] Determining sound features of the audio data based on the audio data corresponding to the image to be recognized;
[0024] The sound abnormality degree of the audio data is determined based on the sound feature of the audio data.
[0025] In one embodiment, determining the thermal intensity of the limb contact area of the detection target in the image to be identified based on the temperature field data corresponding to the image to be identified includes:
[0026] Extracting gradient features of temperature field data corresponding to the image to be identified;
[0027] Determining the thermal intensity of the temperature field data corresponding to the image to be identified based on the gradient feature;
[0028] The thermal intensity of the limb contact area of the detection target in the image to be identified is determined based on the thermal intensity of the temperature field data corresponding to the image to be identified.
[0029] In one embodiment, determining, based on the frame sequence of the image to be identified, the entropy value corresponding to the behavior of the detection target in the image to be identified, the degree of contact force field matching between the detection target's action and the preset target action, and the panic index of the non-detection target in the image to be identified includes:
[0030] Based on the frame sequence of the image to be identified, modeling the behavior of the detection target in the image to be identified in the image frame sequence, and obtaining a state transition matrix of the memory target in the image to be identified in the frame sequence;
[0031] Determining, based on the state transition matrix, an entropy value corresponding to a behavior of the detection target in the image to be identified;
[0032] Determining the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified;
[0033] Determining a degree of contact force field matching between a movement of the detection target in the image to be identified and a preset target movement based on the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified and the spatiotemporal characteristics of the joint pressure distribution corresponding to a preset target movement template;
[0034] The panic index of the non-detection target in the image to be identified is calculated based on the facial expression features of the non-detection target in the image to be identified; the facial expression features include pupil dilation rate and limb withdrawal speed.
[0035] In one embodiment, determining the confidence level of the target behavior of the detection target in the image to be identified based on the confidence level of the target behavior of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the body contact area of the detection target, the entropy value corresponding to the detection target's behavior, the contact force field matching degree between the detection target's movement and a preset target movement, and the panic index of the non-detection target includes:
[0036] Determining whether the detection target in the image to be identified meets a preset first judgment condition based on the confidence level of the target action of the detection target in the image to be identified, the panic index of the non-detection target, the thermal intensity of the body contact area of the detection target, and the sound abnormality of the audio data; the preset first judgment condition is that the action meets a preset first action preset condition and the sound meets a preset sound preset condition;
[0037] When the detection target in the image to be identified meets a preset first judgment condition, based on the entropy value corresponding to the behavior of the detection target in the image to be identified, the contact force field matching degree between the movement of the detection target and the preset target movement, and the panic index of the non-detection target, it is determined whether the detection target in the image to be identified meets a preset second judgment condition; the preset second judgment condition is that the movement meets at least one of a preset second movement preset condition or a preset collision condition;
[0038] When the detection target in the image to be identified meets the preset second judgment condition, the confidence of the target behavior of the detection target in the image to be identified is determined based on the confidence of the target action of the detection target in the image to be identified, the entropy value corresponding to the behavior of the detection target, the panic index of the non-detection target, and the contact force field matching degree between the action of the detection target and the preset target action.
[0039] In a second aspect, the present application also provides a target behavior detection device. The device includes:
[0040] a data acquisition module, configured to acquire an image to be recognized, a sequence of frames corresponding to the image to be recognized, audio data corresponding to the image to be recognized, and temperature field data corresponding to the image to be recognized; the sequence of frames corresponding to the image to be recognized is a time-sequential frame sequence formed by a preset number of image frames adjacent to the image to be recognized and the image to be recognized;
[0041] a confidence determination module, configured to determine a confidence level of a target behavior of a detection target in the image to be identified based on the acquired image to be identified, a sequence of frames corresponding to the image to be identified, audio data corresponding to the image to be identified, and temperature field data corresponding to the image to be identified;
[0042] and a detection module for determining a detection result of the target behavior in the image to be identified based on the confidence level of the target behavior of the detected target in the image to be identified.
[0043] In a third aspect, the present application further provides a target behavior detection system, the system comprising: a data acquisition module, a feature extraction module, a classification decision module, and an optimization deployment module;
[0044] The data acquisition module includes an infrared imaging sensor unit, a skeleton key point trajectory acquisition unit, and a high-frequency voiceprint acquisition unit; the infrared imaging sensor unit is used to acquire instantaneous temperature changes in the limb contact area in the image to be identified, and obtain temperature field data corresponding to the image to be identified; the skeleton key point trajectory acquisition unit is used to perform human key point detection on the detection target in the sequence frame corresponding to the image to be identified, and obtain the coordinates of the human skeleton joint points of the detection target in each image in the sequence frame; the high-frequency voiceprint acquisition unit is used to acquire audio data corresponding to the image to be identified;
[0045] The feature extraction module is used to extract the gradient features of the temperature field data corresponding to the image to be identified; based on the gradient features of the temperature field data corresponding to the image to be identified, determine the thermal intensity of the temperature field data corresponding to the image to be identified; based on the coordinates of the human skeletal joint points of the detection target in each image in the sequence frame, determine the fitting degree of the preset curve, the number of joint acceleration mutations, and the variance of the trunk inclination angle of the detection target in the image to be identified; the fitting degree of the preset curve is the degree of fitting between the action curve of the detection target and the curve of the preset target action; based on the fitting degree of the preset curve of the detection target in the image to be identified, the number of joint acceleration mutations, and the variance of the trunk inclination angle, determine the confidence of the target action of the detection target in the image to be identified; based on the audio data corresponding to the image to be identified, determine the sound features of the audio data; based on the sound features of the audio data, determine the sound abnormality of the audio data;
[0046] The feature extraction module is further used to calculate the KL divergence of the state transition matrix of the detection target in the sequence frame through the hidden Markov model to obtain the entropy value corresponding to the behavior of the detection target; use the contact force field model to determine the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified; based on the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified and the spatiotemporal characteristics of the joint pressure distribution corresponding to the preset target action template, determine the contact force field matching degree between the action of the detection target in the image to be identified and the preset target action; use the micro-expression extraction model to extract the expression features of the faces of people other than the non-detection targets in the image to be identified; the expression features include pupil dilation rate and limb withdrawal speed; based on the expression features of the faces of the non-detection targets in the image to be identified, determine the panic index of the non-detection targets in the image to be identified;
[0047] The classification decision module is used to determine whether the detection target in the image to be identified meets a preset first judgment condition based on the confidence of the target action of the detection target in the image to be identified, the panic index of the non-detection target, the thermal intensity of the limb contact area of the detection target, and the sound abnormality of the audio data; the preset first judgment condition is that the action meets the preset first action preset condition and the sound meets the preset sound preset condition; when the detection target in the image to be identified meets the preset first judgment condition, based on the entropy value corresponding to the behavior of the detection target in the image to be identified, the action of the detection target is matched with the contact force field of the preset target action. degree, and the panic index of the non-detected target in the image to be identified, to determine whether the detected target in the image to be identified meets a preset second judgment condition; the preset second judgment condition is that the action meets at least one of a preset second action preset condition or a preset collision condition; when the detected target in the image to be identified meets the preset second judgment condition, determining the confidence of the target behavior of the detected target in the image to be identified based on the confidence of the target action of the detected target in the image to be identified, the entropy value corresponding to the behavior of the detected target, the panic index of the non-detected target, and the contact force field matching degree between the action of the detected target and the preset target action;
[0048] An optimization deployment module is used to determine the detection result of the target behavior of the detection target in the image to be identified based on the confidence of the target behavior of the detection target in the image to be identified; generate an adversarial sample of the target behavior based on the detection result of the target behavior of the detection target in the image to be identified and the manual detection result; and use the adversarial sample of the target behavior to update the model in the feature extraction module.
[0049] The above-mentioned target behavior detection method, device and a target behavior detection system determine the confidence of the target behavior of the detection target in the image to be identified by obtaining the image to be identified, the sequence frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified. Then, based on the confidence of the target behavior of the detection target in the image to be identified, the detection result of the target behavior in the image to be identified is determined. It determines the confidence of the target behavior of the detection target in the image to be identified by combining the image to be identified, the sequence frames corresponding to the image to be identified, the audio data and the temperature field data, and multiple data. This makes the detection result of the target behavior more accurate, avoids missed detection caused by single data detection, and when performing target behavior detection, there is no need to transmit the data to be detected to the cloud, which solves the problems of existing target behavior detection methods, such as high missed detection rate, low target recognition accuracy and security risks of data leakage and privacy leakage.
[0050] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0052] Figure 1 A hardware structure block diagram of a terminal for a target behavior detection method provided in one embodiment of the present application;
[0053] Figure 2 A flowchart of a target behavior detection method provided in one embodiment of the present application;
[0054] Figure 3 A flowchart of a target behavior detection method provided in a preferred embodiment of the present application;
[0055] Figure 4 This is a structural block diagram of a target behavior detection device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0056] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0057] Unless otherwise defined, technical or scientific terms used in this application shall have the ordinary meanings as understood by persons of ordinary skill in the art to which this application belongs. The terms "a," "an," "the," "these," and similar expressions in this application do not denote limitations on quantity and may be singular or plural. The terms "comprise," "include," "have," and any variations thereof, as used in this application, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include unlisted steps or modules (units) or other steps or modules (units) inherent to the process, method, product, or device. The terms "connected," "connected," "coupled," and similar expressions used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. As used in this application, "plurality" means two or more. "And / or" describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone; A and B exist simultaneously; or B exists alone. Generally, the character " / " indicates that the objects in the preceding and following relationship are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0058] The method embodiment provided in this embodiment can be executed in a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1 FIG is a block diagram of the hardware structure of the terminal of the target behavior detection method of this embodiment. Figure 1 As shown, the terminal may include one or more ( Figure 1 The processor 102 (only one is shown) and a memory 104 for storing data, wherein the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA. The terminal may also include a transmission device 106 for communication functions and an input / output device 108. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0059] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the target behavior detection method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0060] Transmission device 106 is used to receive or transmit data via a network. This network may include a wireless network provided by the terminal's communications provider. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0061] In this embodiment, a target behavior detection method is provided. Figure 2 is a flow chart of the target behavior detection method of this embodiment. Figure 2 As shown, the process includes the following steps:
[0062] Step S210, obtaining the image to be recognized, the sequence frames corresponding to the image to be recognized, the audio data corresponding to the image to be recognized, and the temperature field data corresponding to the image to be recognized; the sequence frames corresponding to the image to be recognized are a time-series frame sequence composed of a preset number of images adjacent to the image to be recognized and the image to be recognized.
[0063] Acquiring the image to be identified may involve capturing image frames of the target area using an image acquisition device. The image acquisition device may be a visible light imaging device (such as a SLR or micro-single-lens camera), an infrared or thermal imaging device (such as an infrared camera), a hyperspectral camera, or the like. Acquiring the image to be identified may also involve real-time acquisition of the streaming media of the target area captured by the camera (according to the set parameters of 1080P resolution and 25fps frame rate) to obtain key frame images from the video stream. Specifically, key frame images from the video stream may be extracted at preset time intervals. The preset time interval may be set based on specific needs and is not specifically limited in this embodiment. For example, the preset time interval may be 200ms.
[0064] Furthermore, the target area can be specifically defined based on the specific scenario. For example, it can be an area such as a hallway or classroom in a school setting. The preset number of images can be set based on specific needs and is not specifically limited in this embodiment. As long as the preset number of images adjacent to the image to be identified, together with the image to be identified, can form a sequence of frames that can determine the confidence level of the target behavior of the detected object in the image to be identified, it is sufficient. Acquiring audio data corresponding to the image to be identified can involve using a high-frequency voiceprint acquisition module to capture the audio stream from the streaming media captured by the camera, or using a separately installed audio acquisition device to capture the audio data corresponding to the image to be identified. The audio acquisition device can be a MEMS (Micro-Electro-Mechanical Systems) ultrasonic microphone. It should be noted that the audio acquisition device can also be other microphones or other audio acquisition devices, as long as they can capture the audio data corresponding to the image to be identified. This is not specifically limited in this embodiment. Acquiring temperature field data corresponding to the image to be identified can involve using an infrared thermal imaging sensor to acquire temperature field data corresponding to the image to be identified. The temperature field data can be the complete distribution of temperature values at each point in space at the time corresponding to the image to be recognized. The resolution of the temperature field data can be set based on specific needs. For example, the resolution of the temperature field data can be set to 320×240. The infrared thermal imaging sensor can be one or more of an uncooled infrared microbolometer, a cooled infrared focal plane array, an infrared thermal imager module, and an infrared thermal imaging chip.
[0065] Step S220 , based on the acquired image to be identified, the sequence frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified, the confidence level of the target behavior of the detection target in the image to be identified is determined.
[0066] In this step, the confidence level of the target behavior of the detection target can be used to characterize the degree of confidence that the target behavior is the target behavior. The higher the confidence level, the greater the likelihood that the target behavior is the target behavior. Conversely, the lower the confidence level, the less likely the target behavior is the target behavior. The target behavior can be specifically set based on specific needs or application scenarios and is not specifically limited in this embodiment. For example, the target behavior is fighting. The above-mentioned determination of the confidence of the target behavior of the detection target in the image to be identified based on the acquired image to be identified, the sequence frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified can be based on the acquired image to be identified, the sequence frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified, to determine the confidence of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the action of the detection target and the preset target action, and the panic index of the non-detected target; and further, based on the confidence of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the action of the detection target and the preset target action, and the panic index of the non-detected target.
[0067] Step S230 : determining a detection result of the target behavior in the image to be identified based on the confidence level of the target behavior of the detected target in the image to be identified.
[0068] The above detection results include the presence of target behavior, the absence of target behavior, and suspicious behavior where it is uncertain whether the target behavior exists. The above detection result of the target behavior in the image to be identified is determined based on the confidence level of the target behavior of the detected target in the image to be identified. When the confidence level of the target behavior of the detected target in the image to be identified is greater than a preset first confidence threshold, it is determined that the target behavior exists in the image to be identified; when the confidence level of the target behavior of the detected target in the image to be identified is less than a preset second confidence threshold, it is determined that the target behavior does not exist in the image to be identified; when the confidence level of the target behavior of the detected target in the image to be identified is between the preset second confidence threshold and the preset first confidence threshold, it is determined that the behavior is suspicious and requires manual review. The above preset second confidence threshold is less than the preset first confidence threshold. The above preset second confidence threshold and the preset first confidence threshold can be specifically set according to the specific situation and are not specifically limited in this embodiment. For example, the above preset second confidence threshold is 0.6, and the above preset first confidence threshold is 0.8.
[0069] It should be noted that after determining the detection results of the target behavior in the image to be identified, if the system outputs a high-confidence alarm three times in a row but is manually reviewed as a false alarm, the multimodal features of the misjudged samples will be automatically collected, and the detection threshold and feature weight will be adjusted after comparative analysis. The misjudged features will be added to the non-target behavior template library, and by generating adversarial network enhanced samples and performing incremental training on edge devices, the feature library will be officially updated and synchronized to all nodes after backtesting verifies that the false alarm rate has been reduced by more than 80%.
[0070] The above steps S210 to S230 determine the confidence of the target behavior of the detection target in the image to be identified by obtaining the image to be identified, the sequence frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified. Then, based on the confidence of the target behavior of the detection target in the image to be identified, the detection result of the target behavior in the image to be identified is determined. The confidence of the target behavior of the detection target in the image to be identified is determined by combining the image to be identified, the sequence frames corresponding to the image to be identified, the audio data and the temperature field data, and multiple data. This makes the detection result of the target behavior more accurate, avoids missed detection caused by single data detection, and when performing target behavior detection, there is no need to transmit the data to be detected to the cloud, which solves the problems of existing target behavior detection methods, such as high missed detection rate, low target recognition accuracy, and security risks of data leakage and privacy leakage.
[0071] In one embodiment, step S220, based on the acquired image to be identified, the sequence of frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified, determines the confidence level of the target behavior of the detection target in the image to be identified, including:
[0072] Step S222, based on the acquired image to be identified, the sequence frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified, determine the confidence of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the detection target action and the preset target action, and the panic index of the non-detection target.
[0073] The confidence level of the above-mentioned target action can be used to characterize the degree of credibility of the action of the detection target being the target action. The higher the confidence level, the greater the possibility that the action of the detection target is the target action, and conversely, the lower the confidence level, the smaller the possibility that the action of the detection target is the target action. The above-mentioned target action can be specifically set according to specific needs or application scenarios, and this embodiment does not make specific limitations here. For example, the target action is a punching action in a fight. The above-mentioned confidence level of the target action of the detection target in the image to be identified can be based on the acquired image to be identified, determining the detection target in the image to be identified, and then performing human key point detection on the detection target in the sequence frame corresponding to the image to be identified to obtain the confidence level of the target action of the detection target in the image to be identified.
[0074] The above-mentioned sound abnormality degree can be used to characterize the degree of deviation of the audio data from the standard audio data of the current scene. The above-mentioned standard audio data of the current scene can refer to a pre-stored acoustic model corresponding to the current scene. The above-mentioned acoustic model can be a sound model of normal audio in the current scene. The above-mentioned sound abnormality degree can be a scalar value, specifically a value in the range of [0, 1]. The larger the value, the higher the probability of the sound abnormality, and the smaller the value, the lower the probability of the sound abnormality, indicating that the sound in the audio data is more normal.
[0075] Furthermore, the thermal intensity of the limb contact area may be the intensity of a thermal signal in the limb contact area. The intensity of the thermal signal may be the instantaneous thermal power or instantaneous temperature increment per unit area. Specifically, it may be calculated by the product of the temperature rise value of the limb contact area and the contact heat transfer coefficient. The temperature rise value of the limb contact area may be the average temperature rise value or the maximum temperature rise value of the limb contact area. Specifically, the thermal intensity of the limb contact area may be a value in the range of [0, 1].
[0076] The entropy value corresponding to the aforementioned behavior can represent the predictability or regularity of the target's current behavioral pattern, that is, it is used to characterize the severity of the abnormality of the target's current behavioral pattern. A larger entropy value corresponding to the behavior indicates that the current action sequence deviates more from the "normal habitual pattern" and is more unpredictable and abnormal. A smaller entropy value corresponding to the behavior indicates that the action sequence is highly consistent with normal patterns and the behavior is stable and predictable. The entropy value corresponding to the aforementioned behavior can specifically be a value in the range [0, 1].
[0077] The contact force field matching degree can be used to measure the spatiotemporal similarity between the joint-pressure distribution sequence of the target's movement, currently measured in real time, and the contact force template of a preset target movement (e.g., fighting, falling, or carrying). The range of the contact force field matching degree can be set based on specific needs and is not specifically limited in this embodiment. For example, the contact force field matching degree can be set to a value between 0 and 1. When the contact force field matching degree between the target's movement and the preset target movement is closer to 1, the more similar the joint-pressure distribution sequence of the target's movement and the contact force template of the preset target movement are in spatiotemporal morphology, and the closer the target's movement is to the preset target movement. When the contact force field matching degree between the target's movement and the preset target movement is closer to 0, the less similar the joint-pressure distribution sequence of the target's movement and the contact force template of the preset target movement are in spatiotemporal morphology, and the less similar the target's movement is to the preset target movement.
[0078] In this step, the panic index of the non-detected target can be the panic index of onlookers or other people in the environment (other than the detected target), and is used to indicate the degree of panic or fright experienced by the crowd in the current scene due to the sudden event. A larger value for the panic index of the non-detected target indicates more pronounced panic in the crowd and a stronger reaction to the abnormal event. A smaller value for the panic index of the non-detected target indicates more stable emotions in the crowd. The panic index of the non-detected target can be used to determine whether the behavior of the detected target is the target behavior.
[0079] Step S224, based on the confidence of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the detection target's action and the preset target action, and the panic index of the non-detection target, the confidence of the target behavior of the detection target in the image to be identified is determined.
[0080] The above steps S222 to S224 determine the confidence of the target behavior of the detection target in the image to be identified based on the acquired image to be identified, the sequence frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified. By determining the confidence of the target behavior of the detection target in the image to be identified, it is convenient to determine the detection result of the target behavior in the image to be identified according to the confidence of the target behavior of the detection target in the image to be identified.
[0081] Specifically, in one embodiment, step S222 determines, based on the acquired image to be identified, the sequence of frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified, the confidence level of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the detection target action and the preset target action, and the panic index of the non-detection target, including:
[0082] Step S2221: Determine the detection target in the image to be identified based on the acquired image to be identified.
[0083] In this step, the above-mentioned detection target may be a target that needs to be subjected to target behavior detection in the current scenario. For example, the above-mentioned detection target may be a person or animal that needs to be subjected to target behavior detection. The above-mentioned determination of the detection target in the image to be identified based on the acquired image to be identified may be to detect the detection target in the image to be identified using a preset detection method. The above-mentioned preset detection method may be specifically set according to specific needs, and this embodiment does not make any specific limitations here. For example, when performing target behavior detection on a person, a person whose distance is less than a preset distance threshold may be used as a detection target. At this time, the preset detection method may be to use a trained neural network model to identify the image to be identified, and to determine a person whose distance is less than a preset distance threshold as a detection target in the drawing to be identified.
[0084] Step S2222 , performing human body key point detection on the detection target in the sequence frames corresponding to the image to be recognized, and obtaining the confidence level of the target action of the detection target in the image to be recognized.
[0085] The aforementioned method of performing human key point detection on the detection target in the sequence of frames corresponding to the image to be identified and obtaining the confidence level of the target action of the detection target in the image to be identified may involve performing human key point detection on the detection target in the sequence of frames corresponding to the image to be identified, obtaining the coordinates of the detection target's skeletal joints in each image of the sequence, and then determining the fit of a preset curve, the number of joint acceleration mutations, and the variance of the trunk tilt angle of the detection target in the image to be identified based on the coordinates of the detection target's skeletal joints in each image of the sequence. The fit of the preset curve is the degree of fit between the detection target's action curve and the preset target action curve. The confidence level of the target action of the detection target in the image to be identified is determined based on the fit of the preset curve, the number of joint acceleration mutations, and the variance of the trunk tilt angle of the detection target in the image to be identified. The aforementioned detection target action curve may be an action curve determined based on the changes in the position of one or more skeletal joints of the detection target in each image of the sequence. The aforementioned preset target action may be a target action corresponding to the target behavior. For example, when the target behavior is fighting, the target action may be a punch, a slap, or the like. The curve of the preset target action may be a position change curve of one or more skeletal joints corresponding to the preset target action. For example, when the preset target action is punching, the curve of the preset target action may be a position change curve of the skeletal joints of the hand during the punching action.
[0086] Step S2223 : determining the sound abnormality degree of the audio data corresponding to the image to be identified based on the audio data corresponding to the image to be identified.
[0087] The above-mentioned determination of the sound abnormality of the audio data corresponding to the image to be identified based on the audio data corresponding to the image to be identified may be based on the audio data corresponding to the image to be identified, determining the sound features of the audio data, and further, determining the sound abnormality of the audio data based on the sound features of the audio data.
[0088] Step S2224 , determining the thermal intensity of the limb contact area of the detection target in the image to be identified based on the temperature field data corresponding to the image to be identified.
[0089] Among them, the above-mentioned determination of the thermal intensity of the limb contact area of the detection target in the image to be identified based on the temperature field data corresponding to the image to be identified can be to extract the gradient characteristics of the temperature field data corresponding to the image to be identified, and then, based on the gradient characteristics, determine the thermal intensity of the temperature field data corresponding to the image to be identified.
[0090] Step S2225, based on the frame sequence of the image to be identified, determine the entropy value corresponding to the behavior of the detection target in the image to be identified, the contact force field matching degree between the detection target's action and the preset target action, and the panic index of the non-detection target in the image to be identified.
[0091] In this step, the above-mentioned determination of the entropy value corresponding to the behavior of the detection target in the image to be identified, the contact force field matching degree between the detection target's motion and the preset target motion, and the panic index of the non-detection target in the image to be identified based on the frame sequence of the image to be identified can be performed by modeling the behavior of the detection target in the image to be identified in the image frame sequence based on the frame sequence of the image to be identified to obtain a state transition matrix of the memory target in the image to be identified in the frame sequence. Then, based on the state transition matrix, the entropy value corresponding to the behavior of the detection target in the image to be identified is determined. Then, the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified are determined. The contact force field matching degree between the motion of the detection target in the image to be identified and the preset target motion is determined based on the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified and the spatiotemporal characteristics of the joint pressure distribution corresponding to the preset target motion template. The panic index of the non-detection target in the image to be identified is calculated based on the facial expression characteristics (pupil dilation rate and limb withdrawal speed) of the non-detection target in the image to be identified.
[0092] The above steps S2221 to S2225 determine the confidence of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the action of the detection target and the preset target action, and the panic index of the non-detection target through the acquired image to be identified, the sequence frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified. The determination of the confidence of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the action of the detection target and the preset target action, and the panic index of the non-detection target facilitates the subsequent determination of the confidence of the target behavior of the detection target in the image to be identified.
[0093] In addition, in one embodiment, step S2222, performing human key point detection on the detection target in the sequence frames corresponding to the image to be identified to obtain the confidence level of the target action of the detection target in the image to be identified, includes:
[0094] Step S11 , performing human body key point detection on the detection target in the sequence frames corresponding to the image to be recognized, and obtaining the coordinates of the human skeleton joint points of the detection target in each image in the sequence frames.
[0095] In this step, the human body key point detection is performed on the detection target in the sequence frames corresponding to the image to be identified, and the coordinates of the human skeletal joint points of the detection target in each image in the sequence frames are obtained. This can be achieved by using an edge computing device. After receiving the sequence frames corresponding to the image to be identified from the video surveillance platform, the received image to be identified is preprocessed by grayscale normalization, noise reduction, etc., and then using a skeletal joint point trajectory acquisition module to perform human body key point detection on each image in the sequence frames, extracting the two-dimensional coordinates of a preset number of joint points in each image in real time (the extracted coordinate error is required to be ≤ pixels), and obtaining the coordinates of the human skeletal joint points of the detection target in each image in the sequence frames. The preset number can be set according to specific needs, and the location of each joint point can also be set according to specific needs. This is not specifically limited in this embodiment. For example, the preset number can be set to 17, and the locations of each joint point are respectively nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle.
[0096] Step S12, based on the coordinates of the human skeleton joint points of the detection target in each image in the sequence frame, determine the fitting degree of the preset curve, the number of joint acceleration mutations, and the variance of the trunk inclination angle of the detection target in the image to be identified; the fitting degree of the preset curve is the degree of fitting between the action curve of the detection target and the curve of the preset target action; based on the fitting degree of the preset curve, the number of joint acceleration mutations, and the variance of the trunk inclination angle of the detection target in the image to be identified, determine the confidence level of the target action of the detection target in the image to be identified.
[0097] The above-mentioned determination of the degree of fit of the preset curve of the detection target in the image to be identified may include first determining a coordinate sequence of key points of the detected detection target (person or object) in the sequence frames based on the coordinates of the detection target's human skeletal joints in each image in the sequence frames; calculating the mean square error between the coordinate sequence of the key points of the detected detection target (person or object) in the sequence frames and each corresponding point on the preset target action curve; and determining the degree of fit of the detection target's action curve with the preset target action curve based on the mean square error. Specifically, the greater the mean square error, the less the degree of fit of the detection target's action curve with the preset target action curve, and the smaller the mean square error, the greater the degree of fit of the detection target's action curve with the preset target action curve. Alternatively, the degree of fit of the detection target's action curve with the preset target action curve may be determined based on the relationship between the mean square error and a preset mean square error threshold. When the mean square error is greater than or equal to the preset mean square error threshold, it is determined that the detection target's action does not belong to the preset target action; when the mean square error is less than the preset mean square error threshold, it is determined that the detection target's action belongs to the preset target action. The above-mentioned curve may be a parabola. If the target action is a punch, the curve may be a parabola corresponding to the punch. It should be noted that the coordinates of the human skeletal joints of the target in each image in the sequence of frames may be used to determine whether the target's action trajectory is an abnormal curve. Specifically, the curvature of the target's action curve may be determined based on the coordinates of the human skeletal joints of the target in each image in the sequence of frames. When the curvature of the action curve is greater than a preset threshold, the target's action curve is determined to be an abnormal parabola, indicating that the target's action is likely a target action such as a punch.
[0098] Furthermore, the number of joint acceleration mutations can be the cumulative number of significant, discontinuous step changes in the joint angular acceleration (or linear acceleration) in a short period of time within a continuous period of time (a preset time length within the acquisition time of the sequence frame, for example, a time length of 100ms). The above determination of the number of joint acceleration mutations can be based on the coordinates of the human skeletal joint points of the detection target in each image in the sequence frame, and the velocity change rate (acceleration) of the target joint point within a preset time range is determined. When the acceleration is greater than the preset mutation threshold, it is recorded as a mutation, and the number of mutations within the preset time is counted to obtain the number of joint acceleration mutations. The above mutation threshold can be a preset value that can measure whether the acceleration of the joint point has mutated. When the acceleration is greater than the mutation threshold, it can be determined that a mutation has occurred at this time. When the acceleration is less than or equal to the mutation threshold, it can be determined that the acceleration is normal at this time. The above mutation threshold can be specifically set according to the specific scenario, and this embodiment does not make specific limitations here. For example, the above mutation threshold can be set to .
[0099] Additionally, the trunk tilt angle variance can be a statistic that measures the fluctuation in the trunk tilt angle over a continuous period of time, used to quantify body stability or the intensity of movement. Determining the trunk tilt angle variance can be based on the coordinates of the target's skeletal joints in each image of the sequence frame, determining the positional changes in the target's scapula and hip joint coordinates within the sequence frame, determining the target's trunk axis offset during this period based on the positional changes in the target's scapula and hip joint coordinates within the sequence frame, and determining the trunk tilt angle variance of the target in the image to be identified over the time length of the sequence frame (or a predetermined time length) based on the target's trunk axis offset during this period. If the trunk tilt angle variance of the target in the image to be identified over the time length of the sequence frame (or a predetermined time length) is greater than a preset threshold, the target is determined to be unbalanced, and an imbalance warning is triggered.
[0100] In this step, the process of determining the confidence level of the target action of the detection target in the image to be identified based on the degree of fit of the preset curve, the number of joint acceleration mutations, and the variance of the trunk inclination angle of the detection target in the image to be identified can be to set corresponding weights for the degree of fit of the preset curve, the number of joint acceleration mutations, and the variance of the trunk inclination angle according to the specific application scenario, and determine the confidence level S of the target action of the detection target in the image to be identified based on the weights of the degree of fit of the preset curve, the number of joint acceleration mutations, and the variance of the trunk inclination angle, and the values of the degree of fit of the preset curve, the number of joint acceleration mutations, and the variance of the trunk inclination angle. The specific calculation process is:
[0101] ;
[0102] In the above, A is the weight of the fitting degree of the preset curve, is the fitting degree of the preset curve of the detection target, B is the weight of the number of joint acceleration mutations, is the number of joint acceleration mutations of the detected target, C is the weight of the trunk tilt angle variance, is the trunk tilt angle variance.
[0103] It should be noted that the weight of the fitting degree of the preset curve, the weight of the number of joint acceleration mutations, and the weight of the trunk tilt angle variance can be set based on the specific situation. This embodiment does not make specific limitations here. For example, A can be set to 0.2, B to 0.3, and C to 0.5.
[0104] The above steps S11 to S12 perform human key point detection on the detection target in the sequence frames corresponding to the image to be identified, obtain the coordinates of the human skeletal joint points of the detection target in each image in the sequence frames, and then obtain the confidence of the target action of the detection target in the image to be identified. By determining the confidence of the target action of the detection target in the image to be identified, it is convenient to determine the confidence of the target behavior of the detection target in the image to be identified through the confidence of the target action of the detection target.
[0105] In one embodiment, step S2223, based on the audio data corresponding to the image to be identified, determining the sound abnormality of the audio data corresponding to the image to be identified, includes:
[0106] Step S21 : determining the sound features of the audio data based on the audio data corresponding to the image to be recognized.
[0107] The above-mentioned determination of the sound features of the audio data based on the audio data corresponding to the image to be recognized may be to extract a 13-dimensional feature vector of the audio data using Mel-frequency cepstral coefficients from the audio data.
[0108] Step S22: determining the sound abnormality of the audio data based on the sound features of the audio data.
[0109] Among them, the above-mentioned determination of the sound abnormality of the audio data based on the sound characteristics of the audio data can be based on the sound characteristics of the audio data, analyzing the energy ratio of the target frequency band in the sound characteristics of the audio data, and if the energy ratio of the target frequency band exceeds the preset ratio threshold of the total energy, the sound abnormality of the audio data is calculated by the Gaussian mixture model. The above-mentioned target frequency band can be specifically set based on the specific scenario, and this embodiment does not make specific restrictions here. For example, the above-mentioned target frequency band can be The aforementioned percentage threshold can be set based on specific scenarios and is not specifically limited in this embodiment. For example, the aforementioned percentage threshold can be 35%. The aforementioned Gaussian mixture model can be a Gaussian mixture model trained using a large amount of normal audio and can be used to determine the degree of sound abnormality in audio data.
[0110] The above steps S21 to S22 determine the sound abnormality of the audio data corresponding to the image to be identified based on the audio data corresponding to the image to be identified. The determination of the sound abnormality facilitates the subsequent determination of the confidence level of the target behavior of the detection target in the image to be identified based on the sound abnormality of the audio data.
[0111] In addition, in one embodiment, step S2224, determining the thermal intensity of the limb contact area of the detection target in the image to be identified based on the temperature field data corresponding to the image to be identified, includes:
[0112] Step S31: extracting the gradient features of the temperature field data corresponding to the image to be identified.
[0113] The above-mentioned extraction of the gradient features of the temperature field data corresponding to the image to be identified may be performed by using a 5×5 convolution kernel to extract the gradient features of the temperature field data.
[0114] Step S32: determining the thermal intensity of the temperature field data corresponding to the image to be identified based on the gradient feature.
[0115] In this step, determining the thermal intensity of the temperature field data corresponding to the image to be identified based on the gradient feature may be performed by using a preset convolutional neural network based on the gradient feature to generate the thermal intensity of the temperature field data corresponding to the image to be identified. The preset convolutional neural network may be a pre-trained network that can be used to generate the thermal intensity of temperature field data.
[0116] Step S33 : determining the thermal intensity of the limb contact area of the detection target in the image to be identified based on the thermal intensity of the temperature field data corresponding to the image to be identified.
[0117] The above-mentioned determination of the thermal intensity of the limb contact area of the detection target in the image to be identified based on the thermal intensity of the temperature field data corresponding to the image to be identified can be based on the thermal intensity of the temperature field data corresponding to the image to be identified and a preset contact area temperature mutation threshold to determine the contact area, and based on the determined contact area, determine the thermal intensity of the limb contact area of the detection target in the image to be identified. The above-mentioned preset contact area temperature mutation threshold can be specifically set based on the specific scenario, and this embodiment does not make specific limitations here. For example, the above-mentioned preset contact area temperature mutation threshold can be , that is, when the temperature mutation is greater than or equal to the preset contact area temperature mutation threshold, this area is determined to be a contact area.
[0118] The above steps S31 to S33 determine the thermal intensity of the limb contact area of the detection target in the image to be identified based on the temperature field data corresponding to the image to be identified. The determination of the thermal intensity of the limb contact area of the detection target in the image to be identified facilitates the subsequent determination of the confidence level of the target behavior of the detection target in the image to be identified based on the thermal intensity of the limb contact area of the detection target in the image to be identified.
[0119] Furthermore, in one embodiment, step S2225, based on the frame sequence of the image to be identified, determines the entropy value corresponding to the behavior of the detection target in the image to be identified, the contact force field matching degree between the detection target's action and the preset target action, and the panic index of the non-detection target in the image to be identified, including:
[0120] Step S41 , based on the frame sequence of the image to be identified, modeling the behavior of the detection target in the image to be identified in the image frame sequence, and obtaining a state transfer matrix of the memory target in the image to be identified in the frame sequence.
[0121] The state transition matrix can be an N×N probability table in a hidden Markov model, representing the probability of the system jumping to hidden state j at the next moment if it is currently in hidden state i. The above-mentioned modeling of the behavior of the detection target in the image to be identified in the image frame sequence based on the frame sequence of the image to be identified, thereby obtaining the state transition matrix of the memory target in the image to be identified in the frame sequence, can be achieved by using a hidden Markov model to model the behavior of the detection target in the image frame sequence and obtain the state transition matrix of the memory target in the image to be identified in the frame sequence. It should be noted that a preset number of images in the sequence of frames can be selected, and the selected images can be modeled using a hidden Markov model to obtain the state transition matrix of the detection target in the image to be identified in the frame sequence. The above-mentioned preset number can be specifically set according to specific needs and is not specifically limited in this embodiment. For example, the above-mentioned preset number can be 10.
[0122] Step S42: determining the entropy value corresponding to the behavior of the detection target in the image to be identified based on the state transition matrix.
[0123] Determining the entropy value corresponding to the behavior of the target in the image to be identified based on the state transition matrix may involve calculating the KL divergence (Kulbeck-Leibler divergence) between the state transition matrix and a normal behavior model, and determining the entropy value corresponding to the behavior of the target in the image to be identified based on the KL divergence between the state transition matrix and the normal behavior model. Determining the entropy value corresponding to the behavior of the target in the image to be identified based on the KL divergence between the state transition matrix and the normal behavior model may involve, when the KL divergence between the state transition matrix and the normal behavior model exceeds a preset divergence threshold, converting the difference between the state transition matrix and the normal behavior model by the amount greater than the preset divergence threshold into the entropy value corresponding to the behavior of the target in the image to be identified. The preset divergence threshold can be set based on specific needs and is not specifically limited in this embodiment. For example, the preset divergence threshold can be set to 1.2. When the KL divergence between the state transition matrix and the normal behavior model exceeds the preset divergence threshold, the target's behavior is determined to be abnormal, and the degree of abnormality of the abnormal behavior is represented by the entropy value.
[0124] The calculation process of converting the difference greater than the preset divergence threshold into the entropy value E corresponding to the behavior of the detection target in the image to be identified is:
[0125] ;
[0126] Where D is the KL divergence between the state transition matrix and the normal behavior model, and μ is the mean KL divergence of the normal behavior. The above mean KL divergence of the normal behavior can be obtained empirically.
[0127] Step S43: determining the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified.
[0128] The above-mentioned determination of the spatiotemporal characteristics of the pressure distribution of the joints of the detection target in the image to be identified may be performed by using a preset model to determine the spatiotemporal characteristics of the pressure distribution of the joints of the detection target in the image to be identified.
[0129] Step S44, based on the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified and the spatiotemporal characteristics of the joint pressure distribution corresponding to the preset target action template, determine the contact force field matching degree between the action of the detection target in the image to be identified and the preset target action.
[0130] In this step, the contact force field matching degree between the movement of the target in the image to be identified and the preset target movement is determined based on the spatiotemporal characteristics of the joint pressure distribution of the target in the image to be identified and the spatiotemporal characteristics of the joint pressure distribution corresponding to the preset target movement template. This can be achieved by using an improved dynamic time warping algorithm to match the spatiotemporal characteristics of the joint pressure distribution of the target in the image to be identified with the spatiotemporal characteristics of the joint pressure distribution corresponding to the preset target movement template to obtain the contact force field matching degree between the movement of the target in the image to be identified and the preset target movement. The preset target movement can be a standard fighting movement (e.g., pushing, punching, kicking, etc.).
[0131] The calculation formula for determining the contact force field matching degree M between the motion of the detection target in the image to be identified and the preset target motion is:
[0132] ;
[0133] Among them, the above The temporal and spatial characteristics of the joint pressure distribution corresponding to the preset target action template, It is the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified.
[0134] When M is greater than a preset contact threshold C1, contact verification is triggered. The preset contact threshold can be set according to specific needs and is not specifically limited in this embodiment. For example, the preset contact threshold can be 0.5.
[0135] Step S45 , based on the facial expression features of the non-detection target in the image to be identified, including pupil dilation rate and limb retraction speed, calculate the panic index of the non-detection target in the image to be identified.
[0136] The non-detected target may be any target other than the detected target in the image to be identified, specifically a person other than the detected target in the image to be identified. Calculating the panic index of the non-detected target in the image to be identified based on the facial expression features of the non-detected target in the image to be identified may involve extracting the non-detected target's facial expression features (pupil dilation rate and limb withdrawal speed) using a Vision Transformer, and then performing a weighted calculation based on the non-detected target's facial expression features and the weights of the facial expression features (pupil dilation rate and limb withdrawal speed) to determine the panic index D of the non-detected target in the image to be identified.
[0137] When pupil dilation rate , and the limb withdrawal speed is greater than 0.8 m / s, the D value exceeds the warning threshold of 0.7. At this point, a target behavior such as fighting may have occurred. The weights of the pupil dilation rate and the limb withdrawal speed can be set according to the specific situation and are not specifically limited in this embodiment. For example, the weight of the pupil dilation rate can be 0.4, and the weight of the limb withdrawal speed can be 0.6.
[0138] The above steps S41 to S45 determine the entropy value corresponding to the behavior of the detection target in the image to be identified, the contact force field matching degree between the movement of the detection target and the preset target movement, and the panic index of the non-detection target in the image to be identified based on the frame sequence of the image to be identified. The determination of the entropy value corresponding to the behavior of the detection target in the image to be identified, the contact force field matching degree between the movement of the detection target and the preset target movement, and the panic index of the non-detection target in the image to be identified facilitates the subsequent determination of the confidence level of the target behavior of the detection target in the image to be identified.
[0139] In one embodiment, step S224 determines the confidence level of the target behavior of the detection target in the image to be identified based on the confidence level of the target behavior of the detection target in the image to be identified, the degree of sound abnormality of the audio data, the thermal intensity of the contact area of the detection target's body, the entropy value corresponding to the detection target's behavior, the degree of contact force field matching between the detection target's movement and a preset target movement, and the panic index of the non-detection target, including:
[0140] Step S2242, based on the confidence of the target action of the detection target in the image to be identified, the panic index of the non-detection target, the thermal intensity of the body contact area of the detection target, and the sound abnormality of the audio data, determines whether the detection target in the image to be identified meets the preset first judgment condition; the preset first judgment condition is that the action meets the preset first action preset condition, and the sound meets the preset sound preset condition.
[0141] The above-mentioned preset first action preset condition can be specifically set according to the specific scenario, and is not specifically limited in this embodiment.
[0142] For example, the above-mentioned preset first action preset condition may be:
[0143] ;
[0144] Where S is the confidence of the target action of the detected target in the image to be identified, and D is the KL divergence between the state transition matrix and the normal behavior model.
[0145] The above-mentioned preset sound conditions can be set according to specific scenarios and are not specifically limited in this embodiment. D is the panic index of the non-detected target, which is used to compensate for environmental interference. The higher the panic level of the surrounding people, the lower the trigger threshold of S.
[0146] For example, the preset sound conditions may be:
[0147] ;
[0148] Among them, A1 is the sound abnormality of the audio data, and H is the thermal intensity of the body contact area of the detection target.
[0149] It should be noted that when the thermal imaging detection device detects physical contact, the weight of the audio data's sound abnormality is increased. If the target in the image to be identified meets the preset first judgment condition, the target's behavior may be targeted behavior or high-energy physical interaction (such as strenuous exercise). If the target in the image to be identified does not meet the preset first judgment condition, the target's behavior is normal (such as walking, talking, etc.).
[0150] Step S2244, when the detection target in the image to be identified meets the preset first judgment condition, based on the entropy value corresponding to the behavior of the detection target in the image to be identified, the contact force field matching degree between the action of the detection target and the preset target action, and the panic index of the non-detection target, it is judged whether the detection target in the image to be identified meets the preset second judgment condition; the preset second judgment condition is that the action meets at least one of the preset second action preset condition or the preset collision condition.
[0151] The above-mentioned preset second action preset condition can be specifically set according to the specific scenario, and is not specifically limited in this embodiment.
[0152] For example, the above-mentioned preset second action preset condition may be:
[0153] ;
[0154] Wherein, C1 is a preset contact threshold, and E is the entropy value corresponding to the behavior of the detection target in the image to be identified.
[0155] The above-mentioned preset collision conditions can be specifically set according to specific scenarios and are not specifically limited in this embodiment.
[0156] For example, the above-mentioned preset collision condition may be:
[0157] ;
[0158] Among them, C1 is the preset contact threshold, and D is the KL divergence between the state transition matrix and the normal behavior model.
[0159] If the detected object in the image to be identified meets the preset second judgment condition, it can be determined that the current detected object is highly suspected of target behavior and may be in a conflict or accidental violent collision. If the preset second judgment condition is not met, it is determined to be non-target behavior.
[0160] Step S2246, when the detection target in the image to be identified meets the preset second judgment condition, the confidence of the target behavior of the detection target in the image to be identified is determined based on the confidence of the target action of the detection target in the image to be identified, the entropy value corresponding to the behavior of the detection target, the panic index of the non-detection target, and the contact force field matching degree between the action of the detection target and the preset target action.
[0161] The calculation process for determining the confidence level P of the target action of the detection target in the image to be identified based on the confidence level of the target action of the detection target in the image to be identified, the entropy value corresponding to the detection target's behavior, the panic index of the non-detection target, and the contact force field matching degree between the detection target's action and the preset target action is as follows:
[0162] P=0.4S'+0.3E'+0.2D'+0.1C';
[0163] Among them, S' is the anti-misjudgment mechanism for cross-thermal environments, preventing misjudgments caused by excessively fast motion trajectories in some non-collision-free environments, such as dancing and basketball. S' = S ÷ (1 + 0.2|1 - H|). Through cross-validation of thermal imaging and motion modalities, when the H value is high (contact), S' is amplified, increasing the confidence weight of the target motion detected. It is a dynamic trajectory and thermal sensing anti-misjudgment mechanism to prevent misjudgment caused by hugging, supporting and other behaviors that have contact but low thermal sensitivity and no violent trajectory. ; D' is an anti-misjudgment mechanism that cross-validates panic level and movement trajectory to prevent misjudgment caused by violent behavior. ; E' is a multi-dimensional confidence mechanism, which introduces the mixed behavior entropy of force field matching and acoustic effect to reflect the multiple dimensions that affect confidence. .
[0164] The weights of the above parameters can be set according to actual scenarios, and are not specifically limited in this embodiment.
[0165] The above steps S2242 to S2246 determine the confidence of the target behavior of the detection target in the image to be identified based on the confidence of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the action of the detection target and the preset target action, and the panic index of the non-detection target. By determining the confidence of the target behavior of the detection target in the image to be identified, it is convenient to subsequently determine the detection result of the target behavior in the image to be identified based on the confidence of the target behavior of the detection target in the image to be identified.
[0166] The present embodiment is described and illustrated below through preferred embodiments.
[0167] Figure 3 This is a flow chart of a target behavior detection method provided by a preferred embodiment of the present application. Figure 3 As shown, the target behavior detection method includes the following steps:
[0168] Step S301: Acquire an image to be recognized, a sequence of frames corresponding to the image to be recognized, audio data corresponding to the image to be recognized, and temperature field data corresponding to the image to be recognized; the sequence of frames corresponding to the image to be recognized is a time-sequential frame sequence formed by a preset number of images adjacent to the image to be recognized and the image to be recognized;
[0169] Step S302: determining a detection target in the image to be identified based on the acquired image to be identified;
[0170] Step S303, performing human key point detection on the detection target in the sequence frames corresponding to the image to be identified, and obtaining the confidence level of the target action of the detection target in the image to be identified;
[0171] Step S304, determining the sound abnormality degree of the audio data corresponding to the image to be identified based on the audio data corresponding to the image to be identified;
[0172] Step S305, determining the thermal intensity of the limb contact area of the detection target in the image to be identified based on the temperature field data corresponding to the image to be identified;
[0173] Step S306 , determining, based on the frame sequence of the image to be identified, an entropy value corresponding to the behavior of the detection target in the image to be identified, a degree of contact force field matching between the detection target's action and a preset target action, and a panic index of non-detection targets in the image to be identified;
[0174] Step S307: Determine whether the target in the image to be identified meets a preset first judgment condition based on the confidence level of the target action of the target in the image to be identified, the panic index of the non-detected target, the thermal intensity of the body contact area of the target, and the sound abnormality of the audio data; the preset first judgment condition is that the action meets a preset first action preset condition and the sound meets a preset sound preset condition;
[0175] Step S308: When the detection target in the image to be identified meets the preset first judgment condition, based on the entropy value corresponding to the behavior of the detection target in the image to be identified, the contact force field matching degree between the detection target's action and the preset target action, and the panic index of the non-detection target, it is determined whether the detection target in the image to be identified meets the preset second judgment condition; the preset second judgment condition is that the action meets at least one of the preset second action preset condition or the preset collision condition;
[0176] Step S309: When the detection target in the image to be identified meets the preset second judgment condition, the confidence level of the target behavior of the detection target in the image to be identified is determined based on the confidence level of the target action of the detection target in the image to be identified, the entropy value corresponding to the target action, the panic index of the non-detected targets, and the contact force field matching degree between the target action and the preset target action.
[0177] Step S310 : determining a detection result of the target behavior in the image to be identified based on the confidence level of the target behavior of the detected target in the image to be identified.
[0178] The above steps S301 to S310 determine the confidence of the target behavior of the detection target in the image to be identified by obtaining the image to be identified, the sequence frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified. Then, based on the confidence of the target behavior of the detection target in the image to be identified, the detection result of the target behavior in the image to be identified is determined. The confidence of the target behavior of the detection target in the image to be identified is determined by combining the image to be identified, the sequence frames corresponding to the image to be identified, the audio data and the temperature field data, and multiple data. This makes the detection result of the target behavior more accurate, avoids missed detection caused by single data detection, and when performing target behavior detection, there is no need to transmit the data to be detected to the cloud, which solves the problems of existing target behavior detection methods, such as high missed detection rate, low target recognition accuracy, and security risks of data leakage and privacy leakage.
[0179] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0180] Based on the same inventive concept, a target behavior detection device is also provided in this embodiment, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the details that have been described will not be repeated. The terms "module", "unit", "sub-unit", etc. used below can implement a combination of software and / or hardware for predetermined functions. Although the devices described in the following embodiments are preferably implemented in software, implementation by hardware, or a combination of software and hardware, is also possible and conceivable.
[0181] In one embodiment, Figure 4 This is a structural block diagram of a target behavior detection device provided by an embodiment of the present application. Figure 4 As shown, the target behavior detection device includes:
[0182] The data acquisition module 42 is used to acquire the image to be recognized, the sequence of frames corresponding to the image to be recognized, the audio data corresponding to the image to be recognized, and the temperature field data corresponding to the image to be recognized; the sequence of frames corresponding to the image to be recognized is a time-sequential frame sequence consisting of a preset number of image frames adjacent to the image to be recognized and the image to be recognized;
[0183] a confidence determination module 44 for determining the confidence of the target behavior of the detection target in the image to be identified based on the acquired image to be identified, the sequence of frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified;
[0184] and a detection module 46 for determining a detection result of the target behavior in the image to be identified based on the confidence level of the target behavior of the detected target in the image to be identified.
[0185] The above-mentioned target behavior detection device determines the confidence of the target behavior of the detection target in the image to be identified by obtaining the image to be identified, the sequence frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified. Then, based on the confidence of the target behavior of the detection target in the image to be identified, the detection result of the target behavior in the image to be identified is determined. It determines the confidence of the target behavior of the detection target in the image to be identified by combining the image to be identified, the sequence frames corresponding to the image to be identified, the audio data and the temperature field data, and multiple data. This makes the detection result of the target behavior more accurate, avoids missed detection caused by single data detection, and when performing target behavior detection, there is no need to transmit the data to be detected to the cloud, which solves the problems of existing target behavior detection methods, such as high missed detection rate, low target recognition accuracy, and security risks of data leakage and privacy leakage.
[0186] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0187] In one embodiment, a target behavior detection system is provided, which includes: a data acquisition module, a feature extraction module, a classification decision module, and an optimization deployment module.
[0188] The data acquisition module includes an infrared imaging sensor unit, a skeleton key point trajectory acquisition unit and a high-frequency voiceprint acquisition unit; the infrared imaging sensor unit is used to acquire the instantaneous temperature changes in the limb contact area in the image to be identified, and obtain the temperature field data corresponding to the image to be identified; the skeleton key point trajectory acquisition unit is used to perform human key point detection on the detection target in the sequence frame corresponding to the image to be identified, and obtain the coordinates of the human skeleton joint points of the detection target in each image in the sequence frame; the high-frequency voiceprint acquisition unit is used to acquire the audio data corresponding to the image to be identified.
[0189] The above-mentioned skeleton key point trajectory acquisition unit can also include multiple cameras for real-time acquisition of coordinate information of human skeleton key points. The high-frequency voiceprint acquisition unit is used to identify and acquire sound features such as high-frequency screams, heavy object impact sounds, and glass breaking sounds.
[0190] A feature extraction module is used to extract the gradient features of the temperature field data corresponding to the image to be identified; based on the gradient features of the temperature field data corresponding to the image to be identified, determine the thermal intensity of the temperature field data corresponding to the image to be identified; based on the coordinates of the human skeletal joint points of the detection target in each image in the sequence frame, determine the fitting degree of the preset curve, the number of joint acceleration mutations, and the variance of the trunk inclination angle of the detection target in the image to be identified; the fitting degree of the preset curve is the degree of fitting between the action curve of the detection target and the curve of the preset target action; based on the fitting degree of the preset curve, the number of joint acceleration mutations, and the variance of the trunk inclination angle of the detection target in the image to be identified, determine the confidence level of the target action of the detection target in the image to be identified; based on the audio data corresponding to the image to be identified, determine the sound features of the audio data; based on the sound features of the audio data, determine the sound abnormality of the audio data.
[0191] The feature extraction module is used to preprocess and standardize the collected data. In this module, the system preprocesses and standardizes the collected data to extract key features of target behaviors such as fighting. This feature extraction module specifically utilizes a multi-head self-attention mechanism and Transformer network to enable cross-modal correlation analysis of multimodal data, enhancing feature representation capabilities.
[0192] In the feature extraction module, the system uses a weighted fusion algorithm to convert skeletal trajectory features into an action confidence score S. The fit of the preset curve is assigned a weight of 0.2 to reflect the standardization of the action, the number of joint acceleration mutations is assigned a weight of 0.3 to capture violent characteristics, and the variance of the trunk tilt angle is assigned a weight of 0.5 to quantify the degree of body imbalance. Voiceprint features use Mel-frequency cepstral coefficients to extract the energy ratio of abnormal frequency bands to construct the sound abnormality level of the audio data. For thermal imaging data, a convolutional neural network is used to extract the temperature gradient characteristics of the contact area to generate the thermal intensity H of the temperature field data.
[0193] The feature extraction module is also used to calculate the KL divergence of the state transition matrix of the detection target in the sequence frame through the hidden Markov model to obtain the entropy value corresponding to the behavior of the detection target; use the contact force field model to determine the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified; based on the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified, determine the action of the detection target in the image to be identified; based on the action of the detection target in the image to be identified and the preset target action template, determine the contact force field matching degree between the action of the detection target in the image to be identified and the preset target action; use the micro-expression extraction model to extract the expression features of the faces of non-detection targets in the image to be identified; the expression features include pupil dilation rate and limb withdrawal speed; based on the expression features of the faces of non-detection targets in the image to be identified, determine the panic index of the non-detection targets in the image to be identified.
[0194] This module uses a dynamic threshold classifier to generate a multi-scenario noise data training model based on adversarial learning. Using a generative adversarial network, the original features are mapped to a space suitable for classification. The classifier dynamically adjusts detection sensitivity based on environmental parameters and crowd density, setting appropriate threshold rules for different scenarios and effectively reducing false alarm rates.
[0195] The classification decision module is configured to determine whether the detection target in the image to be identified satisfies a preset first judgment condition based on the confidence level of the detection target's target action in the image to be identified, the panic index of non-detected targets, the thermal intensity of the detection target's limb contact area, and the sound abnormality of the audio data. The preset first judgment condition is that the action satisfies a preset first action preset condition and the sound satisfies a preset sound preset condition. When the detection target in the image to be identified satisfies the preset first judgment condition, the module determines whether the detection target in the image to be identified satisfies a preset second judgment condition based on the entropy value corresponding to the detection target's action in the image to be identified, the degree of contact force field matching between the detection target's action and the preset target action, and the panic index of non-detected targets in the image to be identified. The preset second judgment condition is that the action satisfies at least one of a preset second action preset condition or a preset collision condition. When the detection target in the image to be identified satisfies the preset second judgment condition, the module determines the confidence level of the detection target's target action in the image to be identified based on the confidence level of the detection target's target action in the image to be identified, the entropy value corresponding to the detection target's action, the panic index of non-detected targets, and the degree of contact force field matching between the detection target's action and the preset target action.
[0196] The classification decision module is the core of the system, introducing a multi-dimensional dynamic decision-making mechanism. First, the system constructs a basic feature layer (feature extraction module) to quantify skeletal trajectories, voiceprint features, and thermal imaging data into action confidence, acoustic anomaly, and thermal intensity, respectively. The calculation of action confidence S combines the parabolic fit of the punching trajectory, the number of joint acceleration mutations, and the variance of the torso tilt angle to form a composite indicator. This module also introduces an advanced feature layer. By analyzing the Markov transition probability matrix of the action sequence, it calculates the KL divergence E of the behavior pattern deviation from the normal state, and constructs a contact force field model in real time to output the contact matching degree C1. In addition, the system also integrates a scene semantic analysis module to identify the panic index D of the surrounding people (non-detection targets).
[0197] The final decision-making process utilizes a three-level fusion strategy, comprehensively considering multiple indicators at both the basic and advanced feature layers. Primary threshold screening requires that action confidence and acoustic anomaly levels meet certain thresholds; intermediate verification requires that behavioral entropy assessment or contact matching meet certain criteria; and advanced confirmation uses environmental factors for correction and calculation of the final confidence level. Experiments have shown that this approach significantly improves the accuracy of identifying fighting behaviors in specific scenarios, while significantly reducing false positives.
[0198] The decision-making mechanism adopts a three-level cascade verification architecture: primary screening sets dynamic thresholds (Environmental Disturbance Compensation), (Thermal enhancement); intermediate verification introduces composite conditions ;Confidence in high-level decisions ,in, , achieving modal cross-validation. Experimental data shows that this mechanism maintains an 89.7% accuracy rate even when the crowd density exceeds 3 people / m2, a 23.6% improvement over traditional methods. The system integrates an online adversarial training module. When three consecutive discrepancies with manual judgment are detected, it automatically generates adversarial samples to update the feature library, reducing the false alarm rate by 8.2% per month.
[0199] The optimized deployment layer utilizes an end-to-end lightweight approach, migrating the key features of a large 3D CNN teacher model to a lightweight student model through layered knowledge distillation. This layer also includes a model compression and acceleration module, with an 8-bit integer quantization scheme customized for edge devices, significantly improving inference speed. The system utilizes low-cost hardware built on domestic HiSilicon chips, transmitting only skeletal point coordinates and voiceprint spectrograms. ISO 31700-2024 privacy certification ensures data security and privacy.
[0200] The optimization deployment module is used to determine the detection result of the target behavior of the detection target in the image to be identified based on the confidence of the target behavior of the detection target in the image to be identified; generate an adversarial sample of the target behavior based on the detection result of the target behavior of the detection target in the image to be identified and the manual detection result; and use the adversarial sample of the target behavior to update the model in the feature extraction module.
[0201] This optimized deployment module utilizes an end-to-end lightweight approach, migrating the key features of a large 3D CNN (Three-Dimensional Convolutional Neural Network) teacher model to a lightweight student model through layered knowledge distillation. This layer also includes a model compression and acceleration module, using an 8-bit integer quantization scheme customized for edge devices to significantly improve inference speed. The system utilizes hardware in the low-cost range of hundreds of yuan, built on traditional chips. It transmits only skeletal point coordinates and voiceprint spectrograms and is ISO 31700-2024 privacy certified, ensuring data security and privacy.
[0202] The system also incorporates a dynamic learning mechanism. When multiple high-confidence detections are detected but manually verified as false positives, the adversarial sample generator is automatically triggered to update the feature library, further enhancing the system's adaptability and robustness. This multimodal edge intelligence-based detection system not only enables accurate real-time detection of fighting in complex scenarios, but also effectively reduces missed detection and false positive rates, providing strong support for public safety monitoring.
[0203] In one embodiment, the target behavior detection system provided in the embodiment of the present application can be applied in school security monitoring scenarios, with the corridors of teaching buildings as the specific monitoring areas. After the system is connected to various sensors, it collects multimodal data and analyzes and processes it to achieve accurate real-time detection of fighting behaviors.
[0204] In one example, infrared thermal imaging sensors, skeletal joint trajectory acquisition modules, and high-frequency voiceprint acquisition modules can be deployed at the corners and intersections of teaching building corridors, with the hardware integrated with existing surveillance poles. At the software level, interfaces are established with the school's video surveillance platform to implement protocol integration. The surveillance platform then captures real-time streaming media from the corridor cameras (1080P resolution, 25fps), transmits keyframe images (one frame every 200ms) from the video stream to the edge computing device via the RTSP protocol, and simultaneously outputs an audio stream to the high-frequency voiceprint acquisition module, forming a real-time access link for multimodal data.
[0205] After receiving keyframe images from the video surveillance platform using an edge computing device, an infrared thermal imaging sensor simultaneously collects temperature field data (320×240 resolution) for the same scene. The skeletal joint trajectory acquisition module detects key human body points in the images and extracts the two-dimensional coordinates of 17 joints in real time (with an error of ≤ 2 pixels). The collected image data undergoes grayscale normalization and noise reduction in a preprocessing module. The temperature field data and joint coordinate information are packaged into a multimodal data frame (JSON format, including a timestamp, coordinate matrix, and temperature matrix). This data is transmitted to the feature extraction layer via the edge node's Gigabit Ethernet port using UDP (User Datagram Protocol). Transmission latency is kept within 50ms, ensuring real-time data delivery.
[0206] The feature extraction module is deployed in the high-performance computing unit of the edge computing device, utilizing a heterogeneous computing architecture to enable parallel processing of multimodal data. The basic feature extraction layer receives multimodal data frames transmitted by the data acquisition layer and then extracts skeletal features, voiceprint features, and thermal imaging features. The advanced feature extraction layer constructs a spatiotemporal dynamic analysis model based on these basic features. The specific processing logic includes the calculation of behavioral entropy E, contact force field matching C1, and panic index D.
[0207] In addition, the system will automatically adjust the detection sensitivity and resource allocation according to the environmental characteristics and human activity patterns in different time periods: during daytime classes (8:00-12:00, 14:00-18:00), each sensor maintains a standard acquisition frequency, and the feature extraction layer prioritizes high-frequency voiceprints and bone trajectory features. At the same time, the primary threshold of the action confidence S is set to 0.7 to balance detection accuracy and computing power consumption; during the break period (10 minutes between classes), the mobility of the crowd increases, and the system automatically reduces the sensor acquisition frequency by 20%, and sets S The threshold is raised to 0.85, and the weight of the variance of the torso inclination angle is reduced at the same time to reduce misjudgments caused by normal crowd pushing. At this time, the computing power resources of the edge device prioritize multi-target detection; during the night period (22:00 to 6:00 the next day), the infrared thermal imaging sensor switches to night vision mode, and the voiceprint module starts the noise reduction algorithm, and simultaneously reduces the warning threshold of the thermal intensity H from 0.6 to 0.45, thereby improving the sensitivity of contact detection in dim environments. At the same time, the dynamic quantization technology is used to reduce the model computing power consumption by 40%, achieving a balance between energy saving and high sensitivity.
[0208] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0209] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0210] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A target behavior detection method, characterized in that: The method comprises: Acquire an image to be recognized, a sequence of frames corresponding to the image to be recognized, audio data corresponding to the image to be recognized, and temperature field data corresponding to the image to be recognized; the sequence of frames corresponding to the image to be recognized is a time-sequential frame sequence formed by a preset number of images adjacent to the image to be recognized and the image to be recognized; Determining a confidence level of a target behavior of a detection target in the image to be identified based on the acquired image to be identified, a sequence of frames corresponding to the image to be identified, audio data corresponding to the image to be identified, and temperature field data corresponding to the image to be identified; A detection result of the target behavior in the image to be recognized is determined based on the confidence level of the target behavior of the detected target in the image to be recognized.
2. The target behavior detection method according to claim 1, characterized in that: The determining of the confidence level of the target behavior of the detection target in the image to be identified based on the acquired image to be identified, the sequence of frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified includes: Based on the acquired image to be identified, the sequence of frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified, determine the confidence level of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the detection target action and the preset target action, and the panic index of the non-detection target; The confidence level of the target action of the detection target in the image to be identified is determined based on the confidence level of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the detection target's action and the preset target action, and the panic index of the non-detection target.
3. The target behavior detection method according to claim 2, characterized in that: The method of determining, based on the acquired image to be identified, the sequence of frames corresponding to the image to be identified, the audio data corresponding to the image to be identified, and the temperature field data corresponding to the image to be identified, the confidence level of the target action of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the detection target action and the preset target action, and the panic index of the non-detection target, includes: Determining a detection target in the image to be identified based on the acquired image to be identified; Performing human key point detection on the detection target in the sequence frames corresponding to the image to be identified, and obtaining a confidence level of a target action of the detection target in the image to be identified; determining, based on the audio data corresponding to the image to be identified, the degree of sound abnormality of the audio data corresponding to the image to be identified; Determining the thermal intensity of the limb contact area of the detection target in the image to be identified based on the temperature field data corresponding to the image to be identified; Based on the frame sequence of the image to be identified, the entropy value corresponding to the behavior of the detection target in the image to be identified, the contact force field matching degree between the detection target's action and the preset target action, and the panic index of the non-detection target in the image to be identified are determined.
4. The target behavior detection method according to claim 3, characterized in that: The performing of human key point detection on the detection target in the sequence frames corresponding to the image to be identified to obtain the confidence level of the target action of the detection target in the image to be identified includes: Performing human body key point detection on the detection target in the sequence frames corresponding to the image to be identified, respectively, to obtain the coordinates of the human skeleton joint points of the detection target in each image in the sequence frames; Based on the coordinates of the human skeletal joint points of the detection target in each image in the sequence frame, the fitting degree of the preset curve, the number of joint acceleration mutations, and the variance of the trunk inclination angle of the detection target in the image to be identified are determined; the fitting degree of the preset curve is the degree of fitting between the action curve of the detection target and the curve of the preset target action; based on the fitting degree of the preset curve of the detection target in the image to be identified, the number of joint acceleration mutations, and the variance of the trunk inclination angle, the confidence level of the target action of the detection target in the image to be identified is determined.
5. The target behavior detection method according to claim 3, characterized in that: The determining, based on the audio data corresponding to the image to be recognized, the sound abnormality degree of the audio data corresponding to the image to be recognized includes: Determining sound features of the audio data based on the audio data corresponding to the image to be recognized; The sound abnormality degree of the audio data is determined based on the sound feature of the audio data.
6. The target behavior detection method according to claim 3, characterized in that: The determining, based on the temperature field data corresponding to the image to be identified, the thermal intensity of the limb contact area of the detection target in the image to be identified, includes: Extracting gradient features of temperature field data corresponding to the image to be identified; Determining the thermal intensity of the temperature field data corresponding to the image to be identified based on the gradient feature; The thermal intensity of the limb contact area of the detection target in the image to be identified is determined based on the thermal intensity of the temperature field data corresponding to the image to be identified.
7. The target behavior detection method according to claim 3, characterized in that: The determining, based on the frame sequence of the image to be identified, an entropy value corresponding to the behavior of the detection target in the image to be identified, a contact force field matching degree between the movement of the detection target and a preset target movement, and a panic index of a non-detection target in the image to be identified includes: Based on the frame sequence of the image to be identified, modeling the behavior of the detection target in the image to be identified in the image frame sequence, and obtaining a state transition matrix of the memory target in the image to be identified in the frame sequence; Determining, based on the state transition matrix, an entropy value corresponding to a behavior of the detection target in the image to be identified; Determining the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified; Determining a degree of contact force field matching between a movement of the detection target in the image to be identified and a preset target movement based on the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified and the spatiotemporal characteristics of the joint pressure distribution corresponding to a preset target movement template; The panic index of the non-detection target in the image to be identified is calculated based on the facial expression features of the non-detection target in the image to be identified; the facial expression features include pupil dilation rate and limb withdrawal speed.
8. The target behavior detection method according to claim 2, characterized in that: The step of determining the confidence level of the target behavior of the detection target in the image to be identified based on the confidence level of the target behavior of the detection target in the image to be identified, the sound abnormality of the audio data, the thermal intensity of the limb contact area of the detection target, the entropy value corresponding to the behavior of the detection target, the contact force field matching degree between the detection target's action and a preset target action, and the panic index of the non-detection target includes: Determining whether the detection target in the image to be identified meets a preset first judgment condition based on the confidence level of the target action of the detection target in the image to be identified, the panic index of the non-detection target, the thermal intensity of the body contact area of the detection target, and the sound abnormality of the audio data; the preset first judgment condition is that the action meets a preset first action preset condition and the sound meets a preset sound preset condition; When the detection target in the image to be identified meets a preset first judgment condition, based on the entropy value corresponding to the behavior of the detection target in the image to be identified, the contact force field matching degree between the movement of the detection target and the preset target movement, and the panic index of the non-detection target, it is determined whether the detection target in the image to be identified meets a preset second judgment condition; the preset second judgment condition is that the movement meets at least one of a preset second movement preset condition or a preset collision condition; When the detection target in the image to be identified meets the preset second judgment condition, the confidence of the target behavior of the detection target in the image to be identified is determined based on the confidence of the target action of the detection target in the image to be identified, the entropy value corresponding to the behavior of the detection target, the panic index of the non-detection target, and the contact force field matching degree between the action of the detection target and the preset target action.
9. A target behavior detection device, characterized in that: The device comprises: a data acquisition module, configured to acquire an image to be recognized, a sequence of frames corresponding to the image to be recognized, audio data corresponding to the image to be recognized, and temperature field data corresponding to the image to be recognized; the sequence of frames corresponding to the image to be recognized is a time-sequential frame sequence formed by a preset number of image frames adjacent to the image to be recognized and the image to be recognized; a confidence determination module, configured to determine a confidence level of a target behavior of a detection target in the image to be identified based on the acquired image to be identified, a sequence of frames corresponding to the image to be identified, audio data corresponding to the image to be identified, and temperature field data corresponding to the image to be identified; and a detection module for determining a detection result of the target behavior in the image to be identified based on the confidence level of the target behavior of the detected target in the image to be identified.
10. A target behavior detection system, characterized in that: The system includes: a data acquisition module, a feature extraction module, a classification decision module, and an optimization deployment module; The data acquisition module includes an infrared imaging sensor unit, a skeleton key point trajectory acquisition unit, and a high-frequency voiceprint acquisition unit; the infrared imaging sensor unit is used to acquire instantaneous temperature changes in the limb contact area in the image to be identified, and obtain temperature field data corresponding to the image to be identified; the skeleton key point trajectory acquisition unit is used to perform human key point detection on the detection target in the sequence frame corresponding to the image to be identified, and obtain the coordinates of the human skeleton joint points of the detection target in each image in the sequence frame; the high-frequency voiceprint acquisition unit is used to acquire audio data corresponding to the image to be identified; The feature extraction module is used to extract the gradient features of the temperature field data corresponding to the image to be identified; based on the gradient features of the temperature field data corresponding to the image to be identified, determine the thermal intensity of the temperature field data corresponding to the image to be identified; based on the coordinates of the human skeletal joint points of the detection target in each image in the sequence frame, determine the fitting degree of the preset curve, the number of joint acceleration mutations, and the variance of the trunk inclination angle of the detection target in the image to be identified; the fitting degree of the preset curve is the degree of fitting between the action curve of the detection target and the curve of the preset target action; based on the fitting degree of the preset curve of the detection target in the image to be identified, the number of joint acceleration mutations, and the variance of the trunk inclination angle, determine the confidence of the target action of the detection target in the image to be identified; based on the audio data corresponding to the image to be identified, determine the sound features of the audio data; based on the sound features of the audio data, determine the sound abnormality of the audio data; The feature extraction module is further used to calculate the KL divergence of the state transition matrix of the detection target in the sequence frame through the hidden Markov model to obtain the entropy value corresponding to the behavior of the detection target; use the contact force field model to determine the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified; based on the spatiotemporal characteristics of the joint pressure distribution of the detection target in the image to be identified and the spatiotemporal characteristics of the joint pressure distribution corresponding to the preset target action template, determine the contact force field matching degree between the action of the detection target in the image to be identified and the preset target action; use the micro-expression extraction model to extract the expression features of the faces of people other than the non-detection targets in the image to be identified; the expression features include pupil dilation rate and limb withdrawal speed; based on the expression features of the faces of the non-detection targets in the image to be identified, determine the panic index of the non-detection targets in the image to be identified; The classification decision module is used to determine whether the detection target in the image to be identified meets a preset first judgment condition based on the confidence of the target action of the detection target in the image to be identified, the panic index of the non-detection target, the thermal intensity of the limb contact area of the detection target, and the sound abnormality of the audio data; the preset first judgment condition is that the action meets the preset first action preset condition and the sound meets the preset sound preset condition; when the detection target in the image to be identified meets the preset first judgment condition, based on the entropy value corresponding to the behavior of the detection target in the image to be identified, the action of the detection target is matched with the contact force field of the preset target action. degree, and the panic index of the non-detected target in the image to be identified, to determine whether the detected target in the image to be identified meets a preset second judgment condition; the preset second judgment condition is that the action meets at least one of a preset second action preset condition or a preset collision condition; when the detected target in the image to be identified meets the preset second judgment condition, determining the confidence of the target behavior of the detected target in the image to be identified based on the confidence of the target action of the detected target in the image to be identified, the entropy value corresponding to the behavior of the detected target, the panic index of the non-detected target, and the contact force field matching degree between the action of the detected target and the preset target action; An optimization deployment module is used to determine the detection result of the target behavior of the detection target in the image to be identified based on the confidence of the target behavior of the detection target in the image to be identified; generate an adversarial sample of the target behavior based on the detection result of the target behavior of the detection target in the image to be identified and the manual detection result; and use the adversarial sample of the target behavior to update the model in the feature extraction module.
Citation Information
Patent Citations
Skeleton sequence-based old person behavior identification method in infrared video
CN114724251A
Multi-mode integrated traffic abnormal event detection method
CN118298628A
Infrared image target detection method and system based on deep learning
CN118644723A
Behavior detection method, device and system, computer equipment and storage medium
CN119577651A
Action recognition system
WO2025053080A1
Cited By
Video abnormal behavior detection method and system based on image recognition
CN121505698A