Abnormal behavior identification and early warning method based on cloud edge cooperation mechanism and related system

Through the abnormal behavior recognition and early warning system of the cloud-edge collaboration mechanism, the coordinated work of edge IoT agents and cloud IoT agents is used to realize intelligent identification and early warning of smart security scenarios, improve the accuracy and security of early warnings, and ensure public safety.

CN120337107AActive Publication Date: 2025-07-18SICHUAN RUITING ZHIHUI TECH CO LTD

Patent Information

Application Number
CN202510825189.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The identification and early warning technology of smart security scenarios lacks intelligence, resulting in reduced security and inability to effectively ensure public safety.

Method used

An abnormal behavior recognition and early warning system based on the cloud-edge collaboration mechanism is adopted, and personnel images and environmental data are obtained through edge IoT agents for intelligent analysis. Cloud IoT agents perform preprocessing and drive multimodal large model diagnosis, determine early warning parameters, and conduct early warning operations and reinforcement learning.

Benefits of technology

It improves the accuracy of early warning operations and the intelligence of edge IoT intelligent bodies, improves the intelligence of smart security scenarios, and ensures public safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337107A_ABST
    Figure CN120337107A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and the technical field of computers, and provides an abnormal behavior recognition and early warning method and related system based on a cloud-edge cooperation mechanism, a first edge Internet of Things intelligent agent obtains personnel image data and environment data, and intelligently analyzes the personnel image data and the environment data; obtaining a first analysis result; the cloud internet-of-things intelligent agent is used for preprocessing the personnel image data and the environment data to obtain first personnel image data and first environment data; driving a multi-modal large model to diagnose the first person image data, the first environment data and the first analysis result to obtain a first diagnosis result, and determining a first early warning parameter according to the first diagnosis result; and the first edge internet-of-things intelligent agent performs early warning operation and reinforcement learning according to the first diagnosis result and the first early warning parameter. According to the embodiment of the invention, the intelligence of an intelligent security scene can be improved, so that the public safety is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of artificial intelligence technology and computer technology, and specifically to an abnormal behavior identification and early warning method and related system based on a cloud-edge collaboration mechanism. Background Art

[0002] Smart security scenarios have built a multi-dimensional active protection system by integrating technologies such as artificial intelligence, the Internet of Things, and big data. At present, the identification and early warning technologies of smart security scenarios are not intelligent, which reduces the security of smart security scenarios. Therefore, how to improve the intelligence of smart security scenarios to ensure public safety needs to be solved urgently. Summary of the invention

[0003] The embodiments of the present application provide an abnormal behavior identification and early warning method and related system based on a cloud-edge collaboration mechanism, which can enhance the intelligence of smart security scenarios to ensure public safety.

[0004] In a first aspect, an embodiment of the present application provides an abnormal behavior recognition and early warning system based on a cloud-edge collaboration mechanism, the system comprising: a plurality of edge IoT agents and a cloud IoT agent, wherein the cloud IoT agent allows a multimodal large model to be pre-configured in a driving cloud platform; wherein: A first edge IoT agent is used to obtain personnel image data and environmental data, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result; the first edge IoT agent is any edge IoT agent among the multiple edge IoT agents; The cloud IoT intelligent agent is used to pre-process the personnel image data and the environmental data to obtain first personnel image data and first environmental data; drive the multimodal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain a first diagnosis result, and determine a first warning parameter according to the first diagnosis result; The first edge IoT agent is used to perform warning operations and reinforcement learning based on the first diagnostic result and the first warning parameter.

[0005] In a second aspect, an embodiment of the present application provides an abnormal behavior recognition and early warning method based on a cloud-edge collaborative mechanism, which is applied to an abnormal behavior recognition and early warning system based on a cloud-edge collaborative mechanism, the system comprising: a plurality of edge IoT agents, a cloud IoT agent, the cloud IoT agent allowing a multimodal large model to be pre-configured in a driving cloud platform; the method comprising: Obtain personnel image data and environmental data through the first edge Internet of Things intelligent agent, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result; the first edge Internet of Things intelligent agent is any one of the multiple edge Internet of Things intelligent agents; Preprocess the personnel image data and the environmental data through the cloud Internet of Things intelligent agent to obtain first personnel image data and first environmental data; drive the multi-modal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain a first diagnosis result, and determine a first warning parameter according to the first diagnosis result; Perform a warning operation and reinforcement learning through the first edge Internet of Things intelligent agent according to the first diagnosis result and the first warning parameter.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, including a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the processor, and the programs include instructions for executing the steps in the second aspect of the embodiment of the present application.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute some or all of the steps described in the second aspect of the embodiment of the present application.

[0008] In a fifth aspect, an embodiment of the present application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute some or all of the steps described in the second aspect of the embodiment of the present application. The computer program product can be a software installation package.

[0009] Implementing the embodiments of the present application has the following beneficial effects: It can be seen that the abnormal behavior recognition and warning method and related system described in the embodiments of the present application based on the cloud-edge collaboration mechanism are applied to an abnormal behavior recognition and warning system based on the cloud-edge collaboration mechanism. The system includes: multiple edge IoT agents and a cloud IoT agent. The cloud IoT agent is allowed to drive a multi-modal large model pre-configured in the cloud platform. Among them, the first edge IoT agent obtains personnel image data and environmental data, and performs intelligent analysis on the personnel image data and environmental data to obtain a first analysis result. The first edge IoT agent is any one of the multiple edge IoT agents. The cloud IoT agent pre-processes the personnel image data and environmental data to obtain first personnel image data and first environmental data. The multi-modal large model is driven to diagnose the first personnel image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, and a first warning parameter is determined according to the first diagnosis result. The first edge IoT agent performs warning operations and reinforcement learning according to the first diagnosis result and the first warning parameter. On the one hand, the accuracy of the warning operation is improved. On the other hand, the intelligence of the edge IoT agent can be improved through reinforcement learning. Furthermore, the intelligence of the intelligent security scenario can be improved to ensure public safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0011] Figure 1 is a schematic diagram of the architecture of an abnormal behavior recognition and warning system based on the cloud-edge collaboration mechanism provided by the embodiments of the present application; Figure 2 is a schematic diagram of the scenario demonstration of a cloud IoT agent provided by the embodiments of the present application; Figure 3 is another schematic diagram of the scenario demonstration of a cloud IoT agent provided by the embodiments of the present application; Figure 4 is a schematic diagram of the scenario demonstration of an edge IoT agent provided by the embodiments of the present application; Figure 5 is a schematic diagram of the flow of an abnormal behavior recognition and warning method based on the cloud-edge collaboration mechanism provided by the embodiments of the present application; Figure 6 is a schematic diagram of the structure of an electronic device provided by the embodiments of the present application; Figure 7It is a functional unit composition block diagram of an abnormal behavior recognition and warning device based on a cloud-edge collaboration mechanism provided by an embodiment of the present application. Detailed implementation manners

[0012] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may also include unlisted steps or units in a possible example, or other steps or units inherent to these processes, methods, products or devices in a possible example.

[0013] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0014] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0015] The edge IoT agent and the cloud IoT agent involved in the embodiments of the present application can both be understood as devices including agents. In specific implementation, an agent can be specifically understood as: a software program, a robot or other automated devices, which have a certain degree of autonomy and intelligence.

[0016] In specific implementation, an agent can continuously learn and adapt through interaction with the environment to achieve specific goals. The core of an agent lies in its autonomy. An agent can adjust its behavior according to changes in the environment and show a certain level of intelligence.

[0017] Among them, the edge IoT agent and the cloud IoT agent may include at least one of the following devices: intelligent cameras, smartphones, tablets, intelligent robots, smart home devices, in-vehicle devices, intelligent driving recorders, wearable devices, computing devices, or other processing devices connected to a wireless modem, as well as various forms of user equipment (UE), mobile stations (MS), terminal devices, etc., which are not limited herein.

[0018] Among them, the edge IoT agent and the cloud IoT agent may also be servers. For example, the edge IoT agent may include an edge server, or an edge IoT device. The cloud IoT agent may be deployed in the cloud and may be a cloud server.

[0019] Please refer to Figure 1 , Figure 1 FIG. is a schematic diagram of the architecture of a personnel abnormal behavior recognition and early warning system based on the cloud-edge collaboration mechanism artificial intelligence technology provided by an embodiment of the present application. The system includes: a plurality of edge IoT agents and cloud IoT agents. The cloud IoT agent allows driving a multi-modal large model pre-configured in the cloud platform. Among them, The first edge IoT agent is used to obtain personnel image data and environmental data, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result; the first edge IoT agent is any one of the plurality of edge IoT agents; The cloud IoT agent is used to preprocess the personnel image data and the environmental data to obtain first personnel image data and the first environmental data; drive the multi-modal large model to diagnose the first personnel image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, and determine a first early warning parameter according to the first diagnosis result; The first edge IoT agent is used to perform early warning operations and reinforcement learning according to the first diagnosis result and the first early warning parameter.

[0020] In the embodiment of the present application, the personnel abnormal behavior recognition and early warning system based on the cloud-edge collaboration mechanism artificial intelligence technology may include a plurality of edge IoT agents and cloud IoT agents. The cloud IoT agent allows driving a multi-modal large model pre-configured in the cloud platform. As Figure 2 shown, the cloud IoT agent and the cloud platform may be two devices. The cloud IoT agent and the cloud platform are both deployed in the cloud, and a multi-modal large model is configured in the cloud platform. As Figure 3 shown, the cloud IoT agent and the cloud platform may be the same device, that is, a multi-modal large model may be configured in the cloud IoT agent.

[0021] Among them, the multimodal big model can be preset or the system defaults, for example, the multimodal big model can include the deepseek big model, the ChatGpt big model, etc., which are not limited here. The multimodal big model can also include a neural network model, a deep learning model, etc., which are not limited here.

[0022] In specific implementation, the abnormal behavior recognition and early warning system for personnel based on cloud-edge collaborative mechanism artificial intelligence technology can include multiple edge IoT agents, a cloud IoT agent, and a multimodal large model of the smart security scene of the cloud platform. Multiple edge IoT agents perceive on-site personnel and environmental data, conduct intelligent analysis, obtain reasoning results, and can also generate warnings and actions. At the same time, the data (on-site personnel and environmental data) and reasoning results can be transmitted to the cloud IoT agent. The cloud IoT agent pre-processes the data and drives the multimodal large model of the cloud platform to analyze and diagnose the warning events, whether there are abnormalities and abnormal classification, and the action strategy taken. After the cloud IoT agent converts and analyzes, it transmits the feedback information to the corresponding edge IoT agent, and lets the edge IoT agent execute the decision to carry out reinforcement learning.

[0023] Among them, personnel image data can be obtained by cameras, and environmental data can be obtained by environmental sensors.

[0024] The environmental sensor may include at least one of the following: a sound sensor, an odor sensor, a temperature sensor, a humidity sensor, a meteorological sensor, a light sensor, etc., which are not limited here. Correspondingly, the environmental data may include at least one of the following: sound data, odor data, temperature data, humidity data, meteorological data, light brightness data, etc., which are not limited here.

[0025] Among them, Figure 4 As shown, the edge IoT intelligent body can be connected to communicate with multiple sensors, which may include at least one of the following: sound sensor (such as microphone array), odor sensor, temperature sensor, humidity sensor, meteorological sensor, light sensor, camera, electronic fence, infrared dual-detection sensor, access control device, etc., without limitation here.

[0026] In a specific implementation, taking the first edge IoT agent as an example, the first edge IoT agent is any edge IoT agent among multiple edge IoT agents. The first edge IoT agent can obtain personnel image data and environmental data, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result, which may include at least one of the following: whether there is abnormal behavior of personnel, abnormal behavior category, action decision information, etc., which is not limited here.

[0027] In the specific implementation, in the field of behavior recognition, action decision information can be understood as mainly referring to: in the process of identifying and analyzing human movements and behaviors, key information used to judge and determine the specific behavior type. For example, in visual human motion behavior recognition, the positions of key points of the human body and the movement conditions of the movements obtained through the analysis of human posture and movements are all action decision information.

[0028] The first warning parameter may include at least one of the following: warning method, warning level, warning personnel, etc., which are not limited here. Different warning parameters can be pre-set for different abnormal behaviors of different personnel, that is, the mapping relationship between preset abnormal behaviors of personnel and warning parameters can be pre-stored.

[0029] Next, the cloud IoT agent can pre-process the personnel image data and environmental data to obtain the first personnel image data and the first environmental data, and then drive the multimodal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain the first diagnosis result, and determine the first warning parameter based on the first diagnosis result. The first edge IoT agent then performs warning operations and reinforcement learning based on the first diagnosis result and the first warning parameters. On the one hand, the accuracy of the warning operation is improved, and on the other hand, the intelligence of the edge IoT agent can be improved.

[0030] In the embodiment of the present application, each of the multiple edge IoT intelligent agents is an edge IoT intelligent agent with autonomous learning capabilities.

[0031] Among them, the edge IoT agent can be a device with hardware and software entities and analysis capabilities, access data from multiple sensors on site, perform intelligent analysis of sensor data, and execute actions (warning operations) after warnings. The edge IoT agent can have autonomous learning capabilities and use reinforcement learning to learn the best action decisions in a specific environment.

[0032] In the specific implementation, the edge IoT intelligent agent can be responsible for collecting and perceiving various types of on-site personnel and environmental information data. For example, the camera collects on-site video image data, the microphone array collects on-site sound signals, the electronic fence and infrared dual-detection sensors monitor personnel intrusion, and the access control device is responsible for recording personnel entry and exit events, and has strategies for triggering analysis and reasoning.

[0033] Optionally, the multimodal large model includes a first layer architecture and a second layer architecture, the first layer architecture adopts an encoder-decoder architecture to achieve recognition and understanding tasks of sound, image, and text, and outputs first text description information, the first text description information includes text description information for at least one of the sound, image, and text; The second-layer architecture includes an autoregressive Transformer architecture that encodes and decodes the first text description information and outputs the decoded second text description information, which is used to represent at least one of the following: whether there is abnormal behavior of personnel, the category of abnormal behavior, and action decision information.

[0034] Among them, the category of abnormal behavior can be preset or the system default. The setting of the category of abnormal behavior can be based on experience, and different scenarios can correspond to different categories of abnormal behavior. The category of abnormal behavior can include at least one of the following: fighting, brawling, quarreling, chatting, etc., which are not limited here.

[0035] In the embodiment of the present application, the multimodal large model is divided into two-layer network architectures. The first-layer architecture adopts an encoder-decoder architecture (Encoder-Decoder Transformer) to complete the recognition and understanding tasks of sound, image, and text, and outputs the text description information of sound and image. The second-layer architecture adopts an autoregressive Transformer architecture to separately encode and decode the text description information output by the first-layer architecture, and outputs the decoded text description information, which can include whether there is abnormal behavior of personnel, the category of abnormal behavior, and what action decision information should be taken.

[0036] In the embodiment of the present application, the multimodal large model can be trained, deployed, and run on the cloud platform side.

[0037] Optionally, when the environmental data includes sound data and sensor data, in terms of preprocessing the personnel image data and the environmental data to obtain the first personnel image data and the first environmental data, the cloud IoT agent is specifically used for: Processing the sound data into a mel spectrogram; Performing semantic processing on the sensor data to automatically generate the descriptive semantics of structured data to obtain the first descriptive semantics; Determining the first environmental data according to the mel spectrogram and the first descriptive semantics; Performing standardization processing on the personnel image data to obtain the first personnel image data.

[0038] In the embodiments of the present application, when the environmental data includes sound data and sensor data, the sound data can be processed into a mel spectrogram, and the sensor data can be semantically processed to automatically generate a descriptive semantics of the structured data, obtaining a first descriptive semantics. Then, the first environmental data can be determined based on the mel spectrogram and the first descriptive semantics, and the personnel image data can be normalized to obtain the first personnel image data. Thus, the feature quality can be improved, which helps to improve the subsequent diagnosis accuracy. Moreover, due to the preprocessing of the features, the diagnosis speed can be increased.

[0039] For example, the cloud Internet of Things intelligent agent receives the sound, image signals, sensor structured data, and text data of the edge inference result transmitted by the edge Internet of Things intelligent agent. The cloud Internet of Things intelligent agent preprocesses the information, processes the sound signal into a mel spectrogram, normalizes the image data, analyzes the sensor structured data, and automatically generates a descriptive semantics of the structured data. For example, the electronic fence alarm signal automatically generates a semantic description of "suspected intrusion".

[0040] Optionally, in terms of driving the multimodal large model to diagnose the first personnel image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, the cloud Internet of Things intelligent agent is specifically configured to: Determine a first set of relevant prompt word dictionaries according to the first descriptive semantics and the first analysis result; Input the mel spectrogram, the first personnel image data, and the first set of relevant prompt word dictionaries into the multimodal large model for diagnosis to obtain the first diagnosis result.

[0041] In specific implementation, a first set of relevant prompt word dictionaries can be determined according to the first descriptive semantics and the first analysis result, and then the mel spectrogram, the first personnel image data, and the first set of relevant prompt word dictionaries are input into the multimodal large model for diagnosis to obtain the first diagnosis result. For example, specifically, the generated semantic description can be combined with the text data of the edge inference result. For example, if there is a preset abnormal behavior of the personnel, a relevant set of prompt word dictionaries can be generated. After the preprocessing is completed, the preprocessed sound, image data, and relevant prompt word dictionary information are input into the multimodal large model for analysis, that is, auxiliary prompts can be made based on the analysis result of the edge Internet of Things intelligent agent, and on this basis, in-depth analysis is carried out again to deeply determine the authenticity of the abnormal behavior of the personnel, which can deeply improve the intelligence of the intelligent security scenario to ensure public safety.

[0042] For example, after receiving the output analysis and decision-making information of the multi-modal large model, the cloud IoT agent can analyze and compare the semantic similarity between the edge inference result text (the first analysis result) and the multi-modal large model analysis text (the first diagnosis result), evaluate whether the inference of the edge IoT agent is correct or incorrect, transmit the inference evaluation result and the decision-making information to the edge IoT agent through the network, wait for the actions and data feedback of the edge IoT agent, and give corresponding rewards to the edge IoT agent for the impact of this action after receiving the feedback.

[0043] Optionally, the cloud IoT agent is further specifically configured to: Determine a first similarity between the first analysis result and the first diagnosis result; Determine an inference evaluation result of the first edge IoT agent regarding the first analysis result according to the first similarity; Determine a first reward according to the inference evaluation result; Push the first reward to the first edge IoT agent.

[0044] In specific implementation, the first similarity between the first analysis result and the first diagnosis result can be determined. For example, the keywords of the first analysis result can be extracted to obtain the first keywords, the keywords of the first diagnosis result can be extracted to obtain the second keywords, and the similarity between the first keywords and the second keywords can be determined to obtain the first similarity. Another example is that the first target text description information corresponding to the first analysis result can be determined, the second target text description information corresponding to the first diagnosis result can be determined, and the similarity between the first target text description information and the second target text description information can be determined to obtain the first similarity. Another example is that the first vector corresponding to the first analysis result can be determined, the second vector corresponding to the first diagnosis result can be determined, and the similarity between the first vector and the second vector can be determined to obtain the first similarity.

[0045] Then, the inference evaluation result of the first edge IoT agent regarding the first analysis result can be determined according to the first similarity. The mapping relationship between the preset similarity and the inference evaluation result can be stored in advance, and the inference evaluation result corresponding to the first similarity can be determined based on this method. For example, a similarity threshold, that is, a preset value, can be preset. When the first similarity is greater than or equal to the preset value, it can be determined that the inference evaluation result is that the first analysis result is correct. On the contrary, when the first similarity is less than the preset value, it can be determined that the inference evaluation result is that the first analysis result is incorrect.

[0046] Next, the mapping relationship between the preset inference evaluation results and rewards can also be pre-stored. Based on this mapping relationship, the first reward corresponding to the corresponding inference evaluation result can be determined, and then the first reward is pushed to the first edge IoT intelligent agent. The reward provides immediate feedback on the intelligent agent behavior of the first edge IoT intelligent agent, which is used to evaluate the quality of an action in a certain state, thereby affecting its future decisions. Through continuous trial and error and adjustment, the first edge IoT intelligent agent learns the behavior strategies that can obtain high rewards in different states. Furthermore, the intelligence of the intelligent security scenario can be improved to ensure public safety.

[0047] Optionally, the cloud IoT intelligent agent is further specifically configured to: When the inference evaluation result includes that the first analysis result is incorrect, execute the step of determining the first warning parameter according to the first diagnosis result; The first edge IoT intelligent agent is further specifically configured to: When the inference evaluation result includes that the first analysis result is correct, determine the second warning parameter corresponding to the first analysis result, and perform a warning operation according to the second warning parameter.

[0048] In specific implementation, when the inference evaluation result includes that the first analysis result is incorrect, the step of determining the first warning parameter according to the first diagnosis result can be executed, that is, if the first edge IoT intelligent agent makes an incorrect inference, the corresponding warning operation is performed based on the first diagnosis result of the cloud to ensure the accuracy of the warning.

[0049] Among them, the second warning parameter may include at least one of the following: warning method, warning level, warning personnel, etc., which are not limited here. Different abnormal behaviors of personnel can be pre-set with different warning parameters, that is, the mapping relationship between the preset abnormal behaviors of personnel and warning parameters can be pre-stored.

[0050] Correspondingly, when the inference evaluation result includes that the first analysis result is correct, it means that the first edge IoT intelligent agent makes a correct inference, and then the corresponding warning operation can be directly performed based on the first analysis result to improve the warning speed. Different analysis results can correspond to different warning parameters, and then the second warning parameter corresponding to the first analysis result is determined, and the warning operation is performed according to the second warning parameter. Thus, the warning speed can be ensured, and furthermore, the intelligence of the intelligent security scenario can be improved to ensure public safety.

[0051] Optionally, the environmental data includes sound data; in terms of performing intelligent analysis on the personnel image data and the environmental data to obtain the first analysis result, the first edge IoT intelligent agent is specifically configured to: Extract features from the personnel image data to obtain a first feature set; Input the first feature set into the small behavior recognition model to obtain a first recognition result; Extract features from the voice data to obtain a second feature set; Input the second feature set into the small voiceprint recognition model to obtain a second recognition result; Determine the first analysis result according to the first recognition result and the second recognition result.

[0052] Among them, the small behavior recognition model can be preset or the system default, and the small voiceprint recognition model can be preset or the system default.

[0053] In the embodiments of the present application, feature extraction can be performed on the personnel image data to obtain a first feature set, and the first feature set can include at least one of the following: feature points, feature vectors, eigenvalues, feature patterns, etc., which are not limited herein.

[0054] In specific implementation, the first feature set can be input into the small behavior recognition model to obtain a first recognition result, and then, a corresponding image recognition result can be obtained.

[0055] Correspondingly, feature extraction can be performed on the voice data to obtain a second feature set, and the second feature set can include at least one of the following: feature points, feature vectors, eigenvalues, etc., which are not limited herein. Then, the second feature set is input into the small voiceprint recognition model to obtain a second recognition result, and then, a corresponding speech recognition result can be obtained. Then, the first analysis result is determined according to the first recognition result and the second recognition result, that is, the final first analysis result can be comprehensively determined by the image recognition result and the speech recognition result, that is, by combining the image and the environment, the accuracy of personnel abnormal behavior recognition can be guaranteed to a certain extent. Furthermore, the intelligence of the intelligent security scenario can be improved to ensure public safety.

[0056] For example, in specific implementation, the edge IoT intelligent body can quickly identify and reason about the suspicious abnormal behavior of on-site personnel according to a variety of edge-side small models configured. For example, in the case of illegal intrusion on-site, the camera or access control is damaged, and the abnormal environmental sound is collected by the microphone array. The edge IoT intelligent body calls the small voiceprint recognition model to judge the intrusion event, drives the on-site alarm and protection device, and uploads the data and reasoning results to the cloud IoT intelligent body in real time.

[0057] In specific implementation, the edge IoT intelligent body can receive the reasoning results fed back by the cloud IoT intelligent body afterwards, use them to update the managed edge small model sample library, and trigger the edge small model to perform training and learning updates. Receive the action decision information fed back by the cloud IoT intelligent body afterwards, execute the recommended actions, observe the results, collect environmental data again and upload it to the cloud IoT intelligent body in real time, and the cloud IoT intelligent body gives corresponding rewards to the edge IoT intelligent body for the impact of this action after analysis.

[0058] In specific implementation, for the personnel image data, feature extraction is performed to obtain the first feature set, and then the action speed of the personnel can be determined. For example, the mapping relationship between the preset action speed and the feature extraction algorithm can be pre-stored. Furthermore, based on this mapping relationship, the first feature extraction algorithm corresponding to the work speed of the personnel can be determined. The first screen occupancy ratio of the personnel can also be determined, and the mapping relationship between the preset screen occupancy ratio and the control parameters of the first feature extraction algorithm can be pre-stored. Based on this mapping relationship, the first control parameter corresponding to the first screen occupancy ratio can be determined. In this way, on the one hand, the corresponding feature extraction algorithm can be adapted based on the action speed, which helps to accurately capture key actions. On the other hand, the control parameters of the first feature extraction algorithm are adapted based on the screen occupancy ratio, which can capture subtle actions. Furthermore, it helps to improve the accuracy of subsequent abnormal behavior recognition. The first screen occupancy ratio can be understood as the average screen occupancy ratio or the screen occupancy ratio of a specified person. The screen occupancy ratio can be understood as the ratio between the personnel area and the entire image area. The specified person can be pre-set or defaulted by the system, and the specified person can be one or more. The control parameters of the first feature extraction algorithm are used to control the algorithm effect of the first feature extraction algorithm, and the algorithm effect can include at least one of the following: algorithm speed, accuracy, part, feature quality, etc., which are not limited here.

[0059] Optionally, for determining the first analysis result according to the first recognition result and the second recognition result, the first edge IoT agent is specifically used for: Determine the second similarity between the first recognition result and the second recognition result; When the second similarity is greater than the preset similarity, determine the first intersection result of the first recognition result and the second recognition result, and determine the first analysis result according to the first intersection result; When the second similarity is less than or equal to the preset similarity, splice the first feature set and the second feature set to obtain a third feature set; Input the third feature set into the local model of the multi-modal large model to obtain the first analysis result.

[0060] Among them, the preset similarity can be pre-set or defaulted by the system.

[0061] In specific implementation, the second similarity between the first recognition result and the second recognition result can be determined. For example, the keywords of the first recognition result can be extracted to obtain one keyword, and the keywords of the second analysis result can be extracted to obtain another keyword. The similarity between the two keywords can be determined to obtain the second similarity.

[0062] Next, when the second similarity is greater than the preset similarity, the first intersection result of the first recognition result and the second recognition result can be determined, and the first analysis result can be determined according to the first intersection result. That is, by combining the image and the environment, and the recognition results of both are the same, it indicates that abnormal human behaviors are recognized in both dimensions and with a certain degree of accuracy. Then, the results of both can be integrated, which can ensure the accuracy of abnormal human behavior recognition to a certain extent. Furthermore, the intelligence of the intelligent security scenario can be improved to ensure public safety.

[0063] Correspondingly, when the second similarity is less than or equal to the preset similarity, that is, when combining the image and the environment and the recognition results of both are quite different, the first feature set and the second feature set can be concatenated to obtain a third feature set, and then the third feature set is input into the local model of the multi-modal large model to obtain the first analysis result. That is, by combining the image and the environment, the accuracy of abnormal human behavior recognition can be ensured to a certain extent. Furthermore, the intelligence of the intelligent security scenario can be improved to ensure public safety.

[0064] Among them, the local model of the multi-modal large model can be updated at preset time intervals, and the preset time intervals can be set in advance or default by the system.

[0065] Optionally, in terms of inputting the first feature set into the behavior recognition small model to obtain the first recognition result, the first edge IoT agent is specifically used for: Determine the first environmental complexity corresponding to the sound data; Determine the first model parameter corresponding to the first environmental complexity; Determine the first recognition result according to the first feature set, the first model parameter, and the behavior recognition small model.

[0066] Among them, the sound data can be analyzed to obtain multiple sound analysis results. For example, the sound analysis results can include: environmental noise, number of sound sources, voice quality, etc. The mapping relationship between the preset sound analysis results of each dimension and the environmental complexity can be stored in advance. Based on this mapping relationship, the environmental complexity corresponding to each sound analysis result in the multiple sound analysis results can be determined to obtain multiple environmental complexities. Obtain multiple weights corresponding to the multiple sound analysis results, and the sum of the multiple weights is 1. Then, a weighted operation is performed on the multiple environmental complexities and the multiple weights to obtain the first environmental complexity. The multiple weights can be set in advance or default by the system. For example, the multiple weights can be related to time, and the multiple weights can also be related to weather, which is not limited here.

[0067] In specific implementation, the mapping relationship between the preset environmental complexity and the model parameters of the behavior recognition sub-model can be stored in advance. Furthermore, based on this mapping relationship, the first model parameters corresponding to the first environmental complexity can be determined, and then the behavior recognition sub-model can be configured according to the first model parameters. Next, the first feature set is input into the behavior recognition sub-model configured with the first model parameters to obtain the first recognition result. In this way, the model parameters of the behavior recognition sub-model can be optimized by using environmental factors, so that the model ability depth of the behavior recognition sub-model is adapted to the actual environment, which helps to improve the accuracy of abnormal behavior recognition of the behavior recognition sub-model. For example, if the abnormal behavior of a person is fighting and corresponding quarrels are detected by sound, the weight of the abnormal behavior of "fighting" of the person can be adjusted. Thus, it helps to further ensure the accuracy of abnormal behavior recognition of the person, that is, by deeply combining the image and the environment, the accuracy of abnormal behavior recognition of the person can be ensured to a certain extent. Furthermore, the intelligence of the intelligent security scenario can be improved to ensure public safety.

[0068] Optionally, the first recognition result includes a first abnormal behavior type and a first abnormal probability value; in terms of inputting the second feature set into the voiceprint recognition sub-model to obtain the second recognition result, the first edge IoT intelligent body is specifically used for: Determine the second model parameters corresponding to the first abnormal behavior type; Determine the first adjustment parameter corresponding to the first abnormal probability value; Adjust the second model parameters according to the first adjustment parameter to obtain the target second model parameters; Determine the second recognition result according to the second feature set, the target second model parameters, and the voiceprint recognition sub-model.

[0069] In specific implementation, the first recognition result includes a first abnormal behavior type and a first abnormal probability value. The first abnormal behavior type may include at least one abnormal behavior of a person, and the first abnormal probability value may include at least one probability value, and each abnormal behavior of a person corresponds to a probability value.

[0070] Next, the mapping relationship between the preset abnormal behavior types and the model parameters of the small voiceprint recognition model can also be pre-stored. Furthermore, based on this mapping relationship, the second model parameters corresponding to the first abnormal behavior type can be determined. The mapping relationship between the preset probability values and the adjustment parameters can also be pre-stored. Based on this mapping relationship, the first adjustment parameter corresponding to the first abnormal probability value can be determined. Then, according to the first adjustment parameter, some or all of the model parameters of the second model parameters are adjusted to obtain the target second model parameters. For example, the target second model parameters = (1 + the first adjustment parameter) × the second model parameters. Then, the small voiceprint recognition model is configured according to the target second model parameters, and the second feature set is input into the small voiceprint recognition model configured with the target second model parameters to obtain the second recognition result.

[0071] In this way, the recognition results in the image dimension can be used to optimize the model parameters of the small voiceprint recognition model in turn, so that the model ability of the small voiceprint recognition model conforms to the actual environment depth, which helps to improve the accuracy of abnormal behavior recognition of the small voiceprint recognition model. That is, by combining the image and the environment depth, the accuracy of personnel abnormal behavior recognition can be guaranteed to a certain extent. Furthermore, the intelligence of the intelligent security scenario can be improved to ensure public safety.

[0072] It can be seen that the personnel abnormal behavior recognition and early warning system based on the cloud-edge collaborative mechanism artificial intelligence technology described in the embodiments of the present application includes: multiple edge IoT intelligent agents and a cloud IoT intelligent agent. The cloud IoT intelligent agent is allowed to drive a multi-modal large model pre-configured in the cloud platform. Among them, the first edge IoT intelligent agent obtains personnel image data and environmental data, and performs intelligent analysis on the personnel image data and environmental data to obtain the first analysis result. The first edge IoT intelligent agent is any one of the multiple edge IoT intelligent agents. The cloud IoT intelligent agent pre-processes the personnel image data and environmental data to obtain the first personnel image data and the first environmental data. The multi-modal large model is driven to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain the first diagnosis result, and the first early warning parameter is determined according to the first diagnosis result. The first edge IoT intelligent agent performs early warning operations and reinforcement learning according to the first diagnosis result and the first early warning parameter. On the one hand, the accuracy of the early warning operation is improved. On the other hand, the intelligence of the edge IoT intelligent agent can be improved. Furthermore, the intelligence of the intelligent security scenario can be improved to ensure public safety.

[0073] Please refer to Figure 5 , Figure 5 is a schematic flowchart of a method for recognizing and warning personnel abnormal behavior based on the cloud-edge collaborative mechanism artificial intelligence technology provided by the embodiments of the present application, and is applied to such as Figure 1The described personnel abnormal behavior recognition and early warning system based on the cloud-edge collaborative mechanism artificial intelligence technology, the system includes: a plurality of edge IoT intelligent agents, a cloud IoT intelligent agent, and the cloud IoT intelligent agent is allowed to drive a multi-modal large model pre-configured in the cloud platform; the personnel abnormal behavior recognition and early warning method based on the cloud-edge collaborative mechanism artificial intelligence technology includes: 501. Obtain personnel image data and environmental data through the first edge IoT intelligent agent, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result; the first edge IoT intelligent agent is any one of the plurality of edge IoT intelligent agents; 502. Preprocess the personnel image data and the environmental data through the cloud IoT intelligent agent to obtain first personnel image data and first environmental data; drive the multi-modal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain a first diagnosis result, and determine a first early warning parameter according to the first diagnosis result; 503. Perform early warning operations and reinforcement learning through the first edge IoT intelligent agent according to the first diagnosis result and the first early warning parameter.

[0074] Optionally, the multi-modal large model includes a first-layer architecture and a second-layer architecture. The first-layer architecture adopts an encoder-decoder architecture to implement the recognition and understanding tasks of sound, image, and text, and outputs first text description information, and the first text description information includes text description information for at least one of sound, image, and text; The second-layer architecture includes an autoregressive Transormer architecture, encodes and decodes the first text description information, and outputs the decoded second text description information, and the second text description information is used to represent at least one of the following contents: whether there is a personnel abnormal behavior, the category of abnormal behavior, and action decision information.

[0075] Optionally, when the environmental data includes sound data and sensor data, in step 502 above, preprocessing the personnel image data and the environmental data to obtain first personnel image data and first environmental data can be implemented as follows: Process the sound data into a mel spectrogram; Perform semantic processing on the sensor data to automatically generate a description semantics of structured data to obtain a first description semantics; Determine the first environmental data according to the mel spectrogram and the first description semantics; Perform normalization processing on the personnel image data to obtain the first personnel image data.

[0076] Optionally, for step 502 above, driving the multi-modal large model to diagnose the first personnel image data, the first environmental data, and the first analysis result to obtain a first diagnosis result can be implemented as follows: Determine a first set of relevant prompt word dictionaries according to the first descriptive semantics and the first analysis result; Input the mel spectrogram, the first personnel image data, and the first set of relevant prompt word dictionaries into the multi-modal large model for diagnosis to obtain the first diagnosis result.

[0077] Optionally, the following steps may further be included: Determine a first similarity between the first analysis result and the first diagnosis result; Determine an inference evaluation result of the first edge IoT agent regarding the first analysis result according to the first similarity; Determine a first reward according to the inference evaluation result; Push the first reward to the first edge IoT agent.

[0078] Optionally, the following steps may further be included: When the inference evaluation result includes that the first analysis result is incorrect, execute the step of determining the first warning parameter according to the first diagnosis result; When the inference evaluation result includes that the first analysis result is correct, determine a second warning parameter corresponding to the first analysis result through the first edge IoT agent, and perform a warning operation according to the second warning parameter.

[0079] Optionally, the environmental data includes sound data; for step 501 above, performing intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result can be implemented as follows: Extract features from the personnel image data to obtain a first feature set; Input the first feature set into a small behavior recognition model to obtain a first recognition result; Extract features from the sound data to obtain a second feature set; Input the second feature set into a small voiceprint recognition model to obtain a second recognition result; Determine the first analysis result according to the first recognition result and the second recognition result.

[0080] Optionally, for the above step of determining the first analysis result according to the first recognition result and the second recognition result, it can be implemented as follows: Determine a second similarity between the first recognition result and the second recognition result; When the second similarity is greater than the preset similarity, determine the first intersection result of the first recognition result and the second recognition result, and determine the first analysis result according to the first intersection result; When the second similarity is less than or equal to the preset similarity, splice the first feature set and the second feature set to obtain a third feature set; Input the third feature set into the local model of the multi-modal large model to obtain the first analysis result.

[0081] Optionally, the above step of inputting the first feature set into the behavior recognition small model to obtain the first recognition result can be implemented as follows: Determine the first environmental complexity corresponding to the voice data; Determine the first model parameter corresponding to the first environmental complexity; Determine the first recognition result according to the first feature set, the first model parameter and the behavior recognition small model.

[0082] Optionally, the first recognition result includes a first abnormal behavior type and a first abnormal probability value; the above step of inputting the second feature set into the voiceprint recognition small model to obtain the second recognition result can be implemented as follows: Determine the second model parameter corresponding to the first abnormal behavior type; Determine the first adjustment parameter corresponding to the first abnormal probability value; Adjust the second model parameter according to the first adjustment parameter to obtain the target second model parameter; Determine the second recognition result according to the second feature set, the target second model parameter and the voiceprint recognition small model.

[0083] Among them, the specific relevant descriptions of the above steps can refer to the corresponding parts in the embodiments described above Figure 1 and will not be repeated here.

[0084] It can be seen that the method for identifying and warning of abnormal human behaviors based on the cloud-edge collaboration mechanism artificial intelligence technology described in the embodiments of the present application is applied to a system for identifying and warning of abnormal human behaviors based on the cloud-edge collaboration mechanism artificial intelligence technology. The system includes: a plurality of edge IoT agents and a cloud IoT agent. The cloud IoT agent is allowed to drive a multi-modal large model pre-configured in the cloud platform. Among them, the first edge IoT agent acquires human image data and environmental data, and performs intelligent analysis on the human image data and environmental data to obtain a first analysis result. The first edge IoT agent is any one of the plurality of edge IoT agents. The cloud IoT agent pre-processes the human image data and environmental data to obtain first human image data and first environmental data. The multi-modal large model is driven to diagnose the first human image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, and a first warning parameter is determined according to the first diagnosis result. The first edge IoT agent performs a warning operation and reinforcement learning according to the first diagnosis result and the first warning parameter. On the one hand, the accuracy of the warning operation is improved. On the other hand, the intelligence of the edge IoT agent can be improved. Furthermore, the intelligence of the intelligent security scenario can be improved to ensure public safety.

[0085] Consistent with the above embodiments, please refer to Figure 6 , Figure 6 FIG. is a schematic structural diagram of an electronic device provided in an embodiment of the present application. As shown in the figure, the electronic device includes a processor, a memory, a communication interface, and one or more programs. Among them, the above one or more programs are stored in the above memory and are configured to be executed by the above processor. In the embodiments of the present application, it is applied to a system for identifying and warning of abnormal human behaviors based on the cloud-edge collaboration mechanism artificial intelligence technology. The system includes: a plurality of edge IoT agents and a cloud IoT agent. The cloud IoT agent is allowed to drive a multi-modal large model pre-configured in the cloud platform. The above programs include instructions for performing the following steps: Acquire human image data and environmental data through the first edge IoT agent, and perform intelligent analysis on the human image data and the environmental data to obtain a first analysis result. The first edge IoT agent is any one of the plurality of edge IoT agents. Pre-process the human image data and the environmental data through the cloud IoT agent to obtain first human image data and first environmental data. Drive the multi-modal large model to diagnose the first human image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, and determine a first warning parameter according to the first diagnosis result. Perform a warning operation and reinforcement learning through the first edge IoT agent according to the first diagnosis result and the first warning parameter.

[0086] Optionally, the multi-modal large model includes a first-layer architecture and a second-layer architecture. The first-layer architecture adopts an encoder-decoder architecture to implement the recognition and understanding tasks of sound, image, and text, and outputs first text description information, which includes text description information for at least one of sound, image, and text. The second-layer architecture includes an autoregressive Transformer architecture, which encodes and decodes the first text description information and outputs the decoded second text description information. The second text description information is used to represent at least one of the following contents: whether there is abnormal behavior of personnel, the category of abnormal behavior, and action decision information.

[0087] Optionally, when the environmental data includes sound data and sensor data, in the aspect of preprocessing the personnel image data and the environmental data to obtain first personnel image data and first environmental data, the above program includes instructions for performing the following steps: Process the sound data into a mel spectrogram; Semantically process the sensor data to automatically generate a descriptive semantics of structured data, obtaining a first descriptive semantics; Determine the first environmental data according to the mel spectrogram and the first descriptive semantics; Normalize the personnel image data to obtain the first personnel image data.

[0088] Optionally, in the aspect of driving the multi-modal large model to diagnose the first personnel image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, the above program includes instructions for performing the following steps: Determine a first set of relevant prompt word dictionaries according to the first descriptive semantics and the first analysis result; Input the mel spectrogram, the first personnel image data, and the first set of relevant prompt word dictionaries into the multi-modal large model for diagnosis to obtain the first diagnosis result.

[0089] Optionally, the above program further includes instructions for performing the following steps: Determine a first similarity between the first analysis result and the first diagnosis result; Determine an inference evaluation result of the first edge IoT agent regarding the first analysis result according to the first similarity; Determine a first reward according to the inference evaluation result; Push the first reward to the first edge IoT agent.

[0090] Optionally, the above program further includes instructions for performing the following steps: When the inference evaluation result includes an error in the first analysis result, perform the step of determining a first warning parameter according to the first diagnosis result; When the inference evaluation result includes that the first analysis result is correct, determine a second warning parameter corresponding to the first analysis result through the first edge IoT agent, and perform a warning operation according to the second warning parameter.

[0091] Optionally, the environmental data includes sound data; in terms of performing intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result, the above program includes instructions for performing the following steps: Extract features from the personnel image data to obtain a first feature set; Input the first feature set into a small behavior recognition model to obtain a first recognition result; Extract features from the sound data to obtain a second feature set; Input the second feature set into a small voiceprint recognition model to obtain a second recognition result; Determine the first analysis result according to the first recognition result and the second recognition result.

[0092] Optionally, in terms of determining the first analysis result according to the first recognition result and the second recognition result, the above program includes instructions for performing the following steps: Determine a second similarity between the first recognition result and the second recognition result; When the second similarity is greater than a preset similarity, determine a first intersection result between the first recognition result and the second recognition result, and determine the first analysis result according to the first intersection result; When the second similarity is less than or equal to the preset similarity, splice the first feature set and the second feature set to obtain a third feature set; Input the third feature set into the local model of the multi-modal large model to obtain the first analysis result.

[0093] Optionally, in terms of inputting the first feature set into a small behavior recognition model to obtain a first recognition result, the above program includes instructions for performing the following steps: Determine a first environmental complexity corresponding to the sound data; Determine a first model parameter corresponding to the first environmental complexity; Determine the first recognition result according to the first feature set, the first model parameter, and the small behavior recognition model.

[0094] Optionally, the first recognition result includes a first abnormal behavior type and a first abnormal probability value; in terms of inputting the second feature set into the voiceprint recognition small model to obtain a second recognition result, the above program includes instructions for performing the following steps: Determine a second model parameter corresponding to the first abnormal behavior type; Determine a first adjustment parameter corresponding to the first abnormal probability value; Adjust the second model parameter according to the first adjustment parameter to obtain a target second model parameter; Determine the second recognition result according to the second feature set, the target second model parameter, and the voiceprint recognition small model.

[0095] It can be seen that the electronic device described in the embodiments of the present application is applied to a personnel abnormal behavior recognition and early warning system based on the cloud-edge collaboration mechanism artificial intelligence technology. The system includes: multiple edge IoT intelligent agents and a cloud IoT intelligent agent. The cloud IoT intelligent agent allows driving a multi-modal large model pre-configured in the cloud platform; wherein, the first edge IoT intelligent agent acquires personnel image data and environmental data, and performs intelligent analysis on the personnel image data and environmental data to obtain a first analysis result; the first edge IoT intelligent agent is any one of the multiple edge IoT intelligent agents. The cloud IoT intelligent agent preprocesses the personnel image data and environmental data to obtain first personnel image data and first environmental data; drives the multi-modal large model to diagnose the first personnel image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, and determines a first early warning parameter according to the first diagnosis result. The first edge IoT intelligent agent performs early warning operations and reinforcement learning according to the first diagnosis result and the first early warning parameter. On the one hand, the accuracy of the early warning operation is improved. On the other hand, the intelligence of the edge IoT intelligent agent can be improved. Furthermore, the intelligence of the intelligent security scenario can be improved to ensure public safety.

[0096] Among them, the electronic device may include a controller of the edge IoT intelligent agent, or an edge IoT intelligent agent, or a controller of the cloud IoT intelligent agent, or a cloud IoT intelligent agent.

[0097] Figure 7It is a functional unit composition block diagram of a personnel abnormal behavior recognition and early warning device 700 based on cloud-edge collaborative mechanism artificial intelligence technology involved in an embodiment of the present application. The personnel abnormal behavior recognition and early warning device 700 based on cloud-edge collaborative mechanism artificial intelligence technology is applied to a personnel abnormal behavior recognition and early warning system based on cloud-edge collaborative mechanism artificial intelligence technology. The system includes: a plurality of edge IoT agents and a cloud IoT agent. The cloud IoT agent allows driving a pre-configured multi-modal large model in the cloud platform. The personnel abnormal behavior recognition and early warning device 700 based on cloud-edge collaborative mechanism artificial intelligence technology includes: an analysis unit 701, a diagnosis unit 702, and a feedback unit 703. Among them, The analysis unit 701 is configured to obtain personnel image data and environmental data through a first edge IoT agent, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result. The first edge IoT agent is any one of the plurality of edge IoT agents; The diagnosis unit 702 is configured to preprocess the personnel image data and the environmental data through the cloud IoT agent to obtain first personnel image data and first environmental data; drive the multi-modal large model to diagnose the first personnel image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, and determine a first early warning parameter according to the first diagnosis result; The feedback unit 703 is configured to perform an early warning operation and reinforcement learning according to the first diagnosis result and the first early warning parameter through the first edge IoT agent.

[0098] Optionally, the multi-modal large model includes a first-layer architecture and a second-layer architecture. The first-layer architecture adopts an encoder-decoder architecture to implement tasks of recognizing and understanding sound, image, and text, and outputs first text description information. The first text description information includes text description information for at least one of sound, image, and text; The second-layer architecture includes an autoregressive Transformer architecture, encodes and decodes the first text description information, and outputs decoded second text description information. The second text description information is used to represent at least one of the following contents: whether there is a personnel abnormal behavior, the category of abnormal behavior, and action decision information.

[0099] Optionally, when the environmental data includes sound data and sensor data, in terms of preprocessing the personnel image data and the environmental data to obtain first personnel image data and first environmental data, the diagnosis unit 702 is specifically configured to: Process the sound data into a mel spectrogram; Semantically process the sensor data to automatically generate a description semantics of the structured data, obtaining a first description semantics; Determine the first environmental data according to the mel spectrogram and the first description semantics; Standardize the personnel image data to obtain the first personnel image data.

[0100] Optionally, in the aspect of driving the multimodal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain a first diagnosis result, the diagnosis unit 702 is specifically configured to: Determine a first set of relevant prompt word dictionaries according to the first description semantics and the first analysis result; Input the mel spectrogram, the first personnel image data and the first set of relevant prompt word dictionaries into the multimodal large model for diagnosis to obtain the first diagnosis result.

[0101] Optionally, the personnel abnormal behavior recognition and warning 700 based on the cloud-edge collaborative mechanism artificial intelligence technology is further specifically configured to: Determine a first similarity between the first analysis result and the first diagnosis result; Determine an inference evaluation result of the first edge IoT agent regarding the first analysis result according to the first similarity; Determine a first reward according to the inference evaluation result; Push the first reward to the first edge IoT agent.

[0102] Optionally, the personnel abnormal behavior recognition and warning 700 based on the cloud-edge collaborative mechanism artificial intelligence technology is further specifically configured to: When the inference evaluation result includes that the first analysis result is incorrect, execute the step of determining the first warning parameter according to the first diagnosis result; When the inference evaluation result includes that the first analysis result is correct, determine a second warning parameter corresponding to the first analysis result through the first edge IoT agent, and perform a warning operation according to the second warning parameter.

[0103] Optionally, the environmental data includes sound data; in the aspect of intelligently analyzing the personnel image data and the environmental data to obtain a first analysis result, the analysis unit 701 is specifically configured to: Extract features from the personnel image data to obtain a first feature set; Input the first feature set into a behavior recognition small model to obtain a first recognition result; Extract features from the sound data to obtain a second feature set; Input the second feature set into the voiceprint recognition small model to obtain a second recognition result; Determine the first analysis result according to the first recognition result and the second recognition result.

[0104] Optionally, when determining the first analysis result according to the first recognition result and the second recognition result, the analysis unit 701 is specifically configured to: Determine a second similarity between the first recognition result and the second recognition result; When the second similarity is greater than a preset similarity, determine a first intersection result of the first recognition result and the second recognition result, and determine the first analysis result according to the first intersection result; When the second similarity is less than or equal to the preset similarity, splice the first feature set and the second feature set to obtain a third feature set; Input the third feature set into the local model of the multimodal large model to obtain the first analysis result.

[0105] Optionally, in terms of inputting the first feature set into the behavior recognition small model to obtain a first recognition result, the analysis unit 701 is specifically configured to: Determine a first environmental complexity corresponding to the voice data; Determine a first model parameter corresponding to the first environmental complexity; Determine the first recognition result according to the first feature set, the first model parameter, and the behavior recognition small model.

[0106] Optionally, the first recognition result includes a first abnormal behavior type and a first abnormal probability value; in terms of inputting the second feature set into the voiceprint recognition small model to obtain a second recognition result, the analysis unit 701 is specifically configured to: Determine a second model parameter corresponding to the first abnormal behavior type; Determine a first adjustment parameter corresponding to the first abnormal probability value; Adjust the second model parameter according to the first adjustment parameter to obtain a target second model parameter; Determine the second recognition result according to the second feature set, the target second model parameter, and the voiceprint recognition small model.

[0107] It can be seen that the device for identifying and warning of abnormal human behaviors based on the artificial intelligence technology of the cloud-edge collaboration mechanism described in the embodiments of the present application is applied to a system for identifying and warning of abnormal human behaviors based on the artificial intelligence technology of the cloud-edge collaboration mechanism. The system includes: a plurality of edge IoT agents and a cloud IoT agent. The cloud IoT agent is allowed to drive a multi-modal large model pre-configured in the cloud platform. Among them, the first edge IoT agent acquires human image data and environmental data, and performs intelligent analysis on the human image data and environmental data to obtain a first analysis result. The first edge IoT agent is any one of the plurality of edge IoT agents. The cloud IoT agent pre-processes the human image data and environmental data to obtain first human image data and first environmental data. The multi-modal large model is driven to diagnose the first human image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, and first warning parameters are determined according to the first diagnosis result. The first edge IoT agent performs warning operations and reinforcement learning according to the first diagnosis result and the first warning parameters. On the one hand, the accuracy of the warning operation is improved. On the other hand, the intelligence of the edge IoT agent can be improved. Furthermore, the intelligence of the intelligent security scenario can be improved to ensure public safety.

[0108] It can be understood that the functions of the device for identifying and warning of abnormal human behaviors based on the artificial intelligence technology of the cloud-edge collaboration mechanism in this embodiment can be specifically implemented according to the methods in the above method embodiments. The specific implementation process can refer to the relevant descriptions in the above method embodiments and will not be elaborated here.

[0109] The embodiments of the present application further provide a computer storage medium. The computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute some or all of the steps of any of the methods recorded in the above method embodiments. The above computer includes an electronic device.

[0110] The embodiments of the present application further provide a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to enable a computer to execute some or all of the steps of any of the methods recorded in the above method embodiments. The computer program product can be a software installation package, and the above computer includes an electronic device.

[0111] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0112] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0113] In several embodiments provided in the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in an electrical or other form.

[0114] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0115] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0116] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in each embodiment of the present application. And the aforementioned memory includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical disks and other various media that can store program codes.

[0117] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable memory. The memory may include: a flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, an optical disc, etc.

[0118] The above has introduced the embodiments of the present application in detail. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present application.

Claims

1. An abnormal behavior recognition and early warning system based on a cloud-edge collaboration mechanism, characterized in that, The system includes: multiple edge Internet of Things agents and a cloud Internet of Things agent, and the cloud Internet of Things agent allows to drive a pre-configured multi-modal large model in the cloud platform; wherein, The first edge Internet of Things agent is used to obtain personnel image data and environmental data, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result; the first edge Internet of Things agent is any one of the multiple edge Internet of Things agents; The cloud Internet of Things agent is used to preprocess the personnel image data and the environmental data to obtain first personnel image data and first environmental data; drive the multi-modal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain a first diagnosis result, and determine a first warning parameter according to the first diagnosis result; The first edge Internet of Things agent is used to perform a warning operation and reinforcement learning according to the first diagnosis result and the first warning parameter.

2. The abnormal behavior recognition and warning system based on the cloud-edge collaboration mechanism according to claim 1, characterized in that, The multi-modal large model includes a first-layer architecture and a second-layer architecture. The first-layer architecture adopts an encoder-decoder architecture to implement the recognition and understanding tasks of sound, image, and text, and outputs first text description information, where the first text description information includes text description information for at least one of sound, image, and text; The second-layer architecture includes an autoregressive Transormer architecture, encodes and decodes the first text description information, and outputs decoded second text description information, where the second text description information is used to represent at least one of the following contents: whether there is an abnormal behavior of a person, the category of abnormal behavior, and action decision information.

3. The abnormal behavior recognition and early warning system based on the cloud-edge collaboration mechanism according to claim 1 or 2, characterized in that When the environmental data includes sound data and sensor data, in terms of preprocessing the personnel image data and the environmental data to obtain first personnel image data and first environmental data, the cloud Internet of Things agent specifically is used for: Processing the sound data into a mel spectrogram; Performing semantic processing on the sensor data to automatically generate a descriptive semantics of structured data to obtain a first descriptive semantics; Determining the first environmental data according to the mel spectrogram and the first descriptive semantics; Performing normalization processing on the personnel image data to obtain the first personnel image data.

4. The abnormal behavior recognition and early warning system based on the cloud-edge collaboration mechanism according to claim 3, characterized in that, In terms of driving the multi-modal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain a first diagnosis result, the cloud Internet of Things agent specifically is used for: Determining a first set of relevant prompt word dictionaries according to the first descriptive semantics and the first analysis result; Inputting the mel spectrogram, the first personnel image data and the first set of relevant prompt word dictionaries into the multi-modal large model for diagnosis to obtain the first diagnosis result.

5. The abnormal behavior recognition and early warning system based on the cloud-edge collaboration mechanism according to claim 1 or 2, characterized in that, The cloud Internet of Things agent is also specifically used for: Determining a first similarity between the first analysis result and the first diagnosis result; Determining an inference evaluation result of the first edge Internet of Things agent regarding the first analysis result according to the first similarity; Determining a first reward according to the inference evaluation result; Push the first reward to the first edge Internet of Things intelligent agent.

6. The abnormal behavior recognition and warning system based on the cloud-edge collaboration mechanism according to claim 5, characterized in that, The cloud Internet of Things intelligent agent is further specifically configured to: When the inference evaluation result includes that the first analysis result is incorrect, execute the step of determining the first early warning parameter according to the first diagnosis result; The first edge Internet of Things intelligent agent is further specifically configured to: When the inference evaluation result includes that the first analysis result is correct, determine a second early warning parameter corresponding to the first analysis result, and perform an early warning operation according to the second early warning parameter.

7. The abnormal behavior recognition and warning system based on the cloud-edge collaboration mechanism according to claim 1 or 2, characterized in that, The environmental data includes sound data; in terms of performing intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result, the first edge Internet of Things intelligent agent is specifically configured to: Extract features from the personnel image data to obtain a first feature set; Input the first feature set into the behavior recognition small model to obtain a first recognition result; Extract features from the sound data to obtain a second feature set; Input the second feature set into the voiceprint recognition small model to obtain a second recognition result; Determine the first analysis result according to the first recognition result and the second recognition result.

8. The abnormal behavior recognition and early warning system based on the cloud-edge collaboration mechanism according to claim 7, wherein For determining the first analysis result according to the first recognition result and the second recognition result, the first edge Internet of Things intelligent agent is specifically configured to: Determine a second similarity between the first recognition result and the second recognition result; When the second similarity is greater than the preset similarity, determine a first intersection result of the first recognition result and the second recognition result, and determine the first analysis result according to the first intersection result; When the second similarity is less than or equal to the preset similarity, splice the first feature set and the second feature set to obtain a third feature set; Input the third feature set into the local model of the multi-modal large model to obtain the first analysis result.

9. The abnormal behavior recognition and warning system based on the cloud-edge collaboration mechanism according to claim 7, characterized in that, In terms of inputting the first feature set into the behavior recognition small model to obtain a first recognition result, the first edge Internet of Things intelligent agent is specifically configured to: Determine a first environmental complexity corresponding to the sound data; Determine a first model parameter corresponding to the first environmental complexity; Determine the first recognition result according to the first feature set, the first model parameter, and the behavior recognition small model.

10. The abnormal behavior recognition and early warning system based on the cloud-edge collaboration mechanism according to claim 9, characterized in that, The first recognition result includes a first abnormal behavior type and a first abnormal probability value; in terms of inputting the second feature set into the voiceprint recognition small model to obtain a second recognition result, the first edge Internet of Things intelligent agent is specifically configured to: Determine a second model parameter corresponding to the first abnormal behavior type; Determine a first adjustment parameter corresponding to the first abnormal probability value; Adjust the second model parameter according to the first adjustment parameter to obtain a target second model parameter; Determine the second recognition result according to the second feature set, the target second model parameter, and the voiceprint recognition small model.

11. An abnormal behavior recognition and early warning method based on a cloud-edge collaboration mechanism, characterized in that, Applied to an abnormal behavior recognition and early warning system based on a cloud-edge collaboration mechanism, the system includes: multiple edge Internet of Things intelligent agents, a cloud Internet of Things intelligent agent, and the cloud Internet of Things intelligent agent is allowed to drive a multi-modal large model pre-configured in the cloud platform; the method includes: Obtain personnel image data and environmental data through the first edge IoT intelligent agent, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result; the first edge IoT intelligent agent is any one of the multiple edge IoT intelligent agents; Preprocess the personnel image data and the environmental data through the cloud IoT intelligent agent to obtain first personnel image data and first environmental data; drive the multi-modal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain a first diagnosis result, and determine a first warning parameter according to the first diagnosis result; Perform a warning operation and reinforcement learning through the first edge IoT intelligent agent according to the first diagnosis result and the first warning parameter.

12. The abnormal behavior recognition and early warning method based on the cloud-edge collaboration mechanism according to claim 11, wherein The multi-modal large model includes a first-layer architecture and a second-layer architecture. The first-layer architecture adopts an encoder-decoder architecture to implement the recognition and understanding tasks of sound, image, and text, and outputs first text description information, where the first text description information includes text description information for at least one of sound, image, and text; The second-layer architecture includes an autoregressive Transormer architecture, encodes and decodes the first text description information, and outputs the decoded second text description information, where the second text description information is used to represent at least one of the following: whether there is an abnormal behavior of a person, the category of abnormal behavior, and action decision information.

13. The method for abnormal behavior recognition and early warning based on the cloud-edge collaboration mechanism according to claim 11 or 12, characterized in that, When the environmental data includes sound data and sensor data, the preprocessing of the personnel image data and the environmental data to obtain first personnel image data and first environmental data includes: Process the sound data into a mel spectrogram; Perform semantic processing on the sensor data to automatically generate a descriptive semantics of structured data to obtain a first descriptive semantics; Determine the first environmental data according to the mel spectrogram and the first descriptive semantics; Perform normalization processing on the personnel image data to obtain the first personnel image data.

14. The abnormal behavior recognition and early warning method based on the cloud-edge collaboration mechanism according to claim 13, characterized in that, The driving the multi-modal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain a first diagnosis result includes: Determine a first set of relevant prompt word dictionaries according to the first descriptive semantics and the first analysis result; Input the mel spectrogram, the first personnel image data and the first set of relevant prompt word dictionaries into the multi-modal large model for diagnosis to obtain the first diagnosis result.

15. The method for abnormal behavior recognition and early warning based on the cloud-edge collaboration mechanism according to claim 11 or 12, characterized in that The abnormal behavior recognition and warning method based on the cloud-edge collaboration mechanism further includes: Determine a first similarity between the first analysis result and the first diagnosis result; Determine an inference evaluation result of the first edge IoT intelligent agent regarding the first analysis result according to the first similarity; Determine a first reward according to the inference evaluation result; Push the first reward to the first edge IoT intelligent agent.

16. The abnormal behavior recognition and warning method based on the cloud-edge collaboration mechanism according to claim 15, characterized in that, The abnormal behavior recognition and warning method based on the cloud-edge collaboration mechanism further includes: When the inference evaluation result includes that the first analysis result is incorrect, execute the step of determining the first warning parameter according to the first diagnosis result; When the inference evaluation result includes that the first analysis result is correct, determine, by the first edge Internet of Things intelligent agent, a second warning parameter corresponding to the first analysis result, and perform a warning operation according to the second warning parameter.

17. The method for abnormal behavior recognition and early warning based on the cloud-edge collaboration mechanism according to claim 11 or 12, wherein The environmental data includes sound data; the intelligent analysis of the personnel image data and the environmental data to obtain a first analysis result includes: Extract features from the personnel image data to obtain a first feature set; Input the first feature set into a small behavior recognition model to obtain a first recognition result; Extract features from the sound data to obtain a second feature set; Input the second feature set into a small voiceprint recognition model to obtain a second recognition result; Determine the first analysis result according to the first recognition result and the second recognition result.

18. The method for abnormal behavior recognition and early warning based on the cloud-edge collaboration mechanism according to claim 17, wherein The determining the first analysis result according to the first recognition result and the second recognition result includes: Determine a second similarity between the first recognition result and the second recognition result; When the second similarity is greater than a preset similarity, determine a first intersection result of the first recognition result and the second recognition result, and determine the first analysis result according to the first intersection result; When the second similarity is less than or equal to the preset similarity, splice the first feature set and the second feature set to obtain a third feature set; Input the third feature set into the local model of the multi-modal large model to obtain the first analysis result.

19. The abnormal behavior recognition and early warning method based on the cloud-edge collaboration mechanism according to claim 17, characterized in that, The inputting the first feature set into a small behavior recognition model to obtain a first recognition result includes: Determine a first environmental complexity corresponding to the sound data; Determine a first model parameter corresponding to the first environmental complexity; Determine the first recognition result according to the first feature set, the first model parameter, and the small behavior recognition model.

20. The method for abnormal behavior recognition and early warning based on the cloud-edge collaboration mechanism according to claim 19, wherein The first recognition result includes a first abnormal behavior type and a first abnormal probability value; the inputting the second feature set into a small voiceprint recognition model to obtain a second recognition result includes: Determine a second model parameter corresponding to the first abnormal behavior type; Determine a first adjustment parameter corresponding to the first abnormal probability value; Adjust the second model parameter according to the first adjustment parameter to obtain a target second model parameter; Determine the second recognition result according to the second feature set, the target second model parameter, and the small voiceprint recognition model.

Citation Information

Patent Citations

  • Human body behavior recognition and data acquisition system based on artificial intelligence

    CN118430070A

  • Intelligent construction site construction safety monitoring cloud edge collaborative early warning system based on edge mobile monitoring station and control system and control method of intelligent construction site construction safety monitoring cloud edge collaborative early warning system

    CN118433231A

  • Intelligent construction safety management system and method based on large-scale multi-modal language model

    CN118735732A

  • Abnormity detection and response method and system based on intelligent perception driving

    CN119544388A

  • Storage security protection and fire-fighting early warning method based on edge AI mode

    CN119694065A

Cited By

  • Household anti-fraud real-time semantic recognition method and system

    CN120998186A

  • Device predictive maintenance system and method based on large model and agent cooperation

    CN121509265A

  • Device predictive maintenance system and method based on large model and agent cooperation

    CN121509265B