Abnormal behavior identification and early warning method and related system based on cloud-edge collaboration mechanism

Through the abnormal behavior recognition and early warning system of the cloud-edge collaboration mechanism, the coordinated work of edge IoT agents and cloud IoT agents is used to achieve high-precision identification and early warning of abnormal behaviors in smart security scenarios, improving the intelligence and security of the system.

CN120337107BActive Publication Date: 2025-09-02SICHUAN RUITING ZHIHUI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510825189.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-02
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

In smart security scenarios, existing identification and early warning technologies lack intelligence, resulting in reduced security.

Method used

An abnormal behavior recognition and early warning system based on the cloud-edge collaboration mechanism is adopted, and images and environmental data are obtained through edge IoT agents for intelligent analysis. Cloud IoT agents perform preprocessing and drive multimodal large model diagnosis, and edge IoT agents perform early warning operations and reinforcement learning.

Benefits of technology

It improves the accuracy of early warning operations and the intelligence of edge IoT intelligent bodies, and improves the intelligence and public safety of smart security scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337107B_ABST
    Figure CN120337107B_ABST
Patent Text Reader

Abstract

The present application relates to the fields of artificial intelligence technology and computer technology. The present application provides an abnormal behavior recognition and warning method and related system based on a cloud-edge collaborative mechanism. The first edge IoT agent obtains personnel image data and environmental data, and performs intelligent analysis on the personnel image data and environmental data to obtain a first analysis result. The cloud IoT agent pre-processes the personnel image data and environmental data to obtain first personnel image data and first environmental data. The multimodal large model is driven to diagnose the first personnel image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, and a first warning parameter is determined based on the first diagnosis result. The first edge IoT agent performs warning operations and reinforcement learning based on the first diagnosis result and the first warning parameter. The use of the embodiments of the present application can enhance the intelligence of smart security scenarios to ensure public safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of artificial intelligence technology and computer technology, and specifically to an abnormal behavior identification and early warning method and related system based on a cloud-edge collaboration mechanism. Background Art

[0002] Smart security scenarios have built a multi-dimensional active protection system by integrating technologies such as artificial intelligence, the Internet of Things, and big data. At present, the identification and early warning technologies of smart security scenarios are not intelligent, which reduces the security of smart security scenarios. Therefore, how to improve the intelligence of smart security scenarios to ensure public safety needs to be solved urgently. Summary of the Invention

[0003] The embodiments of the present application provide an abnormal behavior identification and early warning method and related system based on a cloud-edge collaboration mechanism, which can enhance the intelligence of smart security scenarios to ensure public safety.

[0004] In a first aspect, an embodiment of the present application provides an abnormal behavior recognition and early warning system based on a cloud-edge collaboration mechanism, the system comprising: multiple edge IoT agents and cloud IoT agents, wherein the cloud IoT agents allow driving a pre-configured multimodal large model in a cloud platform; wherein,

[0005] a first edge IoT agent, configured to obtain personnel image data and environmental data, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result; the first edge IoT agent is any one of the multiple edge IoT agents;

[0006] The cloud IoT agent is configured to preprocess the person image data and the environment data to obtain first person image data and first environment data; drive the multimodal large model to diagnose the first person image data, the first environment data, and the first analysis result to obtain a first diagnosis result; and determine a first warning parameter based on the first diagnosis result;

[0007] The first edge IoT agent is used to perform warning operations and reinforcement learning based on the first diagnostic result and the first warning parameter.

[0008] In a second aspect, an embodiment of the present application provides an abnormal behavior identification and early warning method based on a cloud-edge collaborative mechanism, which is applied to an abnormal behavior identification and early warning system based on a cloud-edge collaborative mechanism. The system includes: multiple edge IoT agents and a cloud IoT agent. The cloud IoT agent allows driving a pre-configured multimodal large model in a cloud platform; the method includes:

[0009] Acquire personnel image data and environmental data through a first edge IoT agent, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result; the first edge IoT agent is any one of the multiple edge IoT agents;

[0010] Preprocessing the person image data and the environmental data by the cloud IoT agent to obtain first person image data and first environmental data; driving the multimodal large model to diagnose the first person image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, and determining a first warning parameter based on the first diagnosis result;

[0011] The first edge IoT agent performs warning operations and reinforcement learning based on the first diagnostic result and the first warning parameters.

[0012] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the program comprises instructions for executing the steps in the second aspect of the embodiment of the present application.

[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the above-mentioned computer-readable storage medium stores a computer program for electronic data exchange, wherein the above-mentioned computer program enables a computer to execute some or all of the steps described in the second aspect of the embodiment of the present application.

[0014] In a fifth aspect, embodiments of the present application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to execute some or all of the steps described in the second aspect of the embodiments of the present application. The computer program product may be a software installation package.

[0015] The implementation of the embodiments of this application has the following beneficial effects:

[0016] It can be seen that the abnormal behavior recognition and early warning method and related system based on the cloud-edge collaborative mechanism described in the embodiment of the present application are applied to the abnormal behavior recognition and early warning system based on the cloud-edge collaborative mechanism, which includes: multiple edge IoT intelligent agents and cloud IoT intelligent agents. The cloud IoT intelligent agent allows the pre-configured multimodal large model in the cloud platform to be driven; wherein, the personnel image data and environmental data are obtained through the first edge IoT intelligent agent, and the personnel image data and environmental data are intelligently analyzed to obtain the first analysis result; the first edge IoT intelligent agent is any edge IoT intelligent agent among the multiple edge IoT intelligent agents, and the cloud IoT intelligent agent is connected to the edge IoT intelligent agent through the cloud IoT intelligent agent. The intelligent agent pre-processes the personnel image data and the environmental data to obtain the first personnel image data and the first environmental data; drives the multimodal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain the first diagnosis result, and determines the first warning parameter according to the first diagnosis result. The first edge IoT intelligent agent performs warning operations and reinforcement learning according to the first diagnosis result and the first warning parameter. On the one hand, the accuracy of the warning operation is improved. On the other hand, the intelligence of the edge IoT intelligent agent can be improved through reinforcement learning, and thus the intelligence of the smart security scene can be improved to ensure public safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 This is a schematic diagram of the architecture of an abnormal behavior identification and early warning system based on a cloud-edge collaboration mechanism provided in an embodiment of the present application;

[0019] Figure 2 This is a schematic diagram of a scene demonstration of a cloud IoT intelligent entity provided by an embodiment of the present application;

[0020] Figure 3 This is a schematic diagram of another scenario demonstration of a cloud IoT intelligent entity provided by an embodiment of the present application;

[0021] Figure 4 This is a schematic diagram of a scenario demonstration of an edge IoT intelligent agent provided by an embodiment of the present application;

[0022] Figure 5 This is a flow chart of an abnormal behavior identification and early warning method based on a cloud-edge collaboration mechanism provided in an embodiment of the present application;

[0023] Figure 6This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0024] Figure 7 This is a block diagram of the functional units of an abnormal behavior recognition and early warning device based on a cloud-edge collaboration mechanism provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may also include steps or elements not listed, or may include other steps or elements inherent to the process, method, product, or apparatus.

[0026] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0027] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0028] The edge IoT intelligent agent and cloud IoT intelligent agent involved in the embodiments of this application can be understood as devices that include intelligent agents. In specific implementations, intelligent agents can be specifically understood as: software programs, robots or other automated equipment, which have certain autonomy and intelligence.

[0029] In practice, intelligent agents can achieve specific goals by continuously learning and adapting through interaction with their environment. The core of intelligent agents lies in their autonomy, which allows them to adjust their behavior based on environmental changes, demonstrating a certain level of intelligence.

[0030] Among them, edge IoT agents and cloud IoT agents may include at least one of the following devices: smart cameras, smart phones, tablets, smart robots, smart home devices, vehicle-mounted devices, smart driving recorders, wearable devices, computing devices or other processing devices connected to wireless modems, as well as various forms of user equipment (UE), mobile stations (MS), terminal devices, etc., without limitation here.

[0031] Among them, edge IoT agents and cloud IoT agents can also be servers. For example, edge IoT agents can include edge servers or edge IoT devices. Cloud IoT agents can be deployed in the cloud and can be cloud servers.

[0032] See also Figure 1 , Figure 1 This is a schematic diagram of the architecture of a human abnormal behavior recognition and early warning system based on cloud-edge collaborative mechanism artificial intelligence technology provided by an embodiment of the present application. The system includes: multiple edge IoT agents and cloud IoT agents. The cloud IoT agents allow the pre-configured multimodal large model to be driven in the cloud platform; wherein,

[0033] a first edge IoT agent, configured to obtain personnel image data and environmental data, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result; the first edge IoT agent is any one of the multiple edge IoT agents;

[0034] The cloud IoT agent is configured to preprocess the person image data and the environment data to obtain first person image data and the first environment data; drive the multimodal large model to diagnose the first person image data, the first environment data, and the first analysis result to obtain a first diagnosis result; and determine a first warning parameter based on the first diagnosis result;

[0035] The first edge IoT agent is used to perform warning operations and reinforcement learning based on the first diagnostic result and the first warning parameter.

[0036] In the embodiment of the present application, the abnormal behavior recognition and warning system based on cloud-edge collaborative mechanism artificial intelligence technology can include multiple edge IoT agents and cloud IoT agents. The cloud IoT agent allows the pre-configured multimodal large model to be driven in the cloud platform. Figure 2 As shown in , the cloud IoT agent and the cloud platform can be two devices, both of which are set up in the cloud, and a multimodal large model is configured in the cloud platform. Figure 3As shown, the cloud IoT agent and the cloud platform can be the same device, that is, a multimodal large model can be configured in the cloud IoT agent.

[0037] The multimodal large model may be pre-set or set by system default. For example, the multimodal large model may include a deepseek large model, a ChatGpt large model, etc., without limitation. The multimodal large model may also include a neural network model, a deep learning model, etc., without limitation.

[0038] In specific implementations, a system for identifying and warning abnormal human behavior based on cloud-edge collaborative AI technology can include multiple edge IoT agents, a cloud IoT agent, and a multimodal large-scale model of smart security scenarios on the cloud platform. Multiple edge IoT agents perceive on-site personnel and environmental data, perform intelligent analysis, and obtain inference results. They can also generate warnings and take actions. Simultaneously, these data (on-site personnel and environmental data) and inference results can be transmitted to the cloud IoT agent. The cloud IoT agent pre-processes the data and drives the cloud platform's multimodal large-scale model to analyze and diagnose warning events, identify any anomalies, classify them, and determine the action strategy to be taken. After conversion and analysis, the cloud IoT agent transmits feedback information to the corresponding edge IoT agent, which then executes decisions for reinforcement learning.

[0039] Among them, personnel image data can be obtained by cameras, and environmental data can be obtained by environmental sensors.

[0040] The environmental sensor may include at least one of the following: a sound sensor, an odor sensor, a temperature sensor, a humidity sensor, a weather sensor, a light sensor, etc., without limitation herein. Accordingly, the environmental data may include at least one of the following: sound data, odor data, temperature data, humidity data, weather data, light brightness data, etc., without limitation herein.

[0041] Among them, such as Figure 4 As shown, the edge IoT intelligent body can communicate with multiple sensors, which may include at least one of the following: sound sensor (such as microphone array), odor sensor, temperature sensor, humidity sensor, meteorological sensor, light sensor, camera, electronic fence, infrared dual-detection sensor, access control device, etc., without limitation here.

[0042] In a specific implementation, taking a first edge IoT agent as an example, the first edge IoT agent is any edge IoT agent among multiple edge IoT agents. The first edge IoT agent can obtain personnel image data and environmental data, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result. The first analysis result may include at least one of the following: whether there is abnormal personnel behavior, the abnormal behavior category, action decision information, etc., which are not limited here.

[0043] In specific implementation, in the field of behavior recognition, action decision information can be understood as mainly referring to: key information used to judge and determine the specific behavior type in the process of identifying and analyzing human actions and behaviors. For example, in visual human action recognition, the positions of key points of the human body and the movement conditions of the actions obtained through the analysis of human posture and actions are all action decision information.

[0044] The first warning parameter may include at least one of the following: warning method, warning level, warning personnel, etc., which are not limited here. Different warning parameters can be pre-set for different abnormal behaviors of different personnel, that is, a mapping relationship between preset abnormal behaviors of personnel and warning parameters can be pre-stored.

[0045] Next, the cloud IoT agent can pre-process the personnel image data and environmental data to obtain the first personnel image data and the first environmental data, and then drive the multimodal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain the first diagnosis result, and determine the first warning parameter based on the first diagnosis result. The first edge IoT agent then performs warning operations and reinforcement learning based on the first diagnosis result and the first warning parameter. On the one hand, it improves the accuracy of the warning operation, and on the other hand, it can improve the intelligence of the edge IoT agent.

[0046] In the embodiment of the present application, each of the multiple edge IoT agents is an edge IoT agent with autonomous learning capabilities.

[0047] Edge IoT agents can be devices with hardware, software, and analytical capabilities. They access data from multiple sensors on-site, perform intelligent analysis of sensor data, and execute actions (warning operations) after warnings. Edge IoT agents can also be autonomously learning, using reinforcement learning to determine optimal action decisions in specific environments.

[0048] In specific implementation, the edge IoT intelligent agent can be responsible for collecting and perceiving various types of on-site personnel and environmental information data. Specifically, for example, the camera collects on-site video image data, the microphone array collects on-site sound signals, the electronic fence and infrared dual-detection sensors monitor personnel intrusion, and the access control equipment is responsible for recording personnel entry and exit events, and has a strategy for triggering analysis and reasoning.

[0049] Optionally, the multimodal large model includes a first-layer architecture and a second-layer architecture, wherein the first-layer architecture adopts an encoder-decoder architecture to implement recognition and understanding tasks of sound, image, and text, and outputs first text description information, wherein the first text description information includes text description information for at least one of the sound, image, and text;

[0050] The second-layer architecture includes an autoregressive Transormer architecture, which encodes and decodes the first text description information and outputs decoded second text description information. The second text description information is used to represent at least one of the following contents: whether there is abnormal behavior of the person, the category of abnormal behavior, and action decision information.

[0051] The abnormal behavior category can be pre-set or set by system default. The abnormal behavior category can be set based on experience, and different scenarios can correspond to different abnormal behavior categories. Abnormal behavior categories can include at least one of the following: fighting, quarreling, arguing, chatting, etc., which are not limited here.

[0052] In the embodiment of the present application, the multimodal large model has a two-layer network architecture. The first layer adopts an encoder-decoder transformer architecture to complete the recognition and understanding tasks of sound, image, and text, and outputs text description information of the sound and image. The second layer adopts an autoregressive transformer architecture to separately encode and decode the text description information output by the first layer architecture, and outputs the decoded text description information. This text description information may include whether there is abnormal behavior, the type of abnormal behavior, and the decision information about what action should be taken.

[0053] In the embodiment of the present application, the multimodal large model can be trained, deployed and run on the cloud platform side.

[0054] Optionally, when the environmental data includes sound data and sensor data, in preprocessing the personnel image data and the environmental data to obtain the first personnel image data and the first environmental data, the cloud IoT agent is specifically configured to:

[0055] Processing the sound data into a mel spectrogram;

[0056] Performing semantic processing on the sensor data to automatically generate descriptive semantics of the structured data to obtain a first descriptive semantics;

[0057] Determine the first environment data according to the mel spectrum graph and the first description semantics;

[0058] The personnel image data is standardized to obtain the first personnel image data.

[0059] In an embodiment of the present application, when the environmental data includes sound data and sensor data, the sound data can be processed into a mel spectrum graph, and the sensor data can be semantically processed to automatically generate descriptive semantics of structured data to obtain a first descriptive semantic. Then, the first environmental data can be determined based on the mel spectrum graph and the first descriptive semantic, and the personnel image data can be standardized to obtain the first personnel image data. Thus, the feature quality can be improved, which helps to improve the accuracy of subsequent diagnosis, and since the features are pre-processed, the diagnosis speed can be improved.

[0060] To give an example, the cloud IoT agent receives the sound, image signals, sensor structured data, and text data of edge inference results transmitted by the edge IoT agent. The cloud IoT agent preprocesses the information, processes the sound signal into a mel spectrum graph, standardizes the image data, analyzes the sensor structured data, and automatically generates descriptive semantics for the structured data. For example, the electronic fence alarm signal automatically generates a semantic description of "suspected intrusion."

[0061] Optionally, in driving the multimodal large model to diagnose the first person image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, the cloud IoT agent is specifically configured to:

[0062] Determine a first related prompt word dictionary set according to the first description semantics and the first analysis result;

[0063] The mel spectrum graph, the first person image data and the first related prompt word dictionary set are input into the multimodal large model for diagnosis to obtain the first diagnosis result.

[0064] In a specific implementation, a first set of related prompt word dictionaries can be determined based on the first description semantics and the first analysis result, and then the mel spectrum graph, the first person image data and the first set of related prompt word dictionaries can be input into the multimodal large model for diagnosis to obtain a first diagnosis result. To illustrate, specifically, the text data of the edge reasoning result can be combined with the generated semantic description. For example, if there is abnormal behavior of preset personnel, a set of related prompt word dictionaries can be generated. After the preprocessing is completed, the preprocessed sound, image data and related prompt word dictionary information are input into the multimodal large model for analysis, that is, auxiliary prompts can be made based on the analysis results of the edge IoT intelligent body, and on this basis, in-depth analysis can be performed again to deeply determine the authenticity of the abnormal behavior of personnel, which can deeply enhance the intelligence of the smart security scene to ensure public safety.

[0065] For example, after receiving the output analysis and decision information of the multimodal large model, the cloud IoT agent can analyze and compare the semantic similarity between the edge reasoning result text (first analysis result) and the multimodal large model analysis text (first diagnosis result), evaluate whether the edge IoT agent's reasoning is correct or wrong, transmit the reasoning evaluation results and decision information to the edge IoT agent through the network, wait for the edge IoT agent's action and data feedback, and give the edge IoT agent corresponding rewards based on the impact of this action after receiving the feedback.

[0066] Optionally, the cloud IoT agent is further specifically used for:

[0067] determining a first similarity between the first analysis result and the first diagnosis result;

[0068] Determining, based on the first similarity, a reasoning evaluation result of the first edge IoT agent regarding the first analysis result;

[0069] determining a first reward according to the reasoning evaluation result;

[0070] Push the first reward to the first edge IoT agent.

[0071] In a specific implementation, the first similarity between the first analysis result and the first diagnosis result can be determined. For example, the keyword of the first analysis result can be extracted to obtain the first keyword, the keyword of the first diagnosis result can be extracted to obtain the second keyword, and the similarity between the first keyword and the second keyword can be determined to obtain the first similarity. For another example, the first target text description information corresponding to the first analysis result can be determined, the second target text description information corresponding to the first diagnosis result can be determined, and the similarity between the first target text description information and the second target text description information can be determined to obtain the first similarity. For another example, the first vector corresponding to the first analysis result can be determined, the second vector corresponding to the first diagnosis result can be determined, and the similarity between the first vector and the second vector can be determined to obtain the first similarity.

[0072] Then, the reasoning evaluation result of the first edge IoT intelligent agent regarding the first analysis result can be determined based on the first similarity, and the mapping relationship between the preset similarity and the reasoning evaluation result can be pre-stored. Based on this method, the reasoning evaluation result corresponding to the first similarity can be determined. For example, a similarity threshold, that is, a preset value, can be preset. When the first similarity is greater than or equal to the preset value, the reasoning evaluation result can be determined to be that the first analysis result is correct. Conversely, when the first similarity is less than the preset value, the reasoning evaluation result can be determined to be that the first analysis result is wrong.

[0073] Next, a pre-set mapping relationship between reasoning evaluation results and rewards can be pre-stored. Based on this mapping relationship, the first reward corresponding to the corresponding reasoning evaluation result can be determined and then pushed to the first edge IoT agent. The reward provides immediate feedback on the first edge IoT agent's behavior, used to evaluate the quality of a specific action in a certain state, thereby influencing its future decision-making. Through continuous trial and error and adjustment, the first edge IoT agent learns to select behavioral strategies that will obtain high rewards in different states, thereby improving the intelligence of smart security scenarios to ensure public safety.

[0074] Optionally, the cloud IoT agent is further specifically used for:

[0075] When the reasoning evaluation result includes that the first analysis result is wrong, performing the step of determining the first warning parameter according to the first diagnosis result;

[0076] The first edge IoT intelligent agent is further specifically configured to:

[0077] When the reasoning evaluation result includes that the first analysis result is correct, a second warning parameter corresponding to the first analysis result is determined, and a warning operation is performed according to the second warning parameter.

[0078] In a specific implementation, when the reasoning evaluation result includes an error in the first analysis result, the step of determining the first warning parameter based on the first diagnostic result can be executed, that is, if the first edge IoT agent has an inference error, the corresponding warning operation is performed based on the first diagnostic result in the cloud to ensure the accuracy of the warning.

[0079] The second warning parameter may include at least one of the following: warning method, warning level, warning personnel, etc., which are not limited here. Different warning parameters can be pre-set for different abnormal behaviors of different personnel, that is, a mapping relationship between preset abnormal behaviors of personnel and warning parameters can be pre-stored.

[0080] Correspondingly, when the reasoning evaluation results include that the first analysis result is correct, it means that the reasoning of the first edge IoT intelligent body is correct, and the corresponding early warning operation can be directly performed based on the first analysis result to improve the early warning speed. Different analysis results can correspond to different early warning parameters, and then the second early warning parameters corresponding to the first analysis result are determined, and the early warning operation is performed according to the second early warning parameters. In this way, the early warning speed can be guaranteed, and then the intelligence of the smart security scene can be improved to ensure public safety.

[0081] Optionally, the environmental data includes sound data; in performing intelligent analysis on the personnel image data and the environmental data to obtain the first analysis result, the first edge IoT agent is specifically configured to:

[0082] Performing feature extraction on the personnel image data to obtain a first feature set;

[0083] Inputting the first feature set into a small behavior recognition model to obtain a first recognition result;

[0084] Performing feature extraction on the sound data to obtain a second feature set;

[0085] Inputting the second feature set into the voiceprint recognition small model to obtain a second recognition result;

[0086] The first analysis result is determined according to the first recognition result and the second recognition result.

[0087] Among them, the behavior recognition small model can be pre-set or system default, and the voiceprint recognition small model can be pre-set or system default.

[0088] In an embodiment of the present application, feature extraction can be performed on personnel image data to obtain a first feature set, which may include at least one of the following: feature points, feature vectors, feature values, feature patterns, etc., which are not limited here.

[0089] In a specific implementation, the first feature set can be input into a small behavior recognition model to obtain a first recognition result, and then a corresponding image recognition result can be obtained.

[0090] Accordingly, feature extraction can be performed on the sound data to obtain a second feature set. This second feature set may include at least one of the following: feature points, feature vectors, feature values, etc., which are not limited here. The second feature set is then input into the voiceprint recognition model to obtain a second recognition result, and then a corresponding voice recognition result can be obtained. The first analysis result is then determined based on the first and second recognition results. That is, the final first analysis result can be determined by combining the image recognition result and the voice recognition result. This combines the image with the environment, which can ensure the accuracy of identifying abnormal human behavior to a certain extent, and thus enhance the intelligence of smart security scenarios to ensure public safety.

[0091] To illustrate, in a specific implementation, the edge IoT agent can quickly identify and infer suspicious abnormal behaviors of on-site personnel based on the configured multiple edge-side small models. For example, if there is an illegal intrusion at the scene, the camera or access control is damaged, and the abnormal environmental sound is collected by the microphone array, the edge IoT agent will call the voiceprint recognition small model to determine the intrusion event, drive the on-site alarm and protection device, and upload the data and inference results to the cloud IoT agent in real time.

[0092] In specific implementations, the edge IoT agent can receive inference results from the cloud IoT agent after the fact, use them to update the managed edge model sample library, and trigger the edge model to train and update. It then receives action decision information from the cloud IoT agent after the fact, executes the recommended action, observes the results, collects environmental data again, and uploads it to the cloud IoT agent in real time. The cloud IoT agent analyzes the impact of this action and rewards the edge IoT agent accordingly.

[0093] In a specific implementation, feature extraction is performed on person image data to obtain a first feature set, which can then determine the person's motion speed. For example, a mapping relationship between preset motion speed and feature extraction algorithms can be pre-stored. Furthermore, a first feature extraction algorithm corresponding to the person's working speed can be determined based on this mapping relationship. A first screen-occupancy ratio of the person can also be determined. A mapping relationship between a preset screen-occupancy ratio and control parameters of the first feature extraction algorithm can also be pre-stored. Based on this mapping relationship, a first control parameter corresponding to the first screen-occupancy ratio can be determined. In this way, on the one hand, the corresponding feature extraction algorithm can be adapted based on motion speed, facilitating accurate capture of key movements. On the other hand, adapting the control parameters of the first feature extraction algorithm based on the screen-occupancy ratio can capture subtle movements, thereby improving the accuracy of subsequent abnormal behavior identification. The first screen-occupancy ratio can be understood as the average screen-occupancy ratio or the screen-occupancy ratio of a specified person. The screen-occupancy ratio can be understood as the ratio between the person area and the entire image area. The specified person can be pre-set or set by the system default, and can be one or more persons. The control parameters of the first feature extraction algorithm are used to control the algorithm performance of the first feature extraction algorithm. This algorithm performance can include at least one of the following: algorithm speed, accuracy, location, feature quality, etc., which are not limited here.

[0094] Optionally, the first analysis result is determined according to the first recognition result and the second recognition result, and the first edge IoT agent is specifically used to:

[0095] determining a second similarity between the first recognition result and the second recognition result;

[0096] When the second similarity is greater than a preset similarity, determining a first intersection result of the first recognition result and the second recognition result, and determining the first analysis result according to the first intersection result;

[0097] When the second similarity is less than or equal to the preset similarity, concatenating the first feature set and the second feature set to obtain a third feature set;

[0098] The third feature set is input into the local model of the multimodal large model to obtain the first analysis result.

[0099] The preset similarity can be preset or set by system default.

[0100] In a specific implementation, the second similarity between the first recognition result and the second recognition result can be determined. For example, the keyword of the first recognition result can be extracted to obtain a keyword, and the keyword of the second analysis result can be extracted to obtain another keyword. The similarity between the two keywords can be determined to obtain the second similarity.

[0101] Next, when the second similarity is greater than the preset similarity, the first intersection result of the first recognition result and the second recognition result can be determined, and the first analysis result can be determined based on the first intersection result, that is, the image and the environment are combined. If the recognition results of the two are the same, it means that both dimensions have identified abnormal behavior of people and have a certain degree of accuracy. The results of the two can be combined to ensure the accuracy of abnormal behavior of people to a certain extent, and thus, the intelligence of the smart security scene can be improved to ensure public safety.

[0102] Correspondingly, when the second similarity is less than or equal to the preset similarity, the image and the environment are combined, and the recognition results of the two are quite different. The first feature set and the second feature set can be spliced ​​to obtain a third feature set, and then the third feature set is input into the local model of the multimodal large model to obtain the first analysis result, that is, the image and the environment are combined. To a certain extent, the accuracy of identifying abnormal behaviors of people can be guaranteed, and then the intelligence of smart security scenes can be improved to ensure public safety.

[0103] The local model of the multimodal large model may be updated at preset time intervals, and the preset time intervals may be pre-set or set by system default.

[0104] Optionally, in inputting the first feature set into the behavior recognition small model to obtain a first recognition result, the first edge IoT agent is specifically configured to:

[0105] Determining a first environment complexity corresponding to the sound data;

[0106] determining a first model parameter corresponding to the first environmental complexity;

[0107] The first recognition result is determined according to the first feature set, the first model parameters and the behavior recognition small model.

[0108] Among them, the sound data can be analyzed to obtain multiple sound analysis results. For example, the sound analysis results may include: environmental noise, number of sound sources, voice quality, etc. The mapping relationship between the preset sound analysis results of each dimension and the environmental complexity can be pre-stored. Based on the mapping relationship, the environmental complexity corresponding to each sound analysis result in the multiple sound analysis results can be determined to obtain multiple environmental complexities, and multiple weights corresponding to the multiple sound analysis results are obtained. The sum of the multiple weights is 1. Then, a weighted operation is performed on the multiple environmental complexities and the multiple weights to obtain a first environmental complexity. The multiple weights can be pre-set or system default. For example, the multiple weights can be related to time or weather, which is not limited here.

[0109] In a specific implementation, a mapping relationship between a preset environmental complexity and a model parameter of a behavior recognition small model can be pre-stored. Then, based on the mapping relationship, the first model parameter corresponding to the first environmental complexity can be determined, and then the behavior recognition small model can be configured according to the first model parameter. Then, the first feature set is input into the behavior recognition small model configured with the first model parameter to obtain a first recognition result. In this way, the environmental factors can be used to optimize the model parameters of the behavior recognition small model in turn, so that the model capability depth of the behavior recognition small model is adapted to the actual environment, which helps to improve the accuracy of abnormal behavior recognition of the behavior recognition small model. For example, if the abnormal behavior of a person is a fight, and the sound detects a corresponding quarrel, the weight of the abnormal behavior "fighting" of the person can be adjusted, thereby helping to further ensure the accuracy of abnormal behavior recognition of the person, that is, the image and the environment are deeply combined, which can ensure the accuracy of abnormal behavior recognition of the person to a certain extent, and then, it can improve the intelligence of the smart security scene to ensure public safety.

[0110] Optionally, the first recognition result includes a first abnormal behavior type and a first abnormal probability value; and in inputting the second feature set into the voiceprint recognition model to obtain the second recognition result, the first edge IoT agent is specifically configured to:

[0111] determining a second model parameter corresponding to the first abnormal behavior type;

[0112] determining a first adjustment parameter corresponding to the first abnormality probability value;

[0113] Adjusting the second model parameters according to the first adjustment parameters to obtain target second model parameters;

[0114] The second recognition result is determined according to the second feature set, the target second model parameters and the voiceprint recognition small model.

[0115] In a specific implementation, the first recognition result includes a first abnormal behavior type and a first abnormal probability value. The first abnormal behavior type may include at least one abnormal behavior of a person. The first abnormal probability value may include at least one probability value. Each abnormal behavior of a person corresponds to a probability value.

[0116] Next, the mapping relationship between the preset abnormal behavior type and the model parameters of the voiceprint recognition small model can be pre-stored, and then the second model parameter corresponding to the first abnormal behavior type can be determined based on the mapping relationship. The mapping relationship between the preset probability value and the adjustment parameter can also be pre-stored. Based on the mapping relationship, the first adjustment parameter corresponding to the first abnormal probability value can be determined, and then part or all of the model parameters of the second model parameter are adjusted according to the first adjustment parameter to obtain the target second model parameter. For example, the target second model parameter = (1 + first adjustment parameter) × second model parameter, and then the voiceprint recognition small model is configured according to the target second model parameter, and the second feature set is input into the voiceprint recognition small model configured with the target second model parameter to obtain the second recognition result.

[0117] In this way, the recognition results of the image dimension can be used to optimize the model parameters of the voiceprint recognition model in turn, so that the model capabilities of the voiceprint recognition model are consistent with the actual environment depth, which helps to improve the accuracy of the abnormal behavior recognition of the voiceprint recognition model. That is, the depth of the image and the environment can be combined to a certain extent, which can ensure the accuracy of abnormal behavior recognition of personnel, and then, can improve the intelligence of the smart security scene to ensure public safety.

[0118] It can be seen that the abnormal behavior identification and early warning system for personnel based on cloud-edge collaborative mechanism artificial intelligence technology described in the embodiment of the present application includes: multiple edge IoT agents and cloud IoT agents. The cloud IoT agent allows the cloud platform to be driven to pre-configure a multimodal large model; wherein, personnel image data and environmental data are obtained through the first edge IoT agent, and the personnel image data and environmental data are intelligently analyzed to obtain a first analysis result; the first edge IoT agent is any edge IoT agent among the multiple edge IoT agents, and the cloud IoT agent is used to pre-process the personnel image data and environmental data to obtain the first personnel image data and the first environmental data; the multimodal large model is driven to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain a first diagnosis result, and the first warning parameter is determined according to the first diagnosis result, and the first edge IoT agent performs early warning operations and reinforcement learning according to the first diagnosis result and the first warning parameters. On the one hand, the accuracy of the early warning operation is improved, and on the other hand, the intelligence of the edge IoT agent can be improved, and then the intelligence of the smart security scene can be improved to ensure public safety.

[0119] See also Figure 5 , Figure 5 This is a flow chart of a method for identifying and warning abnormal human behavior based on cloud-edge collaborative artificial intelligence technology provided by an embodiment of the present application, which is applied to Figure 1The system for identifying and warning abnormal human behavior based on cloud-edge collaborative mechanism artificial intelligence technology shown in the figure includes: multiple edge IoT agents and cloud IoT agents, wherein the cloud IoT agents allow driving a pre-configured multimodal large model in the cloud platform; the method for identifying and warning abnormal human behavior based on cloud-edge collaborative mechanism artificial intelligence technology includes:

[0120] 501. Obtaining personnel image data and environmental data through a first edge IoT agent, and performing intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result; the first edge IoT agent is any one of the multiple edge IoT agents;

[0121] 502. Preprocess the person image data and the environment data by the cloud IoT agent to obtain first person image data and first environment data; drive the multimodal large model to diagnose the first person image data, the first environment data, and the first analysis result to obtain a first diagnosis result; and determine a first warning parameter based on the first diagnosis result.

[0122] 503. Perform warning operations and reinforcement learning according to the first diagnostic result and the first warning parameter through the first edge IoT agent.

[0123] Optionally, the multimodal large model includes a first-layer architecture and a second-layer architecture, wherein the first-layer architecture adopts an encoder-decoder architecture to implement recognition and understanding tasks of sound, image, and text, and outputs first text description information, wherein the first text description information includes text description information for at least one of the sound, image, and text;

[0124] The second-layer architecture includes an autoregressive Transormer architecture, which encodes and decodes the first text description information and outputs decoded second text description information. The second text description information is used to represent at least one of the following contents: whether there is abnormal behavior of the person, the category of abnormal behavior, and action decision information.

[0125] Optionally, when the environmental data includes sound data and sensor data, the above step 502 of preprocessing the person image data and the environmental data to obtain the first person image data and the first environmental data may be implemented as follows:

[0126] Processing the sound data into a mel spectrogram;

[0127] Performing semantic processing on the sensor data to automatically generate descriptive semantics of the structured data to obtain a first descriptive semantics;

[0128] Determine the first environment data according to the mel spectrum graph and the first description semantics;

[0129] The personnel image data is standardized to obtain the first personnel image data.

[0130] Optionally, the above step 502, driving the multimodal large model to diagnose the first person image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, can be implemented as follows:

[0131] Determine a first related prompt word dictionary set according to the first description semantics and the first analysis result;

[0132] The mel spectrum graph, the first person image data and the first related prompt word dictionary set are input into the multimodal large model for diagnosis to obtain the first diagnosis result.

[0133] Optionally, the following steps may also be included:

[0134] determining a first similarity between the first analysis result and the first diagnosis result;

[0135] Determining, based on the first similarity, a reasoning evaluation result of the first edge IoT agent regarding the first analysis result;

[0136] determining a first reward according to the reasoning evaluation result;

[0137] Push the first reward to the first edge IoT agent.

[0138] Optionally, the following steps may also be included:

[0139] When the reasoning evaluation result includes that the first analysis result is wrong, performing the step of determining the first warning parameter according to the first diagnosis result;

[0140] When the reasoning evaluation result includes that the first analysis result is correct, the first edge IoT agent determines a second warning parameter corresponding to the first analysis result, and performs a warning operation according to the second warning parameter.

[0141] Optionally, the environmental data includes sound data; the above step 501 of performing intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result may be implemented as follows:

[0142] Performing feature extraction on the personnel image data to obtain a first feature set;

[0143] Inputting the first feature set into a small behavior recognition model to obtain a first recognition result;

[0144] Performing feature extraction on the sound data to obtain a second feature set;

[0145] Inputting the second feature set into the voiceprint recognition small model to obtain a second recognition result;

[0146] The first analysis result is determined according to the first recognition result and the second recognition result.

[0147] Optionally, the above step of determining the first analysis result based on the first recognition result and the second recognition result may be implemented as follows:

[0148] determining a second similarity between the first recognition result and the second recognition result;

[0149] When the second similarity is greater than a preset similarity, determining a first intersection result of the first recognition result and the second recognition result, and determining the first analysis result according to the first intersection result;

[0150] When the second similarity is less than or equal to the preset similarity, concatenating the first feature set and the second feature set to obtain a third feature set;

[0151] The third feature set is input into the local model of the multimodal large model to obtain the first analysis result.

[0152] Optionally, the above step of inputting the first feature set into the behavior recognition model to obtain the first recognition result can be implemented as follows:

[0153] Determining a first environment complexity corresponding to the sound data;

[0154] determining a first model parameter corresponding to the first environmental complexity;

[0155] The first recognition result is determined according to the first feature set, the first model parameters and the behavior recognition small model.

[0156] Optionally, the first recognition result includes a first abnormal behavior type and a first abnormal probability value; the above step of inputting the second feature set into the voiceprint recognition model to obtain the second recognition result can be implemented as follows:

[0157] determining a second model parameter corresponding to the first abnormal behavior type;

[0158] determining a first adjustment parameter corresponding to the first abnormality probability value;

[0159] Adjusting the second model parameters according to the first adjustment parameters to obtain target second model parameters;

[0160] The second recognition result is determined according to the second feature set, the target second model parameters and the voiceprint recognition small model.

[0161] The specific description of the above steps can refer to the above Figure 1 The corresponding parts in the described embodiments will not be repeated here.

[0162] It can be seen that the method for identifying and warning abnormal behavior of personnel based on cloud-edge collaborative mechanism artificial intelligence technology described in the embodiment of the present application is applied to the system for identifying and warning abnormal behavior of personnel based on cloud-edge collaborative mechanism artificial intelligence technology, which includes: multiple edge IoT intelligent agents and cloud IoT intelligent agents. The cloud IoT intelligent agent allows the pre-configured multimodal large model in the cloud platform to be driven; wherein, the first edge IoT intelligent agent obtains personnel image data and environmental data, and performs intelligent analysis on the personnel image data and environmental data to obtain a first analysis result; the first edge IoT intelligent agent is any edge IoT agent among the multiple edge IoT intelligent agents. The intelligent agent pre-processes the personnel image data and the environmental data through the cloud IoT intelligent agent to obtain the first personnel image data and the first environmental data; drives the multimodal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain the first diagnosis result, and determines the first warning parameter according to the first diagnosis result, and performs warning operations and reinforcement learning according to the first diagnosis result and the first warning parameter through the first edge IoT intelligent agent. On the one hand, the accuracy of the warning operation is improved, and on the other hand, the intelligence of the edge IoT intelligent agent can be improved, and then the intelligence of the smart security scene can be improved to ensure public safety.

[0163] In accordance with the above embodiment, please refer to Figure 6 , Figure 6 This is a structural diagram of an electronic device provided in an embodiment of the present application. As shown in the figure, the electronic device includes a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor. In an embodiment of the present application, a system for identifying and warning abnormal human behavior based on cloud-edge collaborative mechanism artificial intelligence technology is applied. The system includes: multiple edge IoT agents and cloud IoT agents. The cloud IoT agents allow driving a pre-configured multimodal large model in a cloud platform; the program includes instructions for executing the following steps:

[0164] Acquire personnel image data and environmental data through a first edge IoT agent, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result; the first edge IoT agent is any one of the multiple edge IoT agents;

[0165] Preprocessing the person image data and the environmental data by the cloud IoT agent to obtain first person image data and first environmental data; driving the multimodal large model to diagnose the first person image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, and determining a first warning parameter based on the first diagnosis result;

[0166] The first edge IoT agent performs warning operations and reinforcement learning based on the first diagnostic result and the first warning parameters.

[0167] Optionally, the multimodal large model includes a first-layer architecture and a second-layer architecture, wherein the first-layer architecture adopts an encoder-decoder architecture to implement recognition and understanding tasks of sound, image, and text, and outputs first text description information, wherein the first text description information includes text description information for at least one of the sound, image, and text;

[0168] The second-layer architecture includes an autoregressive Transormer architecture, which encodes and decodes the first text description information and outputs decoded second text description information. The second text description information is used to represent at least one of the following contents: whether there is abnormal behavior of the person, the category of abnormal behavior, and action decision information.

[0169] Optionally, when the environmental data includes sound data and sensor data, in the aspect of preprocessing the person image data and the environmental data to obtain the first person image data and the first environmental data, the program includes instructions for performing the following steps:

[0170] Processing the sound data into a mel spectrogram;

[0171] Performing semantic processing on the sensor data to automatically generate descriptive semantics of the structured data to obtain a first descriptive semantics;

[0172] Determine the first environment data according to the mel spectrum graph and the first description semantics;

[0173] The personnel image data is standardized to obtain the first personnel image data.

[0174] Optionally, in terms of driving the multimodal large model to diagnose the first person image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, the program includes instructions for executing the following steps:

[0175] Determine a first related prompt word dictionary set according to the first description semantics and the first analysis result;

[0176] The mel spectrum graph, the first person image data and the first related prompt word dictionary set are input into the multimodal large model for diagnosis to obtain the first diagnosis result.

[0177] Optionally, the program further includes instructions for executing the following steps:

[0178] determining a first similarity between the first analysis result and the first diagnosis result;

[0179] Determining, based on the first similarity, a reasoning evaluation result of the first edge IoT agent regarding the first analysis result;

[0180] determining a first reward according to the reasoning evaluation result;

[0181] Push the first reward to the first edge IoT agent.

[0182] Optionally, the program further includes instructions for executing the following steps:

[0183] When the reasoning evaluation result includes that the first analysis result is wrong, performing the step of determining the first warning parameter according to the first diagnosis result;

[0184] When the reasoning evaluation result includes that the first analysis result is correct, the first edge IoT agent determines a second warning parameter corresponding to the first analysis result, and performs a warning operation according to the second warning parameter.

[0185] Optionally, the environmental data includes sound data; and in terms of performing intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result, the program includes instructions for executing the following steps:

[0186] Performing feature extraction on the personnel image data to obtain a first feature set;

[0187] Inputting the first feature set into a small behavior recognition model to obtain a first recognition result;

[0188] Performing feature extraction on the sound data to obtain a second feature set;

[0189] Inputting the second feature set into the voiceprint recognition small model to obtain a second recognition result;

[0190] The first analysis result is determined according to the first recognition result and the second recognition result.

[0191] Optionally, in determining the first analysis result based on the first recognition result and the second recognition result, the program includes instructions for performing the following steps:

[0192] determining a second similarity between the first recognition result and the second recognition result;

[0193] When the second similarity is greater than a preset similarity, determining a first intersection result of the first recognition result and the second recognition result, and determining the first analysis result according to the first intersection result;

[0194] When the second similarity is less than or equal to the preset similarity, concatenating the first feature set and the second feature set to obtain a third feature set;

[0195] The third feature set is input into the local model of the multimodal large model to obtain the first analysis result.

[0196] Optionally, in the aspect of inputting the first feature set into the behavior recognition small model to obtain a first recognition result, the program includes instructions for executing the following steps:

[0197] Determining a first environment complexity corresponding to the sound data;

[0198] determining a first model parameter corresponding to the first environmental complexity;

[0199] The first recognition result is determined according to the first feature set, the first model parameters and the behavior recognition small model.

[0200] Optionally, the first recognition result includes a first abnormal behavior type and a first abnormal probability value; and in inputting the second feature set into the voiceprint recognition small model to obtain the second recognition result, the program includes instructions for executing the following steps:

[0201] determining a second model parameter corresponding to the first abnormal behavior type;

[0202] determining a first adjustment parameter corresponding to the first abnormality probability value;

[0203] Adjusting the second model parameters according to the first adjustment parameters to obtain target second model parameters;

[0204] The second recognition result is determined according to the second feature set, the target second model parameters and the voiceprint recognition small model.

[0205] It can be seen that the electronic device described in the embodiment of the present application is applied to a personnel abnormal behavior recognition and warning system based on cloud-edge collaborative mechanism artificial intelligence technology, which includes: multiple edge IoT agents and cloud IoT agents. The cloud IoT agent allows the cloud platform to be driven to pre-configure a multimodal large model; wherein, personnel image data and environmental data are obtained through the first edge IoT agent, and the personnel image data and environmental data are intelligently analyzed to obtain a first analysis result; the first edge IoT agent is any edge IoT agent among the multiple edge IoT agents, and the cloud IoT agent is used to pre-process the personnel image data and environmental data to obtain first personnel image data and first environmental data; the multimodal large model is driven to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain a first diagnosis result, and the first warning parameter is determined according to the first diagnosis result, and the first edge IoT agent performs warning operations and reinforcement learning according to the first diagnosis result and the first warning parameters. On the one hand, the accuracy of the warning operation is improved, and on the other hand, the intelligence of the edge IoT agent can be improved, and thus the intelligence of the smart security scene can be improved to ensure public safety.

[0206] The electronic device may include a controller of an edge IoT agent, or an edge IoT agent, or a controller of a cloud IoT agent, or a cloud IoT agent.

[0207] Figure 7 This is a functional unit composition block diagram of a device 700 for identifying and warning abnormal human behavior based on cloud-edge collaborative mechanism artificial intelligence technology involved in an embodiment of the present application. The device 700 for identifying and warning abnormal human behavior based on cloud-edge collaborative mechanism artificial intelligence technology is applied to a system for identifying and warning abnormal human behavior based on cloud-edge collaborative mechanism artificial intelligence technology, which includes: multiple edge IoT intelligent agents and cloud IoT intelligent agents, wherein the cloud IoT intelligent agents allow driving a pre-configured multimodal large model in a cloud platform; the device 700 for identifying and warning abnormal human behavior based on cloud-edge collaborative mechanism artificial intelligence technology includes: an analysis unit 701, a diagnosis unit 702 and a feedback unit 703, wherein,

[0208] The analysis unit 701 is configured to obtain personnel image data and environmental data through a first edge IoT agent, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result; the first edge IoT agent is any edge IoT agent among the multiple edge IoT agents;

[0209] The diagnosis unit 702 is configured to pre-process the person image data and the environment data through the cloud IoT agent to obtain first person image data and first environment data; drive the multimodal large model to diagnose the first person image data, the first environment data, and the first analysis result to obtain a first diagnosis result; and determine a first warning parameter based on the first diagnosis result;

[0210] The feedback unit 703 is used to perform warning operations and reinforcement learning based on the first diagnosis result and the first warning parameter through the first edge IoT agent.

[0211] Optionally, the multimodal large model includes a first-layer architecture and a second-layer architecture, wherein the first-layer architecture adopts an encoder-decoder architecture to implement recognition and understanding tasks of sound, image, and text, and outputs first text description information, wherein the first text description information includes text description information for at least one of the sound, image, and text;

[0212] The second-layer architecture includes an autoregressive Transormer architecture, which encodes and decodes the first text description information and outputs decoded second text description information. The second text description information is used to represent at least one of the following contents: whether there is abnormal behavior of the person, the category of abnormal behavior, and action decision information.

[0213] Optionally, when the environmental data includes sound data and sensor data, in preprocessing the person image data and the environmental data to obtain the first person image data and the first environmental data, the diagnosis unit 702 is specifically configured to:

[0214] Processing the sound data into a mel spectrogram;

[0215] Performing semantic processing on the sensor data to automatically generate descriptive semantics of the structured data to obtain a first descriptive semantics;

[0216] Determine the first environment data according to the mel spectrum graph and the first description semantics;

[0217] The personnel image data is standardized to obtain the first personnel image data.

[0218] Optionally, in driving the multimodal large model to diagnose the first person image data, the first environmental data, and the first analysis result to obtain a first diagnosis result, the diagnosis unit 702 is specifically configured to:

[0219] Determine a first related prompt word dictionary set according to the first description semantics and the first analysis result;

[0220] The mel spectrum graph, the first person image data and the first related prompt word dictionary set are input into the multimodal large model for diagnosis to obtain the first diagnosis result.

[0221] Optionally, the abnormal behavior identification and warning system 700 based on cloud-edge collaborative mechanism artificial intelligence technology is further specifically used for:

[0222] determining a first similarity between the first analysis result and the first diagnosis result;

[0223] Determining, based on the first similarity, a reasoning evaluation result of the first edge IoT agent regarding the first analysis result;

[0224] determining a first reward according to the reasoning evaluation result;

[0225] Push the first reward to the first edge IoT agent.

[0226] Optionally, the abnormal behavior identification and warning system 700 based on cloud-edge collaborative mechanism artificial intelligence technology is further specifically used for:

[0227] When the reasoning evaluation result includes that the first analysis result is wrong, performing the step of determining the first warning parameter according to the first diagnosis result;

[0228] When the reasoning evaluation result includes that the first analysis result is correct, the first edge IoT agent determines a second warning parameter corresponding to the first analysis result, and performs a warning operation according to the second warning parameter.

[0229] Optionally, the environmental data includes sound data; in performing intelligent analysis on the personnel image data and the environmental data to obtain the first analysis result, the analyzing unit 701 is specifically configured to:

[0230] Performing feature extraction on the personnel image data to obtain a first feature set;

[0231] Inputting the first feature set into a small behavior recognition model to obtain a first recognition result;

[0232] Performing feature extraction on the sound data to obtain a second feature set;

[0233] Inputting the second feature set into the voiceprint recognition small model to obtain a second recognition result;

[0234] The first analysis result is determined according to the first recognition result and the second recognition result.

[0235] Optionally, in determining the first analysis result according to the first recognition result and the second recognition result, the analyzing unit 701 is specifically configured to:

[0236] determining a second similarity between the first recognition result and the second recognition result;

[0237] When the second similarity is greater than a preset similarity, determining a first intersection result of the first recognition result and the second recognition result, and determining the first analysis result according to the first intersection result;

[0238] When the second similarity is less than or equal to the preset similarity, concatenating the first feature set and the second feature set to obtain a third feature set;

[0239] The third feature set is input into the local model of the multimodal large model to obtain the first analysis result.

[0240] Optionally, in inputting the first feature set into the behavior recognition small model to obtain a first recognition result, the analyzing unit 701 is specifically configured to:

[0241] Determining a first environment complexity corresponding to the sound data;

[0242] determining a first model parameter corresponding to the first environmental complexity;

[0243] The first recognition result is determined according to the first feature set, the first model parameters and the behavior recognition small model.

[0244] Optionally, the first recognition result includes a first abnormal behavior type and a first abnormal probability value; in inputting the second feature set into the voiceprint recognition small model to obtain the second recognition result, the analyzing unit 701 is specifically configured to:

[0245] determining a second model parameter corresponding to the first abnormal behavior type;

[0246] determining a first adjustment parameter corresponding to the first abnormality probability value;

[0247] Adjusting the second model parameters according to the first adjustment parameters to obtain target second model parameters;

[0248] The second recognition result is determined according to the second feature set, the target second model parameters and the voiceprint recognition small model.

[0249] It can be seen that the abnormal behavior identification and early warning device for personnel based on cloud-edge collaborative mechanism artificial intelligence technology described in the embodiment of the present application is applied to the abnormal behavior identification and early warning system for personnel based on cloud-edge collaborative mechanism artificial intelligence technology, which includes: multiple edge IoT intelligent agents and cloud IoT intelligent agents. The cloud IoT intelligent agent allows the pre-configured multimodal large model in the cloud platform to be driven; wherein, the personnel image data and environmental data are obtained through the first edge IoT intelligent agent, and the personnel image data and environmental data are intelligently analyzed to obtain a first analysis result; the first edge IoT intelligent agent is any edge IoT agent among the multiple edge IoT intelligent agents. The intelligent agent pre-processes the personnel image data and the environmental data through the cloud IoT intelligent agent to obtain the first personnel image data and the first environmental data; drives the multimodal large model to diagnose the first personnel image data, the first environmental data and the first analysis result to obtain the first diagnosis result, and determines the first warning parameter according to the first diagnosis result, and performs warning operations and reinforcement learning according to the first diagnosis result and the first warning parameter through the first edge IoT intelligent agent. On the one hand, the accuracy of the warning operation is improved, and on the other hand, the intelligence of the edge IoT intelligent agent can be improved, and then the intelligence of the smart security scene can be improved to ensure public safety.

[0250] It can be understood that the functions of the abnormal behavior identification and warning device for personnel based on cloud-edge collaborative mechanism artificial intelligence technology in this embodiment can be specifically implemented according to the method in the above method embodiment. Its specific implementation process can refer to the relevant description of the above method embodiment, and will not be repeated here.

[0251] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute part or all of the steps of any method described in the above method embodiments, and the above computer includes an electronic device.

[0252] The present application also provides a computer program product comprising a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may comprise an electronic device.

[0253] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0254] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0255] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0256] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0257] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0258] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the above-mentioned methods in each embodiment of the present application. The aforementioned memory includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program code.

[0259] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing related hardware. The program can be stored in a computer-readable memory, which may include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0260] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. At the same time, for those skilled in the art, according to the idea of ​​the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. An abnormal behavior identification and early warning system based on cloud-edge collaboration mechanism, characterized by: The system includes: multiple edge IoT agents and cloud IoT agents, wherein the cloud IoT agents are capable of driving a pre-configured multimodal large model in a cloud platform; each of the multiple edge IoT agents is an edge IoT agent with autonomous learning capabilities; wherein, a first edge IoT agent, configured to obtain personnel image data and environmental data, and perform intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result, wherein the environmental data includes sound data and sensor data; the first edge IoT agent is any one of the multiple edge IoT agents; The cloud IoT agent is configured to pre-process the person image data and the environment data to obtain first person image data and first environment data, specifically by: processing the sound data into a mel spectrum graph, performing semantic processing on the sensor data to automatically generate descriptive semantics of structured data to obtain first descriptive semantics, determining the first environment data based on the mel spectrum graph and the first descriptive semantics, and performing standardization processing on the person image data to obtain the first person image data; driving the multimodal large model to diagnose the first person image data, the first environment data, and the first analysis result to obtain a first diagnosis result, specifically by: determining a first relevant prompt word dictionary set based on the first descriptive semantics and the first analysis result, inputting the mel spectrum graph, the first person image data, and the first relevant prompt word dictionary set into the multimodal large model for diagnosis to obtain the first diagnosis result; and determining a first warning parameter based on the first diagnosis result; determining a first similarity between the first analysis result and the first diagnosis result; determining a reasoning evaluation result of the first edge IoT agent on the first analysis result based on the first similarity; determining a first reward based on the reasoning evaluation result; and pushing the first reward to the first edge IoT agent; The first edge IoT agent is configured to perform warning operations and reinforcement learning based on the first diagnostic result and the first warning parameter; The cloud IoT agent is further specifically configured to: when the reasoning evaluation result includes that the first analysis result is wrong, execute the step of determining the first warning parameter according to the first diagnosis result; The first edge IoT agent is further specifically configured to: when the reasoning evaluation result includes that the first analysis result is correct, determine a second warning parameter corresponding to the first analysis result, and perform a warning operation according to the second warning parameter; In the aspect of performing intelligent analysis on the personnel image data and the environmental data to obtain the first analysis result, the first edge IoT agent is specifically configured to: Performing feature extraction on the person image data to obtain a first feature set; inputting the first feature set into a small behavior recognition model to obtain a first recognition result; performing feature extraction on the voice data to obtain a second feature set; inputting the second feature set into a small voiceprint recognition model to obtain a second recognition result; and determining the first analysis result based on the first recognition result and the second recognition result; In determining the first analysis result based on the first recognition result and the second recognition result, the first edge IoT agent is specifically configured to: determining a second similarity between the first recognition result and the second recognition result; When the second similarity is greater than a preset similarity, determining a first intersection result of the first recognition result and the second recognition result, and determining the first analysis result according to the first intersection result; When the second similarity is less than or equal to the preset similarity, concatenating the first feature set and the second feature set to obtain a third feature set; The third feature set is input into the local model of the multimodal large model to obtain the first analysis result.

2. The abnormal behavior identification and early warning system based on cloud-edge collaboration mechanism according to claim 1 is characterized in that: The multimodal large model includes a first-layer architecture and a second-layer architecture. The first-layer architecture adopts an encoder-decoder architecture to realize the recognition and understanding tasks of sound, image, and text, and outputs first text description information. The first text description information includes text description information for at least one of the sound, image, and text. The second-layer architecture includes an autoregressive Transormer architecture, which encodes and decodes the first text description information and outputs decoded second text description information. The second text description information is used to represent at least one of the following contents: whether there is abnormal behavior of the person, the category of abnormal behavior, and action decision information.

3. The abnormal behavior identification and early warning system based on cloud-edge collaboration mechanism according to claim 1 is characterized in that: In inputting the first feature set into the behavior recognition model to obtain the first recognition result, the first edge IoT agent is specifically configured to: Determining a first environment complexity corresponding to the sound data; determining a first model parameter corresponding to the first environmental complexity; The first recognition result is determined according to the first feature set, the first model parameters and the behavior recognition small model.

4. The abnormal behavior identification and early warning system based on cloud-edge collaboration mechanism according to claim 3 is characterized in that: The first recognition result includes a first abnormal behavior type and a first abnormal probability value; in inputting the second feature set into the voiceprint recognition model to obtain the second recognition result, the first edge IoT agent is specifically used to: determining a second model parameter corresponding to the first abnormal behavior type; determining a first adjustment parameter corresponding to the first abnormality probability value; Adjusting the second model parameters according to the first adjustment parameters to obtain target second model parameters; The second recognition result is determined according to the second feature set, the target second model parameters and the voiceprint recognition small model.

5. A method for identifying and warning abnormal behaviors based on a cloud-edge collaborative mechanism, characterized in that: The invention is applied to an abnormal behavior recognition and early warning system based on a cloud-edge collaborative mechanism, the system comprising: multiple edge IoT agents and a cloud IoT agent, wherein the cloud IoT agent is allowed to drive a pre-configured multimodal large model in a cloud platform; each of the multiple edge IoT agents is an edge IoT agent with autonomous learning capabilities; the method comprises: Acquiring personnel image data and environmental data through a first edge IoT agent, and performing intelligent analysis on the personnel image data and the environmental data to obtain a first analysis result, wherein the environmental data includes sound data and sensor data; the first edge IoT agent is any one of the multiple edge IoT agents; The cloud IoT agent pre-processes the person image data and the environment data to obtain first person image data and first environment data, specifically by: processing the sound data into a mel spectrum graph, performing semantic processing on the sensor data to automatically generate descriptive semantics of structured data to obtain first descriptive semantics, determining the first environment data based on the mel spectrum graph and the first descriptive semantics, and performing standardization processing on the person image data to obtain the first person image data; driving the multimodal large model to diagnose the first person image data, the first environment data, and the first analysis result to obtain a first diagnosis result, specifically by: determining a first relevant prompt word dictionary set based on the first descriptive semantics and the first analysis result, inputting the mel spectrum graph, the first person image data, and the first relevant prompt word dictionary set into the multimodal large model for diagnosis to obtain the first diagnosis result; and determining a first warning parameter based on the first diagnosis result; determining a first similarity between the first analysis result and the first diagnosis result; determining a reasoning evaluation result of the first edge IoT agent on the first analysis result based on the first similarity; determining a first reward based on the reasoning evaluation result; and pushing the first reward to the first edge IoT agent; Performing warning operations and reinforcement learning according to the first diagnostic result and the first warning parameter by the first edge IoT agent; When the reasoning evaluation result includes an error in the first analysis result, the cloud IoT agent executes the step of determining the first warning parameter according to the first diagnosis result; When the reasoning evaluation result includes the first analysis result being correct, determining, by the first edge IoT agent, a second warning parameter corresponding to the first analysis result, and performing a warning operation according to the second warning parameter; The intelligent analysis of the personnel image data and the environmental data to obtain a first analysis result includes: Performing feature extraction on the person image data to obtain a first feature set; inputting the first feature set into a small behavior recognition model to obtain a first recognition result; performing feature extraction on the voice data to obtain a second feature set; inputting the second feature set into a small voiceprint recognition model to obtain a second recognition result; and determining the first analysis result based on the first recognition result and the second recognition result; The determining the first analysis result according to the first recognition result and the second recognition result includes: determining a second similarity between the first recognition result and the second recognition result; When the second similarity is greater than a preset similarity, determining a first intersection result of the first recognition result and the second recognition result, and determining the first analysis result according to the first intersection result; When the second similarity is less than or equal to the preset similarity, concatenating the first feature set and the second feature set to obtain a third feature set; The third feature set is input into the local model of the multimodal large model to obtain the first analysis result.

6. The abnormal behavior identification and early warning method based on cloud-edge collaboration mechanism according to claim 5 is characterized in that: The multimodal large model includes a first-layer architecture and a second-layer architecture. The first-layer architecture adopts an encoder-decoder architecture to realize the recognition and understanding tasks of sound, image, and text, and outputs first text description information. The first text description information includes text description information for at least one of the sound, image, and text. The second-layer architecture includes an autoregressive Transormer architecture, which encodes and decodes the first text description information and outputs decoded second text description information. The second text description information is used to represent at least one of the following contents: whether there is abnormal behavior of the person, the category of abnormal behavior, and action decision information.

7. The abnormal behavior identification and early warning method based on cloud-edge collaboration mechanism according to claim 5 is characterized in that: Inputting the first feature set into the behavior recognition model to obtain a first recognition result includes: Determining a first environment complexity corresponding to the sound data; determining a first model parameter corresponding to the first environmental complexity; The first recognition result is determined according to the first feature set, the first model parameters and the behavior recognition small model.

8. The abnormal behavior identification and early warning method based on cloud-edge collaboration mechanism according to claim 7 is characterized in that: The first recognition result includes a first abnormal behavior type and a first abnormal probability value; the second feature set is input into the voiceprint recognition model to obtain a second recognition result, including: determining a second model parameter corresponding to the first abnormal behavior type; determining a first adjustment parameter corresponding to the first abnormality probability value; Adjusting the second model parameters according to the first adjustment parameters to obtain target second model parameters; The second recognition result is determined according to the second feature set, the target second model parameters and the voiceprint recognition small model.

Citation Information

Patent Citations

  • Abnormity detection and response method and system based on intelligent perception driving

    CN119544388A

  • Storage security protection and fire-fighting early warning method based on edge AI mode

    CN119694065A

  • Intelligent agent architecture based on multi-modal large model

    CN120046645A