Intelligent behavior identification method and system based on distributed sensor network

By using distributed sensor networks and prompt learning technology on edge computing devices, feature extraction and scenario adaptation of multimodal data is solved, and the shortcomings of existing behavior recognition methods in terms of efficiency and adaptability are achieved, and efficient, real-time and accurate behavior recognition is achieved.

CN120067758APending Publication Date: 2025-05-30INSPUR NETWORK TECH (SHANDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510151015.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing behavior recognition methods have shortcomings in terms of low data processing efficiency, insufficient scenario adaptability and weak multimodal fusion capabilities, which cannot meet the needs of real-time and efficient identification.

Method used

Using an intelligent behavior recognition method based on distributed sensor networks, edge computing and prompt learning technology is used to extract multimodal data through pre-trained behavior feature extraction models and behavior feature alignment modules, and embed scene prompt information into behavior feature vectors for classification recognition and global optimization.

Benefits of technology

It improves the data processing efficiency and real-time nature of behavior recognition, enhances the adaptability to different scenarios, and improves the robustness and recognition accuracy of multimodal data fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067758A_ABST
    Figure CN120067758A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent behavior recognition method and system based on a distributed sensor network, belongs to the technical field of artificial intelligence and edge computing, and is used for solving the defects of an existing behavior recognition method in the aspects of environmental adaptability, real-time performance, multi-modal fusion capability, data privacy and the like. The method comprises the following steps: acquiring multi-modal data which is acquired by multi-modal sensors deployed at different positions and is related to a to-be-identified behavior; inputting the multi-modal data into a pre-trained behavior feature extraction model, and performing feature extraction on the multi-modal data through a behavior feature alignment module of the behavior feature extraction model to obtain a behavior feature vector corresponding to the to-be-identified behavior; scene prompt information related to the behavior to be recognized is received, the scene prompt information is embedded into the behavior feature vector, and the scene prompt information is in a vector form; and performing classification identification on the behavior feature vector embedded with the scene prompt information to obtain an identification result corresponding to the to-be-identified behavior.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of artificial intelligence and edge computing, and particularly to an intelligent behavior recognition method and system based on a distributed sensor network. Background Art

[0002] In recent years, with the advancement of smart city construction, the application requirements of behavior recognition technology in the fields of public safety, industrial production, traffic management, etc. have been increasing day by day.

[0003] Currently, the mainstream behavior recognition methods include recognition methods based on visual features, multi-modal data, and cloud inference. Among them, the method based on visual features extracts spatio-temporal features through video streams for classification; the method based on multi-modal data combines video, audio, and sensor information to improve the recognition ability; the method based on cloud inference relies on large-scale models for centralized processing.

[0004] Although the above methods have achieved certain results in the field of behavior recognition, they still face the following problems in practical applications:

[0005] Low data processing efficiency: Traditional centralized behavior recognition systems need to upload all data to the cloud for processing, resulting in high latency and bandwidth pressure, and cannot meet the real-time requirements.

[0006] Insufficient scene adaptability: Existing models have limited performance when facing different scenarios (such as light changes, occlusion, environmental interference), and the generalization ability of the models needs to be improved urgently.

[0007] Weak multi-modal fusion ability: Fusing multi-modal data requires complex modeling and optimization, but existing methods still have deficiencies in terms of efficiency and robustness. Summary of the Invention

[0008] In order to solve the above problems, this application proposes an intelligent behavior recognition method and system based on a distributed sensor network, which uses edge computing and prompt learning technologies to achieve dynamic adaptation and efficient recognition of the model, and at the same time combines multi-modal fusion technology to further improve the recognition accuracy and application universality.

[0009] This application adopts the following technical solutions:

[0010] On the one hand, the present application provides an intelligent behavior recognition method based on a distributed sensor network, which is applied to an edge computing device. The method includes: obtaining multimodal data related to the behavior to be recognized collected by multimodal sensors deployed at different positions; inputting the multimodal data into a pre-trained behavior feature extraction model, and extracting features from the multimodal data through the behavior feature alignment module of the behavior feature extraction model to obtain a behavior feature vector corresponding to the behavior to be recognized; receiving scene prompt information related to the behavior to be recognized and embedding the scene prompt information into the behavior feature vector, where the scene prompt information is in vector form; classifying and recognizing the behavior feature vector embedded with the scene prompt information to obtain a recognition result corresponding to the behavior to be recognized.

[0011] In a possible implementation manner of the present application, after obtaining the recognition result corresponding to the behavior to be recognized, the method further includes: transmitting the recognition result corresponding to the behavior to be recognized to a central coordinator through a preset communication protocol; establishing a model aggregation mechanism using the federated learning framework on the central coordinator to globally optimize the recognition result corresponding to the behavior to be recognized through the model aggregation mechanism; determining a final recognition result corresponding to the behavior to be recognized according to the result of the global optimization.

[0012] In a possible implementation manner of the present application, obtaining multimodal data related to the behavior to be recognized collected by multimodal sensors deployed at different positions includes: obtaining multimodal data collected by multimodal sensors installed at static positions and dynamic positions; where the multimodal sensors include at least one of a camera, an accelerometer, and a pressure sensor, the static positions include at least one of intersections, public places, and indoor monitoring points, and the dynamic positions include at least one of mobile robots and vehicles.

[0013] In a possible implementation manner of the present application, after obtaining the multimodal data, the method further includes: preprocessing the obtained multimodal data, including: using real-time filtering and denoising algorithms to eliminate data interference in the multimodal data; and / or, performing preliminary feature extraction on the multimodal data by extracting key frames from the video data collected by the camera and detecting peaks in the acceleration data collected by the accelerometer; and / or, compressing the multimodal data.

[0014] In a possible implementation manner of the present application, the behavior feature extraction model adopts a Transformer model; the multi-modal data is subjected to feature extraction through the behavior feature alignment module, including: in the behavior feature alignment module, a multi-layer perceptron MLP and / or a convolutional neural network CNN are used to perform feature extraction on the multi-modal data from the multi-modal sensor.

[0015] In a possible implementation manner of the present application, before receiving the scene prompt information related to the behavior to be recognized, the method further includes: obtaining text information related to the behavior to be recognized, where the text information is related to the behavior scene of the behavior to be recognized; converting the text information into a vector representation through a text encoder to obtain the scene prompt information.

[0016] In a possible implementation manner of the present application, after converting the text information into a vector representation through the text encoder, the method further includes: learning global prompts and local spatial location prompts related to the behavior scene; adjusting the scene prompt information by introducing a prompt dropout strategy.

[0017] In a possible implementation manner of the present application, global optimization is performed on the recognition result corresponding to the behavior to be recognized through the model aggregation mechanism, including: using the model aggregation mechanism to synchronize the model parameters in the behavior feature extraction model.

[0018] In a possible implementation manner of the present application, the multi-modal data is subjected to feature extraction through the behavior feature alignment module of the behavior feature extraction model, including: using the behavior feature alignment module to perform feature extraction on the multi-modal data to obtain a behavior feature vector that fuses the video data and the acceleration data.

[0019] On the other hand, the present application further provides an intelligent behavior recognition system based on a distributed sensor network. The system includes a multi-modal sensor, an edge computing device, and a central coordinator. Among them, the edge computing device can execute: obtaining multi-modal data related to the behavior to be recognized collected by multi-modal sensors deployed at different positions; inputting the multi-modal data into a pre-trained behavior feature extraction model, and performing feature extraction on the multi-modal data through the behavior feature alignment module of the behavior feature extraction model to obtain a behavior feature vector corresponding to the behavior to be recognized; receiving scene prompt information related to the behavior to be recognized and embedding the scene prompt information into the behavior feature vector, where the scene prompt information is in vector form; classifying and recognizing the behavior feature vector embedded with the scene prompt information to obtain a recognition result corresponding to the behavior to be recognized.

[0020] An intelligent behavior recognition method and system based on a distributed sensor network provided by the present application have the following beneficial effects:

[0021] In the present application, the edge computing device processes the multimodal data collected by the distributed sensors (i.e., multimodal sensors), and uses the behavior feature recognition model to extract features from the multimodal data to obtain the behavior feature vector of the behavior to be recognized, avoiding the problem of low data processing efficiency caused by the need to upload data to the cloud in the traditional centralized behavior recognition scheme. The edge computing device can directly extract data features from the multimodal data, improving the data processing efficiency in the behavior recognition process, thereby enhancing the real-time performance of behavior recognition. On the other hand, the present application also embeds the scene prompt information into the behavior feature vector, which can help to adaptively adjust different behavior features when dealing with new scenes, enhancing the scene adaptation ability of the behavior recognition scheme.

[0022] In summary, by combining edge computing and prompt learning technologies, the present application realizes accurate and efficient behavior recognition in dynamic scenarios, effectively enhancing the generalization ability and real-time processing ability of the behavior recognition scheme, while taking into account data privacy and security. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0024] Figure 1 It is a flowchart of an intelligent behavior recognition method based on a distributed sensor network provided by the present application;

[0025] Figure 2 It is Embodiment 1 of the process of an intelligent behavior recognition method based on a distributed sensor network provided by the present application;

[0026] Figure 3 It is an architecture diagram of an intelligent behavior recognition system based on a distributed sensor network provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] To enable those skilled in the art to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in this application in conjunction with the accompanying drawings in this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0028] To address the deficiencies of existing behavior recognition methods in aspects such as environmental adaptability, real-time performance, multi-modal fusion ability, and data privacy, this application provides an intelligent behavior recognition method based on a distributed sensor network. By combining edge computing and prompt learning technologies, this method achieves accurate and efficient behavior recognition in dynamic scenarios, effectively enhancing the generalization ability and real-time processing ability of the model, while taking into account data privacy and security.

[0029] The following will detail the method in this application through the accompanying drawings.

[0030] Figure 1 The flowchart of an intelligent behavior recognition method based on a distributed sensor network provided by this application is as Figure 1 shown. The method in this application at least includes the following execution steps:

[0031] Step 101: Obtain multi-modal data related to the behavior to be recognized collected by multi-modal sensors deployed at different locations.

[0032] The intelligent behavior recognition method proposed in this application is implemented by an edge computing device.

[0033] First, data related to the behavior to be recognized is collected through distributed sensors. Here, the distributed sensors are multi-modal sensors, including at least one of a camera, an accelerometer, and a pressure sensor. Of course, other types of sensors can also be included. These sensors are deployed at different static and dynamic locations. Among them, the static locations include at least one of intersections, public places, and indoor monitoring points, and the dynamic locations include at least one of mobile robots and vehicles. The behavior is recognized through the collected multi-modal data.

[0034] In a possible implementation of the present application, after obtaining multi-modal data, in order to ensure the accuracy and efficiency of subsequent behavior recognition, the obtained multi-modal data will be preprocessed. The preprocessing process includes: using real-time filtering and denoising algorithms to eliminate data interference in the multi-modal data; and / or, performing preliminary feature extraction on the multi-modal data by extracting key frames from the video data collected by the camera and detecting peaks in the acceleration data collected by the accelerometer; and / or, compressing the multi-modal data. It should be noted that the data feature extraction, data compression, etc. processes during the preprocessing of the multi-modal data, as well as the real-time filtering and denoising algorithms used, can all be implemented through existing algorithms or solutions, and will not be elaborated herein.

[0035] Step 102: Input the multi-modal data into the pre-trained behavior feature extraction model, and extract features from the multi-modal data through the behavior feature alignment module of the behavior feature extraction model to obtain a behavior feature vector corresponding to the behavior to be recognized.

[0036] Furthermore, input the preprocessed multi-modal data into the pre-trained behavior feature extraction model. Here, the behavior feature recognition model is implemented using a Transformer model. After the Transformer model is pre-trained, the aforementioned multi-modal data is input into it, and the behavior feature alignment module in the model is used to extract features to obtain a behavior feature vector corresponding to the behavior to be recognized. The present application benefits from the fact that the Transformer model has good spatial feature extraction ability and temporal feature extraction ability, can simultaneously capture the spatial feature changes of video frames and the temporal changes of actions, and the selected Transformer model is usually pre-trained on a large-scale dataset and has good feature extraction ability.

[0037] In a possible implementation of the present application, the aforementioned behavior feature alignment module uses a multi-layer perceptron MLP and / or a convolutional neural network CNN to extract features from the multi-modal data from multi-modal sensors.

[0038] Through the aforementioned feature alignment module, a unified behavior feature vector that can represent all sensor data can be generated. This behavior feature vector synthesizes information from various dimensions such as video data, acceleration data, and environmental sensor data, and can provide a more comprehensive scene description. That is, in the feature alignment module, the multi-layer perceptron MLP or the convolutional neural network CNN is used to extract features from the data from different sensors. Through this process, the model can learn the implicit mapping between the data sources of each sensor.

[0039] Step 103: Receive scene prompt information related to the behavior to be recognized and embed the scene prompt information into the behavior feature vector.

[0040] To improve the scene adaptation ability of the solution of this application, after obtaining the behavior feature vector, scene prompt information will also be embedded into the behavior feature vector.

[0041] Specifically, learn globally consistent prompts and locally spatial positioning prompts that are consistent with vision on different edge computing devices, and strengthen the diversity between them to improve their combination. Introduce a new prompt dropout strategy to introduce diversity through randomization. Convert text information into vector representations through a prompt encoder / text encoder, and then combine these vectors with the aforementioned behavior feature vectors.

[0042] Step 104: Classify and identify the behavior feature vector embedded with scene prompt information to obtain the recognition result corresponding to the behavior to be recognized.

[0043] Furthermore, input the behavior feature vector combined with scene prompt information into a classifier to obtain the recognition result corresponding to the behavior to be recognized. This can help the model make adaptive adjustments to different input features when processing new scenes.

[0044] In a possible implementation manner of this application, after obtaining the recognition result corresponding to the behavior to be recognized, the method further includes: transmitting the recognition result corresponding to the behavior to be recognized to a central coordinator through a preset communication protocol, establishing a model aggregation mechanism using the federated learning framework on the central coordinator, and globally optimizing the recognition result corresponding to the behavior to be recognized through the model aggregation mechanism. For example, synchronize the model parameters in the behavior feature extraction model using the model aggregation mechanism; finally, determine the final recognition result corresponding to the behavior to be recognized according to the result of the global optimization.

[0045] Specifically, the central coordinator can receive the behavior recognition results from all edge computing devices. These behavior recognition results are integrated into a federated learning framework. Using the mechanism of distributed collaborative learning, the federated learning framework synchronizes the model parameters on each edge computing device by establishing a model aggregation mechanism. Finally, through model aggregation, the central coordinator can synthesize the recognition results of each edge computing device and perform global optimization to obtain the final recognition result.

[0046] Figure 2 This is Embodiment 1 of a method flow for intelligent behavior recognition based on a distributed sensor network provided by this application. As Figure 2 shown, the behavior recognition solution in this application may further include the following steps:

[0047] Step 1: Distributed sensor data collection

[0048] Multimodal data related to behaviors are collected through multimodal sensors (such as cameras, accelerometers, pressure sensors, etc.) deployed at different locations. The sensor nodes perform preprocessing on the originally collected multimodal data in an edge computing manner, including operations such as noise reduction, feature extraction, and data compression.

[0049] Step 2: Multimodal data fusion

[0050] The preprocessed multimodal data are fed into a pre-trained Transformer model for data feature extraction, and effective fusion of multimodal data is achieved through the feature alignment module in the model. This feature alignment module uses a multi-layer perceptron (MLP) or a convolutional neural network (CNN) to generate a unified behavioral feature representation.

[0051] Step 3: Prompt learning based on edge computing

[0052] On the edge computing device, the pre-trained model is fine-tuned for the scenario using prompt learning technology. Specifically, the text information is encoded by a text encoder to obtain a text information vector, that is, the prompt information related to the scenario is generated. After being processed by the introduced dropout strategy, it is embedded into the behavioral features to obtain a unified behavioral feature vector, so as to improve the adaptability to dynamic scenarios.

[0053] Step 4: Distributed collaborative inference

[0054] The unified behavioral feature vector is recognized by a classifier to obtain the recognition result.

[0055] The recognition results of each edge computing device are transmitted to the central coordinator through a low-latency communication protocol. The global information is integrated using a federated learning framework to further optimize the behavioral recognition results and obtain the final recognition result, such as Figure 2 "clapping hands" shown in

[0056] Based on the same inventive concept, the present application also provides an intelligent behavior recognition system based on a distributed sensor network, and its architecture is as Figure 3 shown.

[0057] Figure 3 This is the architecture diagram of an intelligent behavior recognition system based on a distributed sensor network provided by the present application. As Figure 3As shown in the figure, the intelligent behavior recognition system 300 in this application includes: a multi-modal sensor 301, an edge computing device 302, and a central coordinator 303. Among them, the edge computing device 302 can execute: obtaining multi-modal data related to the behavior to be recognized collected by multi-modal sensors deployed at different positions; inputting the multi-modal data into a pre-trained behavior feature extraction model, and performing feature extraction on the multi-modal data through the behavior feature alignment module of the behavior feature extraction model to obtain a behavior feature vector corresponding to the behavior to be recognized; receiving scene prompt information related to the behavior to be recognized, and embedding the scene prompt information into the behavior feature vector, where the scene prompt information is in vector form; classifying and recognizing the behavior feature vector embedded with the scene prompt information to obtain a recognition result corresponding to the behavior to be recognized.

[0058] The various embodiments in this application are all described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0059] The devices and methods provided in this application correspond one by one. Therefore, the devices also have beneficial technical effects similar to their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices will not be elaborated here.

[0060] Those skilled in the art should understand that the embodiments of this application can be provided as methods, devices, systems, or computer program products. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0061] It should also be noted that the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, commodity, or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity, or device including the element.

[0062] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. An intelligent behavior recognition method based on a distributed sensor network, applied to edge computing devices, characterized in that: The method comprises: Acquire multimodal data related to the behavior to be identified collected by multimodal sensors deployed at different locations; Inputting the multimodal data into a pre-trained behavior feature extraction model, extracting features from the multimodal data through a behavior feature alignment module of the behavior feature extraction model, and obtaining a behavior feature vector corresponding to the behavior to be identified; Receiving scene prompt information related to the behavior to be identified, and embedding the scene prompt information into the behavior feature vector, wherein the scene prompt information is in vector form; The behavior feature vector embedded with the scene prompt information is classified and identified to obtain an identification result corresponding to the behavior to be identified.

2. The intelligent behavior recognition method based on a distributed sensor network according to claim 1 is characterized in that: After obtaining the recognition result corresponding to the behavior to be recognized, the method further includes: Transmitting the recognition result corresponding to the behavior to be recognized to the central coordinator through a preset communication protocol; A model aggregation mechanism is established by using the federated learning framework on the central coordinator, so as to globally optimize the recognition results corresponding to the behaviors to be recognized through the model aggregation mechanism; According to the result of the global optimization, a final recognition result corresponding to the behavior to be recognized is determined.

3. The intelligent behavior recognition method based on a distributed sensor network according to claim 1 is characterized in that: Obtain multimodal data related to the behavior to be identified collected by multimodal sensors deployed at different locations, including: Acquire multimodal data collected by multimodal sensors installed in static and dynamic positions; Among them, the multimodal sensor includes at least one of a camera, an accelerometer and a pressure sensor, the static position includes at least one of an intersection, a public place and an indoor monitoring point, and the dynamic position includes at least one of a mobile robot and a vehicle.

4. The intelligent behavior recognition method based on a distributed sensor network according to claim 3 is characterized in that: After acquiring the multimodal data, the method further includes: Preprocessing the acquired multimodal data includes: Eliminating data interference in the multimodal data using real-time filtering and denoising algorithms; and / or, performing preliminary feature extraction on the multimodal data by performing key frame extraction on the video data collected by the camera and peak detection on the acceleration data collected by the accelerometer; And / or, performing data compression on the multimodal data.

5. The intelligent behavior recognition method based on a distributed sensor network according to claim 1 is characterized in that: The behavior feature extraction model adopts the Transformer model; Extracting features from the multimodal data by using the behavior feature alignment module includes: In the behavior feature alignment module, a multi-layer perceptron MLP and / or a convolutional neural network CNN are used to extract features from the multi-modal data from the multi-modal sensor.

6. The intelligent behavior recognition method based on a distributed sensor network according to claim 1 is characterized in that: Before receiving scene prompt information related to the behavior to be identified, the method further includes: Acquire text information related to the behavior to be identified, where the text information is related to a behavior scenario of the behavior to be identified; The text information is converted into a vector representation through a text encoder to obtain the scene prompt information.

7. The intelligent behavior recognition method based on a distributed sensor network according to claim 6 is characterized in that: After converting the text information into a vector representation by a text encoder, the method further includes: Learning global cues and local spatial positioning cues relevant to the behavioral scenario; The scene prompt information is adjusted by introducing a prompt dropout strategy.

8. The intelligent behavior recognition method based on a distributed sensor network according to claim 2 is characterized in that: The recognition result corresponding to the behavior to be recognized is globally optimized through the model aggregation mechanism, including: The model aggregation mechanism is used to synchronize model parameters in the behavior feature extraction model.

9. The intelligent behavior recognition method based on a distributed sensor network according to claim 4 is characterized in that: Performing feature extraction on the multimodal data through a behavior feature alignment module of the behavior feature extraction model includes: The behavior feature alignment module is used to extract features from the multimodal data to obtain a behavior feature vector that fuses the video data and the acceleration data.

10. An intelligent behavior recognition system based on a distributed sensor network, characterized in that: The system includes a multimodal sensor, an edge computing device, and a central coordinator, wherein the edge computing device is capable of performing: Acquire multimodal data related to the behavior to be identified collected by multimodal sensors deployed at different locations; Inputting the multimodal data into a pre-trained behavior feature extraction model, extracting features from the multimodal data through a behavior feature alignment module of the behavior feature extraction model, and obtaining a behavior feature vector corresponding to the behavior to be identified; Receiving scene prompt information related to the behavior to be identified, and embedding the scene prompt information into the behavior feature vector, wherein the scene prompt information is in vector form; The behavior feature vector embedded with the scene prompt information is classified and identified to obtain an identification result corresponding to the behavior to be identified.