Behavior recognition large model generation method, behavior recognition method, device, equipment and medium

By fine-tuning and optimizing the pre-trained multimodal large model, a quantized target behavior recognition large model is generated, which solves the problems of insufficient accuracy and scene generalization ability of personnel behavior recognition under multimodal data, and achieves high accuracy and low cost recognition effect.

CN120853252APending Publication Date: 2025-10-28SHENZHEN XIAOPAI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510853796.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of human behavior recognition under multimodal data is low, the scene generalization ability is weak, and it is difficult to achieve accurate recognition in a variety of scenarios.

Method used

By acquiring multimodal training samples, and based on multimodal behavior data and labeled behavior tags, the pre-trained multimodal large model is fine-tuned and optimized to generate a fine-tuned and optimized behavior recognition large model, which is then quantized to generate a target behavior recognition large model.

Benefits of technology

It improves the accuracy and stability of human behavior recognition in multimodal data, enhances the model's scene generalization ability, and reduces storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853252A_ABST
    Figure CN120853252A_ABST
Patent Text Reader

Abstract

The invention discloses a behavior recognition large model generation method, a behavior recognition method, a behavior recognition device, behavior recognition equipment and a medium. The method comprises the following steps: acquiring a plurality of multi-modal training samples, wherein each multi-modal training sample comprises multi-modal behavior data and a labeled behavior label; on the basis of the multi-modal behavior data and the labeled behavior labels in the multiple multi-modal training samples, performing fine tuning on a pre-training weight in the pre-training multi-modal large model, and generating a fine tuning behavior recognition large model; based on the multi-modal behavior data and the labeled behavior labels in the multiple multi-modal training samples, performing optimization training on the fine-tuning behavior recognition large model to generate an optimized behavior recognition large model; and performing quantitative processing on the optimized behavior recognition large model to generate a target behavior recognition large model. The target behavior recognition large model generated by the method has strong scene generalization ability, has high stability and accuracy in recognition of personnel behaviors in multi-modal data, and can also reduce the model storage cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and intelligent sensing technology, and in particular to a method for generating large behavior recognition models, a behavior recognition method, a device, equipment, and a medium. Background Technology

[0002] With the continuous development of technology, scenarios such as smart homes, public safety, traffic management, medical care, and human-computer interaction all require the identification of human behavior in multimodal data to achieve the functions required in these scenarios and improve the convenience and safety of human life. For example, in the smart home scenario, it is necessary to identify human behavior in the data collected within the smart home environment to determine the specific behaviors within that scenario. This allows for the assessment of whether abnormal behaviors, such as fighting or falls, exist, enabling real-time monitoring of the smart home environment.

[0003] In existing technologies, object detection or behavior classification models using a single visual modality are typically used to identify human behavior from multimodal data across multiple scenarios. However, the accuracy of these identifications is low, meaning that existing technologies have weak scene generalization capabilities and poor overall accuracy. Therefore, how to accurately identify human behavior from multimodal data across multiple scenarios is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] This invention provides a method for generating a large behavior recognition model, a behavior recognition method, a device, equipment, and a medium to solve the technical problem of how to accurately identify human behavior in multimodal data under various scenarios.

[0005] A method for generating large behavior recognition models includes: Multiple multimodal training samples are obtained, each of which includes multimodal behavioral data and labeled behavioral tags corresponding to the multimodal behavioral data; Based on the multimodal behavior data and labeled behavior tags in multiple multimodal training samples, the pre-training weights in the pre-trained multimodal large model are fine-tuned to generate a fine-tuned behavior recognition large model. Based on the multimodal behavior data and labeled behavior tags in multiple multimodal training samples, the fine-tuned behavior recognition model is optimized and trained to generate an optimized behavior recognition model. The optimized behavior recognition model is quantized to generate a target behavior recognition model.

[0006] Preferably, acquiring multiple multimodal training samples includes: Obtain an initial multimodal behavior dataset, and perform data filtering processing on the initial multimodal behavior dataset according to multiple preset behaviors to determine the filtered behavior dataset corresponding to each preset behavior; Perform data clustering analysis on each of the aforementioned filtering behavior datasets to determine multiple clustered data subsets corresponding to each of the aforementioned filtering behavior datasets; For each clustered data subset, perform data deduplication to determine the multimodal behavioral data corresponding to each clustered data subset; The multimodal behavior data is labeled to determine the labeled behavior label corresponding to each multimodal behavior data. Based on each of the multimodal behavioral data and the corresponding labeled behavioral tags, multiple multimodal training samples are determined.

[0007] Preferably, the step of fine-tuning the pre-training weights in the pre-trained multimodal large model based on the multimodal behavior data and the labeled behavior tags from multiple multimodal training samples to generate a fine-tuned behavior recognition large model includes: Key layer identification is performed on the pre-trained multimodal large model to determine the target key layer; Based on the pre-trained weights of the target key layer, determine the weight matrix to be updated corresponding to the pre-trained weights; Based on the multiple multimodal behavior data and the labeled behavior tags, the weight matrix to be updated is fine-tuned to generate a fine-tuned behavior recognition model.

[0008] Preferably, the step of fine-tuning the weight matrix to be updated based on multiple sets of multimodal behavior data and the labeled behavior tags to generate a fine-tuned behavior recognition model includes: A preset low-rank matrix is ​​inserted into the target key layer, the multimodal behavior data is input into the pre-trained multimodal large model, and the initial recognition result is output. Based on the initial recognition results and the labeled behavior tags corresponding to the multimodal behavior data, the preset low-rank matrix is ​​updated to determine the first update matrix; Based on the first update matrix, the weight matrix to be updated is fine-tuned to determine the fine-tuned weight matrix and the corresponding update behavior recognition model, and the current number of fine-tunings is determined. If the current number of fine-tunings has not reached the preset number of fine-tunings, then the fine-tuning weight matrix is ​​updated to the preset low-rank matrix, the updated behavior recognition big model is updated to the pre-trained multimodal big model, the process of inserting the preset low-rank matrix in the target key layer is repeated, the multimodal behavior data is input into the pre-trained multimodal big model, and the initial recognition result is output. If the current number of fine-tunings reaches the preset number of fine-tunings, then the updated behavior recognition model is determined as the fine-tuned behavior recognition model.

[0009] Preferably, the step of optimizing and training the fine-tuned behavior recognition model based on the multimodal behavior data and the labeled behavior tags from multiple multimodal training samples to generate an optimized behavior recognition model includes: The multimodal behavior data is input into the fine-tuned behavior recognition model, and the reference log probability corresponding to the fine-tuned behavior recognition model is output. The multimodal behavior data is input into the current policy model, and the K current behavior recognition results corresponding to the multimodal behavior data and the current log probability corresponding to each current behavior recognition result are output, where K > 1; Based on the K current behavior recognition results and the labeled behavior tags, K dominant functions are determined; Based on the reference log probability and K current log probabilities, determine K KL divergences; The target loss function value is determined based on the reference log probability, K of the dominance functions, K of the current log probabilities, and K of the KL divergences. If the target loss function value does not meet the preset convergence condition, then the process of inputting the multimodal behavior data into the fine-tuned behavior recognition model is repeated. If the target loss function value satisfies the preset convergence condition, then the current strategy model is determined as the large-scale optimization behavior recognition model; The current strategy model is the model after each optimization during the optimization training process of the fine-tuned behavior recognition large model.

[0010] Preferably, determining the target loss function value based on the reference log probability, K of the dominance functions, K of the current log probabilities, and K of the KL divergences includes: Based on the reference log probability, the advantage function corresponding to each current behavior recognition result, and the current log probability, determine K first loss function values ​​corresponding to the current behavior recognition results; The KL divergence corresponding to each current behavior recognition result is corrected to determine K second loss function values ​​corresponding to the current behavior recognition results; Based on K first loss function values ​​and K second loss function values, determine K in-group loss function values ​​corresponding to the current behavior recognition results; The target loss function value is determined based on the K intra-group loss function values ​​corresponding to the current behavior recognition results.

[0011] A behavior recognition method, comprising: Acquire the behavior data to be identified; The behavior data to be identified is input into the target behavior recognition big model generated by the above-mentioned behavior recognition big model generation method, and behavior recognition is performed to determine the target behavior corresponding to the behavior data to be identified.

[0012] A device for generating large behavior recognition models, comprising: A multimodal training sample acquisition module is used to acquire multiple multimodal training samples, each of which includes multimodal behavior data and labeled behavior tags corresponding to the multimodal behavior data; The fine-tuning behavior recognition large model generation module fine-tunes the pre-training weights in the pre-trained multimodal large model based on the multimodal behavior data and the labeled behavior tags in multiple multimodal training samples, thereby generating a fine-tuned behavior recognition large model. The optimized behavior recognition large model generation module optimizes and trains the fine-tuned behavior recognition large model based on the multimodal behavior data and the labeled behavior tags in multiple multimodal training samples, thereby generating an optimized behavior recognition large model. The target behavior recognition large model generation module is used to quantize the optimized behavior recognition large model to generate a target behavior recognition large model.

[0013] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described behavior recognition large model generation method, or the processor implements the above-described behavior recognition method when executing the computer program.

[0014] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described behavior recognition large model generation method, or, when executed by a processor, implements the above-described behavior recognition method.

[0015] The aforementioned method, apparatus, and medium for generating large-scale behavior recognition models, through fine-tuning a pre-trained multimodal large-scale model using multimodal behavior data and labeled behavior training samples, effectively improves the accuracy of the fine-tuned behavior recognition model in recognizing human behavior in multimodal behavior data. Further reinforcement training of the fine-tuned model enhances both the accuracy and stability of the optimized model in recognizing human behavior in different scenarios. Quantization of the optimized model ensures that the resulting target behavior recognition model maintains the scene generalization ability and recognition accuracy of the pre-quantization optimized model while reducing its storage cost. The target behavior recognition model generated by this method not only possesses strong scene generalization ability and high stability and accuracy in recognizing human behavior in multimodal data, but also effectively reduces storage costs, demonstrating broad application prospects. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of a method for generating a large behavior recognition model according to an embodiment of the present invention; Figure 2 This is another flowchart of the behavior recognition large model generation method in one embodiment of the present invention; Figure 3 This is another flowchart of the behavior recognition large model generation method in one embodiment of the present invention; Figure 4 This is another flowchart of the behavior recognition large model generation method in one embodiment of the present invention; Figure 5 This is another flowchart of the behavior recognition large model generation method in one embodiment of the present invention; Figure 6 This is another flowchart of the behavior recognition large model generation method in one embodiment of the present invention; Figure 7 This is another flowchart of the behavior recognition large model generation method in one embodiment of the present invention; Figure 8 This is a flowchart of a behavior recognition method according to an embodiment of the present invention; Figure 9This is a schematic diagram of a behavior recognition large model generation device in one embodiment of the present invention; Figure 10 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] This invention provides a method for generating a large-scale behavior recognition model, used to accurately identify human behavior in multimodal data across various scenarios. This method generates a target behavior recognition model with strong scenario generalization ability and high recognition accuracy. It is applicable to various scenarios, including but not limited to smart homes, public safety, traffic management, medical care, and human-computer interaction, and can accurately identify human behavior from large amounts of multimodal data.

[0020] In one embodiment, such as Figure 1 As shown, a method for generating a large behavior recognition model is provided, and this method is applied to... Figure 10 Taking a computer device as an example, the explanation includes the following steps: S101: Obtain multiple multimodal training samples, each multimodal training sample including multimodal behavior data and corresponding labeled behavior tags; S102: Based on multimodal behavior data and labeled behavior tags from multiple multimodal training samples, fine-tune the pre-trained weights in the pre-trained multimodal large model to generate a fine-tuned behavior recognition large model; S103: Based on multimodal behavior data and labeled behavior tags from multiple multimodal training samples, optimize and train the fine-tuned behavior recognition model to generate an optimized behavior recognition model; S104: Quantize the optimized behavior recognition model to generate a target behavior recognition model.

[0021] Multimodal training samples refer to samples used for model fine-tuning and optimization training. Multimodal behavioral data refers to multimodal data corresponding to multiple scenarios, including but not limited to video, images, text, and audio. Annotated behavioral labels refer to the labels corresponding to human behaviors within the multimodal behavioral data.

[0022] As an example, in step S101, the computer device acquires multiple multimodal behavioral data from various scenarios, including but not limited to video, images, text, and audio from scenarios such as smart homes, public safety, traffic management, medical care, and human-computer interaction. The computer device labels the human behavior in each multimodal behavioral data point, determines the corresponding labeled behavior label for each multimodal behavioral data point, and stores a multimodal behavioral data point and its corresponding labeled behavior label to form a multimodal training sample. The above operation is performed on each multimodal behavioral data point and its corresponding labeled behavior label to obtain multiple multimodal training samples. In this example, multimodal training samples, including multimodal behavioral data and corresponding labeled behavior labels, are acquired from various scenarios for subsequent model training to improve the model's scenario generalization ability and the accuracy of human behavior recognition.

[0023] In this context, a pre-trained multimodal large model refers to a large model that has been pre-trained and possesses a strong ability to understand multimodal data. Pre-trained weights refer to the weights in a pre-trained multimodal large model.

[0024] As an example, in step S102, the computer device inputs the multimodal behavior data from the multimodal training samples into the pre-trained multimodal large model, outputs the personnel behavior recognition result corresponding to the multimodal behavior data, determines the loss function value corresponding to the pre-trained multimodal large model based on the personnel recognition result and the labeled behavior label corresponding to the multimodal behavior data, determines the gradient of model fine-tuning based on the loss function value, and fine-tunes the pre-training weights in the pre-trained multimodal large model based on the gradient to obtain the fine-tuned pre-trained multimodal large model. The multimodal behavior data from another multimodal training sample is input into the fine-tuned pre-trained multimodal large model, and the above steps are repeated to fine-tune the pre-training weights in the fine-tuned pre-trained multimodal large model until the preset number of fine-tuning times is reached, then the fine-tuning is stopped, and the most recently fine-tuned pre-trained multimodal large model is determined as the fine-tuned behavior recognition large model. Understandably, pre-trained multimodal large models are obtained through pre-training on massive amounts of multimodal data and have a strong ability to understand multimodal data. However, pre-trained multimodal large models are not pre-trained for human behavior recognition. Therefore, pre-trained multimodal large models have poor ability to recognize human behavior in multimodal behavior data. In this example, by fine-tuning the pre-trained multimodal large model for the task of human behavior recognition, the accuracy of the fine-tuned behavior recognition large model can be effectively improved in recognizing human behavior in multimodal behavior data.

[0025] Among them, the optimized behavior recognition model refers to the model after optimizing and training the fine-tuned behavior recognition model.

[0026] As an example, in step S103, the computer device inputs the multimodal behavior data from the multimodal training samples into the fine-tuned behavior recognition model generated by fine-tuning, outputs the personnel behavior recognition results corresponding to the multimodal behavior data, uses a reinforcement learning algorithm to calculate the loss function value of the multimodal behavior data and the personnel behavior recognition results corresponding to the multimodal behavior data, determines the loss function value corresponding to the fine-tuned behavior recognition model, updates the fine-tuned behavior recognition model according to the loss function value corresponding to the fine-tuned behavior recognition model, obtains the updated fine-tuned behavior recognition model, and determines whether the loss function value corresponding to the fine-tuned behavior recognition model has converged. If it is determined that the loss function value corresponding to the fine-tuned behavior recognition model has not converged, the above steps are repeated until the loss function value converges. The most recently updated fine-tuned behavior recognition model is determined as the optimized behavior recognition model. If it is determined that the loss function value corresponding to the fine-tuned behavior recognition model has converged, the updated fine-tuned behavior recognition model is directly determined as the optimized behavior recognition model. In this example, the fine-tuned behavior recognition model is reinforced through training. This improves the accuracy of the optimized behavior recognition model in recognizing human behavior in different scenarios, while also enhancing the stability of the optimized behavior recognition model. In other words, it enhances the model's generalization ability across multiple scenarios and improves the stability and accuracy of the model in recognizing human behavior.

[0027] Among them, the target behavior recognition big model refers to the model after quantizing the optimized behavior recognition big model.

[0028] As an example, in step S104, the computer device uses a preset quantization method to quantize the optimized behavior recognition model. This ensures the accuracy of the optimized behavior recognition model while compressing its size, resulting in the quantized target behavior recognition model. Understandably, while the optimized behavior recognition model obtained after fine-tuning and optimizing the pre-trained multimodal model possesses strong scene generalization ability and high recognition accuracy, its excessively large parameters lead to problems such as high resource consumption, high response latency, and significant power consumption in actual deployment. Therefore, quantizing the optimized behavior recognition model ensures that the target behavior recognition model maintains the scene generalization ability and recognition accuracy of the pre-quantized optimized behavior recognition model while reducing its storage cost.

[0029] In this embodiment, a pre-trained multimodal large-scale model is fine-tuned using multimodal behavior data and labeled behavior training samples. This effectively improves the accuracy of the fine-tuned behavior recognition model in identifying human behavior in multimodal behavior data. The fine-tuned model is then reinforced through training, enhancing both the accuracy and stability of the optimized model in different scenarios. Finally, the optimized model is quantized, ensuring that the target behavior recognition model maintains the scene generalization ability and accuracy of the pre-quantized optimized model while reducing its storage cost. The target behavior recognition model generated by this method not only possesses strong scene generalization ability and high stability and accuracy in identifying human behavior in multimodal data, but also effectively reduces storage costs, demonstrating broad application prospects.

[0030] In one embodiment, such as Figure 2 As shown, step S101, which involves acquiring multiple multimodal training samples, includes: S201: Obtain the initial multimodal behavior dataset, perform data filtering processing on the initial multimodal behavior dataset according to multiple preset behaviors, and determine the filtering behavior dataset corresponding to each preset behavior; S202: Perform data clustering analysis on each filtering behavior dataset to determine multiple clustered data subsets corresponding to each filtering behavior dataset; S203: Perform data deduplication on each clustered data subset to determine the multimodal behavioral data corresponding to each clustered data subset; S204: Perform annotation processing on multimodal behavior data to determine the annotation behavior label corresponding to each multimodal behavior data; S205: Based on each multimodal behavior data and the corresponding labeled behavior label, determine multiple multimodal training samples.

[0031] The initial multimodal behavior dataset refers to the dataset formed by unprocessed multimodal data corresponding to various scenarios. Preset behaviors refer to human behaviors used for data filtering, such as falling, running, and hitting. The filtered behavior dataset is the dataset obtained after filtering the initial multimodal behavior dataset according to the preset behaviors.

[0032] As an example, in step S201, the computer device obtains multiple multimodal data by crawling the web, legally downloading public datasets, and manually capturing images from multiple scenarios. This multimodal data constitutes an initial multimodal behavior dataset. The computer device then filters the data in the initial multimodal behavior dataset according to multiple preset behaviors, identifying multiple data points corresponding to each preset behavior. This data is then designated as the filtered behavior dataset for each preset behavior. In essence, each preset behavior corresponds to a filtered behavior dataset. For example, when the computer device determines that the preset behavior is a fall, it uses an InternVL multimodal large model to filter the data in the initial multimodal behavior dataset, obtaining multiple data points of different modalities. This multiple data points of different modalities are then designated as the filtered behavior dataset corresponding to the fall behavior. In this example, the initial multimodal behavior dataset is filtered according to preset behaviors, selecting the filtered behavior dataset consistent with the preset behaviors, removing redundant data, and retaining valid data.

[0033] Clustered data subsets refer to subsets of data with high similarity in the filtering behavior dataset.

[0034] As an example, in step S202, the computer device uses a preset clustering algorithm to perform data clustering analysis on the data in each screening behavior dataset, determining multiple clustered data subsets in each screening behavior dataset. In this example, the computer device uses the ResNet50 convolutional neural network model to extract features from the data of multiple modalities in the screening behavior dataset, extracting the features corresponding to the data of multiple modalities in each screening behavior dataset. Then, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm is used to cluster the features corresponding to the data of multiple modalities in the screening behavior dataset, obtaining multiple clustered data subsets corresponding to each screening behavior dataset. In this example, data clustering analysis is performed on each screening behavior dataset to facilitate further redundancy removal processing.

[0035] As an example, in step S203, the computer device performs deduplication on the data in each clustered data subset to determine the multimodal behavioral data corresponding to each clustered data subset. For example, if the data in the clustered data subset is video data, then a video data is selected in each clustered data subset, and frame extraction is performed on each video data according to a predetermined sequence number to determine the image sequence corresponding to each video data. The image sequence corresponding to each video data is then determined as the multimodal behavioral data corresponding to each clustered data subset. As another example, if the data in the clustered data subset is image data, then an image is selected in each clustered data subset as the multimodal behavioral data corresponding to each clustered data subset. In this example, the data in each clustered data subset is further deduplicated to ensure that the subsequently constructed multimodal training samples do not contain redundant data.

[0036] As an example, in step S204, the computer device uses a preset model to annotate each multimodal behavior data point and determines the corresponding labeled behavior label. In this example, the computer device uses the Qwen2.5-VL multimodal large model to annotate each multimodal behavior data point and determines the corresponding labeled behavior label. For example, if a multimodal behavior data point is an image of a person falling, the Qwen2.5-VL multimodal large model is used to annotate the multimodal behavior data, and the labeled behavior label corresponding to the multimodal behavior data is determined to be "person falls".

[0037] As an example, in step S205, the computer device stores each multimodal behavior data and the corresponding labeled behavior data to obtain multiple multimodal training samples, so as to facilitate subsequent model fine-tuning and optimization training based on the multimodal training samples.

[0038] In this embodiment, the data in the initial multimodal behavior dataset is processed through data filtering, data clustering analysis, data deduplication, and data annotation to obtain multimodal training samples. This method requires no manual intervention and is relatively efficient and convenient.

[0039] In one embodiment, such as Figure 3 As shown, step S102, which involves fine-tuning the pre-trained weights in the pre-trained multimodal large model based on multimodal behavior data and labeled behavior data from multiple multimodal training samples to generate a fine-tuned behavior recognition large model, includes: S301: Identify key layers in a pre-trained multimodal large model and determine the target key layers; S302: Determine the weight matrix to be updated corresponding to the pre-trained weights based on the pre-trained weights of the target key layer; S303: Based on multiple multimodal behavioral data and labeled behavioral tags, the weight matrix to be updated is fine-tuned to generate a fine-tuned behavior recognition model.

[0040] Among them, the target key layer refers to the layer used for fine-tuning the model in the pre-trained multimodal large model.

[0041] As an example, in step S301, the computer device performs identification and analysis on each layer in the pre-trained multimodal large model, identifying the attention layer and / or feedforward layer in the pre-trained multimodal large model as the target key layer. Understandably, the attention layer and / or feedforward layer have a clear connection with other layers in the pre-trained multimodal large model. Designating the attention layer and / or feedforward layer as the target key layer facilitates lightweight global fine-tuning of the pre-trained multimodal large model when fine-tuning the target key layer.

[0042] The weight matrix to be updated refers to the matrix composed of the pre-trained weights of the target key layer.

[0043] As an example, in step S302, the computer device arranges the pre-trained weights of the target key layer in a matrix manner to determine the weight matrix to be updated corresponding to the pre-trained weights, so that the weight matrix to be updated can be fine-tuned and updated directly in the future, which is more convenient and faster.

[0044] As an example, in step S303, the computer device inputs multimodal behavior data into a pre-trained multimodal large model, outputs personnel recognition results, determines the loss function value based on the output personnel recognition results and labeled behavior, and then fine-tunes the weight matrix to be updated based on the loss function value to generate a fine-tuned behavior recognition large model. In this example, the fine-tuning of the weight matrix to be updated generates a fine-tuned behavior recognition large model without needing to fine-tune all levels in the pre-trained multimodal large model, which is more convenient and faster. For example, if the weight matrix to be updated is the matrix corresponding to the pre-trained weights of the attention mechanism level, then the weight matrix to be updated is fine-tuned to enhance the alignment ability between elements in the pre-trained multimodal large model, thereby improving the accuracy of the generated fine-tuned behavior recognition large model in terms of personnel behavior in multimodal behavior data.

[0045] In this embodiment, the weight matrix to be updated corresponding to the target key layer in the pre-trained multimodal large model is selected for fine-tuning and updating to determine the fine-tuned behavior recognition large model. This method does not require fine-tuning all layers of the pre-trained multimodal large model. While improving the recognition accuracy of the fine-tuned behavior recognition large model, it is also more convenient and faster, thus improving the fine-tuning efficiency.

[0046] In one embodiment, such as Figure 4As shown, in step S303, based on multiple multimodal behavior data and labeled behavior tags, the weight matrix to be updated is fine-tuned to generate a fine-tuned behavior recognition model, including: S401: Insert a preset low-rank matrix into the target key layer, input multimodal behavior data into the pre-trained multimodal large model, and output the initial recognition result; S402: Based on the initial recognition results and the labeled behavior tags corresponding to the multimodal behavior data, update the preset low-rank matrix to determine the first update matrix; S403: Based on the first update matrix, fine-tune the weight matrix to be updated, determine the fine-tuned weight matrix and the corresponding update behavior of the fine-tuned weight matrix to identify the large model, and determine the current number of fine-tuning steps; S404: If the current number of fine-tuning attempts has not reached the preset number of fine-tuning attempts, then update the fine-tuning weight matrix to the preset low-rank matrix, update the behavior recognition large model to the pre-trained multimodal large model, repeat the process of inserting the preset low-rank matrix into the target key layer, input the multimodal behavior data into the pre-trained multimodal large model, and output the initial recognition result. S405: If the current number of fine-tunings reaches the preset number of fine-tunings, then the updated behavior recognition model will be determined as the fine-tuned behavior recognition model.

[0047] Here, the preset low-rank matrix refers to the preset matrix used for updating. The initial recognition result refers to the result of the pre-trained multimodal large model recognizing human behavior from multimodal behavioral data.

[0048] As an example, in step S401, the computer device inserts a preset low-rank matrix into the target key layer of the pre-trained multimodal large model to facilitate subsequent updates to the preset low-rank matrix. Multimodal behavioral data is input into the pre-trained multimodal large model, and initial recognition results are output to facilitate updates to the preset low-rank matrix based on the initial recognition results.

[0049] The first row matrix refers to the matrix updated from the preset low-rank matrix.

[0050] As an example, in step S402, the computer device calculates the loss function value based on the initial recognition result and the labeled behavior tags corresponding to the multimodal behavior data, determines the update gradient based on the loss function value, and updates the preset low-rank matrix according to the update gradient to obtain the first update matrix. In this example, updating the preset low-rank matrix to obtain the first update matrix freezes all weights in the pre-trained multimodal large model, eliminating the need for complex processing of the pre-trained multimodal large model and making it more efficient.

[0051] Here, the fine-tuned weight matrix refers to the weight matrix after the weight matrix to be updated is updated. The updated behavior recognition large model refers to the pre-trained multimodal large model that includes the fine-tuned weight matrix. The current number of fine-tunings refers to the number of times the pre-trained multimodal large model has been fine-tuned.

[0052] As an example, in step S403, the computer device uses a first update matrix to fine-tune the weight matrix to be updated, determines the updated fine-tuned weight matrix, and updates the pre-trained multimodal large model according to the fine-tuned weight matrix to determine the updated behavior recognition large model. In this example, the computer device sums the first update matrix and the weight matrix to be updated to determine the fine-tuned weight matrix corresponding to the target key layer. According to the fine-tuned weight matrix, the weights corresponding to the target key layer in the pre-trained multimodal large model are updated (that is, the weights corresponding to the target key layer in the pre-trained multimodal large model are determined as the weights in the fine-tuned weight matrix), thus obtaining the updated behavior recognition large model. This method achieves lightweight and efficient fine-tuning of the pre-trained multimodal large model in the person recognition task by fine-tuning the target key layer in the pre-trained multimodal large model. In this example, after determining an updated behavior recognition large model, the computer device increments the current fine-tuning count by one to determine the current fine-tuning count, so as to stop fine-tuning in a timely manner.

[0053] As an example, in step S404, when the computer device determines that the current number of fine-tuning attempts has not reached the preset number of fine-tuning attempts, it determines that the updated behavior recognition model needs to be further fine-tuned. The fine-tuning weight matrix is ​​updated to a preset low-rank matrix, and the updated behavior recognition model is updated to a pre-trained multimodal model. Steps S401 to S403 are repeated until the current number of fine-tuning attempts reaches the preset number of fine-tuning attempts. Understandably, if the current number of fine-tuning attempts has not reached the preset number of fine-tuning attempts, it indicates that the person recognition accuracy of the updated behavior recognition model generated in the most recent fine-tuning is insufficient, and further fine-tuning is needed to improve the model's accuracy.

[0054] As an example, in step S405, when the computer device determines that the current number of fine-tuning attempts has reached the preset number of fine-tuning attempts, it stops fine-tuning the updated behavior recognition model and designates the updated behavior recognition model generated by the most recent fine-tuning attempt as the fine-tuned behavior recognition model. Understandably, if the current number of fine-tuning attempts reaches the preset number of fine-tuning attempts, it indicates that the personnel recognition accuracy of the updated behavior recognition model generated by the most recent fine-tuning attempt is high. Therefore, fine-tuning is stopped, and the updated behavior recognition model generated by the most recent fine-tuning attempt is designated as the fine-tuned behavior recognition model to ensure that the fine-tuned behavior recognition model has high accuracy in personnel behavior recognition tasks.

[0055] In this embodiment, a preset low-rank matrix is ​​inserted into the target key layer. Only this preset low-rank matrix is ​​fine-tuned and updated to determine the first update matrix. This first update matrix is ​​then used to fine-tune the weight matrix to be updated, without needing to fine-tune all layers of the model. This method is more efficient and convenient. When the current number of fine-tuning iterations reaches the preset number of iterations, a fine-tuned behavior recognition model with high accuracy for personnel behavior recognition tasks is obtained. This method does not require complex processing of large amounts of data, making it convenient and efficient.

[0056] In one embodiment, such as Figure 5 As shown, in step S103, based on the multimodal behavior data and labeled behavior tags from multiple multimodal training samples, the fine-tuned behavior recognition model is optimized and trained to generate an optimized behavior recognition model, including: S501: Input multimodal behavior data into the fine-tuned behavior recognition model and output the reference log probability corresponding to the fine-tuned behavior recognition model; S502: Input multimodal behavior data into the current policy model, and output K current behavior recognition results corresponding to the multimodal behavior data and the current log probability corresponding to each current behavior recognition result, where K>1; S503: Based on K current behavior recognition results and labeled behavior tags, determine K dominant functions; S504: Determine K KL divergences based on the reference log probability and K current log probabilities; S505: Determine the target loss function value based on the reference log probability, K dominance functions, K current log probabilities, and K KL divergences; S506: If the target loss function value does not meet the preset convergence condition, repeat the process of inputting multimodal behavior data into the fine-tuned behavior recognition model. S507: If the target loss function value satisfies the preset convergence condition, then the current strategy model is determined as the large model for optimizing behavior recognition; The current strategy model is the model after each optimization during the training process of the large model for fine-tuning behavior recognition.

[0057] The reference log probability refers to the log probability of a person's behavior output by the fine-tuned behavior recognition model after recognizing multimodal behavior data.

[0058] As an example, in step S501, the computer device inputs multimodal behavior data into the fine-tuned behavior recognition model, and the fine-tuned behavior recognition model outputs the corresponding reference log probability. For example, a multimodal behavior data set might consist of an image of a person falling. This image is input into a fine-tuned behavior recognition model. The model identifies the image and outputs the probability that the person has fallen, which corresponds to the probability that the model determined the action to be falling. This probability is the reference logarithmic probability. .

[0059] The current log probability refers to the log probability of the current policy model outputting a current behavior recognition result after recognizing multimodal behavior data.

[0060] As an example, in step S502, the computer device inputs multimodal behavior data into the current policy model and outputs K current behavior recognition results corresponding to the multimodal behavior data and the current log probability corresponding to each current behavior recognition result. Where i represents the i-th current behavior recognition result. In this example, if the fine-tuned behavior recognition model is being optimized and trained for the first time, the current policy model is the fine-tuned behavior recognition model; if it is not the first time the fine-tuned behavior recognition model is being optimized and trained, the current policy model is the model after the fine-tuned behavior recognition model has been optimized and trained.

[0061] As an example, in step S503, the computer device determines the degree of matching between each current behavior recognition result and the labeled behavior label based on the K current behavior recognition results and the model input (multimodal behavior data) of the current policy model. ,in, Let represent the degree of matching between the i-th current behavior identification result and the labeled behavior label among K current behavior identification results. In this example, =Acc( , Among them, Acc ( , )express and The degree of matching, For the i-th current behavior recognition result, The input multimodal behavior data is labeled with action tags. The computer device rewards K current action recognition results, determining the intra-group relative reward for each current action recognition result. for: = in, Let K be the number of current behavior recognition results, excluding the i-th current behavior recognition result, and the degree of matching between any current behavior recognition result and the labeled behavior tag. When the input is multimodal behavior data, it is the average reward of all current behavior recognition results, excluding the i-th current behavior recognition result, among the K current behavior recognition results.

[0062] The computer device processes the mean and standard deviation of the relative rewards within each group corresponding to the K current behavior recognition results, and determines the mean as... The standard deviation is The computer device determines the degree of matching between the i-th current behavior recognition result and the labeled behavior. mean and standard deviation Determine the advantage function corresponding to each of the K current behavior recognition results. , for = ,in, It is a small constant. Let be the advantage function corresponding to the i-th current behavior recognition result among K current behavior recognition results.

[0063] As an example, in step S504, the computer device performs difference processing on the reference log probability and the K current log probabilities to determine the K KL divergences, i.e. ,in, Let be the KL divergence corresponding to the i-th current behavior recognition result.

[0064] The target loss function value refers to the loss function value during the optimization training process of the large-scale behavior recognition model.

[0065] As an example, in step S505, the computer device sets the reference log probability. K of the aforementioned advantage functions K current log probabilities and K KL divergences The process is performed to determine the target loss function value L. This can be understood as referring to the logarithmic probability. This is used to ensure that the model formed during the optimization training process does not deviate from the fine-tuned behavior recognition model for the task of human behavior recognition, based on the current log probability. Compared with the reference log probability The difference is indicated, in this example, by referring to the log probability. K of the aforementioned advantage functions K current log probabilities and K KL divergences Within the factors that determine the target loss function value, it is possible to more accurately judge the degree of model optimization training and ensure the model's recognition accuracy of human behavior.

[0066] Among them, the preset convergence condition refers to the preset condition used to determine whether the target loss function value has converged.

[0067] As an example, in step S506, when the computer device determines that the target loss function value does not meet the preset convergence condition, it repeats steps S501 to S505 until the target loss function value meets the preset convergence condition. In this example, the preset convergence condition includes, but is not limited to, the target loss function value being less than a certain threshold, the difference between the target loss function values ​​of two adjacent optimization training sessions being less than a certain threshold, or the number of times the target loss function value is calculated reaching a certain preset number.

[0068] As an example, in step S507, when the computer device determines that the target loss function value meets the preset convergence condition, it identifies the current policy model as the optimized behavior recognition large model. Understandably, if the target loss function value meets the preset convergence condition, it indicates that the optimized model is relatively stable and has high accuracy in the human behavior recognition task. Furthermore, since the multimodal behavior data used for optimization training comes from different scenarios, the optimized model has strong scenario generalization ability. Therefore, the optimized model is identified as the optimized behavior recognition large model after optimization training, and is subsequently deployed to different scenarios to accurately recognize human behavior from multimodal data.

[0069] In this embodiment, the reference log probability, K of the aforementioned advantage functions, K of the current log probabilities, and K of the KL divergences are considered as factors in determining the target loss function value. This allows for a more accurate assessment of the degree of model optimization training and ensures the model's accuracy in recognizing human behavior. Based on the target loss function value, the fine-tuned behavior recognition model is optimized and trained to obtain an optimized behavior recognition model capable of accurately recognizing human behavior from multimodal data.

[0070] In one embodiment, such as Figure 6 As shown, step S505, which determines the target loss function value based on the reference log probability, K dominance functions, K current log probabilities, and K KL divergences, includes: S601: Based on the reference log probability, the advantage function corresponding to each current behavior recognition result, and the current log probability, determine the first loss function values ​​corresponding to the K current behavior recognition results; S602: Correct the KL divergence corresponding to each current behavior recognition result, and determine the second loss function values ​​corresponding to K current behavior recognition results; S603: Based on K first loss function values ​​and K second loss function values, determine the within-group loss function values ​​corresponding to K current behavior recognition results; S604: Determine the target loss function value based on the within-group loss function values ​​corresponding to the K current behavior recognition results.

[0071] The first loss function value refers to the loss function value that outputs the current behavior recognition result during the model optimization and training process.

[0072] As an example, in step S601, the computer device sets the reference log probability. Advantage function corresponding to each current behavior recognition result and the current log probability corresponding to each current behavior recognition result Processing is performed to determine the first loss function values ​​corresponding to the K current behavior recognition results. In this example, the computer device identifies the current logarithmic probability for each current action recognition result. and reference log probability Perform exponential function processing to determine the current logarithmic probability. The corresponding exponential function exp( ) and reference log probability The corresponding exponential function exp( ), for exp( ) and exp( Perform interpolation processing, determine the interpolation result, and then compare this interpolation result with the advantage function corresponding to each current behavior recognition result. Perform multiplication to determine the first loss function value corresponding to each current behavior recognition result. .in, .

[0073] The second loss function value refers to the loss function value after correcting the KL divergence.

[0074] As an example, in step S602, the computer device uses a preset correction coefficient. K KL divergences Make corrections and determine the second loss function values ​​corresponding to the K current behavior recognition results. In this example, = ,in, Let be the value of the second loss function corresponding to the i-th current behavior recognition result.

[0075] The within-group loss function value refers to the loss function value corresponding to each of the K current behavior recognition results output during the optimization training process, where the input is the same multimodal behavior data.

[0076] As an example, in step S603, the computer device processes the K first loss function values ​​and K second loss function values ​​to determine the within-group loss function values ​​corresponding to the K current behavior recognition results. In this example, for the same multimodal behavior data, the i-th current behavior recognition result among the K current behavior recognition results is output, and its corresponding within-group loss function value is... ,in, = In this example, the first and second loss function values ​​are taken into account to ensure that during the training and optimization process, the current policy model after optimization maintains the ability of the large behavior recognition model to recognize people's behavior while improving the accuracy of the model in recognizing people's behavior.

[0077] As an example, in step S604, the computer device determines the target loss function value L as the mean of the within-group loss function values ​​corresponding to the K current behavior recognition results. Where L = In this example, the target loss function value is determined based on the within-group loss function values ​​corresponding to the K current behavior recognition results for the same input, which can effectively improve the accuracy of model optimization training.

[0078] In this embodiment, the target loss function value is determined based on the dominance function and KL divergence, thereby optimizing the large-scale behavior recognition model through fine-tuning. This stabilizes the training process, enhances the model's responsiveness to reward signals, and avoids drastic deviations in the model's performance for personnel behavior recognition tasks. By combining the advantages of the dominance function and KL divergence, the accuracy of personnel behavior recognition tasks is improved while simultaneously enhancing the stability of optimized training and the model's generalization ability.

[0079] In one embodiment, such as Figure 7 As shown, step S104, which involves quantizing the optimized behavior recognition model to generate the target behavior recognition model, includes: S701: Obtain the optimization weight matrix corresponding to the large-scale optimization behavior recognition model, divide the optimization weight matrix into blocks, and determine multiple optimization weight sub-matrices; S702: Use a preset quantization parameter algorithm to process each optimized weight submatrix and determine the scaling factor and zero point corresponding to each optimized weight submatrix; S703: Based on the scaling factor and zero point, the optimized weight submatrix is ​​quantized to determine the quantized weight submatrix corresponding to the optimized weight submatrix. Based on multiple quantized weight submatrixes, the quantized weight matrix is ​​determined. S704: Based on the quantized weight matrix, the optimized behavior recognition model is quantized and updated to determine the target behavior recognition model.

[0080] The optimized weight matrix is ​​a matrix that contains all the weights in the optimized behavior recognition model. The optimized weight submatrix is ​​a matrix obtained by dividing the optimized weight matrix into blocks.

[0081] As an example, in step S701, the computer device acquires all the weights in the optimized behavior recognition large model, arranges all the weights according to the hierarchical relationship, determines the optimized weight matrix, and divides the optimized weight matrix into blocks according to a preset number of rows and columns to determine multiple optimized weight sub-matrices. In this example, the optimized weight matrix is ​​divided into blocks to determine multiple optimized weight sub-matrices, so as to facilitate subsequent block-by-block quantization of the optimized behavior recognition large model and reduce global error propagation.

[0082] Among them, the preset quantization parameter algorithm refers to the preset algorithm used to determine the scaling factor and zero point corresponding to the optimized weight submatrix.

[0083] As an example, in step S702, the computer device processes each optimized weight submatrix using a preset quantization parameter algorithm to determine the scaling factor and zeros corresponding to each optimized weight submatrix. In this example, the computer device uses the GPTQ-int8 algorithm (an 8-bit version of Gradient Post-Training Quantization) to calculate the scaling factor and zeros for each optimized weight submatrix, thus determining the scaling factor and zeros corresponding to each optimized weight submatrix.

[0084] Here, the quantized weight submatrix refers to the submatrix obtained by quantizing the optimized weight submatrix. The quantized weight matrix refers to the weights composed of the quantized weight submatrixes.

[0085] As an example, in step S703, the computer device uses the scaling factor and zero point corresponding to each optimized weight submatrix to perform quantization processing on each optimized weight submatrix, thereby determining the quantized weight submatrix corresponding to each optimized weight submatrix. In this example, the computer device uses the formula... , , For each optimized weight submatrix, quantization is performed to determine the quantized weight submatrix corresponding to each optimized weight submatrix. ,in, The quantized weights contained in the quantized weight submatrix. To optimize the unquantized weights in the weight submatrix, To optimize the zeros corresponding to the weight submatrix, To optimize the scaling factor corresponding to the weight submatrix, the computer device reassembles the quantized weight submatrix into a quantized weight matrix according to the block order of optimizing the weight matrix in step S701, so as to facilitate subsequent quantization updates of the optimized behavior recognition large model based on the quantized weight matrix.

[0086] As an example, in step S704, the computer device uses the quantized weights from the quantized weight matrix to update the corresponding unquantized weights in the optimized behavior recognition model, thus obtaining the target behavior recognition model containing the quantized weights. Understandably, although the optimized behavior recognition model has strong scene generalization ability and relatively accurate personnel behavior recognition ability, its large parameter scale leads to problems such as high resource consumption, high response latency, and significant power consumption in actual deployment, severely limiting the model's practical application. In this example, by quantizing and updating the optimized behavior recognition model, the parameter scale is reduced while ensuring scene generalization ability and personnel behavior recognition ability, effectively reducing the model's storage and deployment costs.

[0087] In this embodiment, by dividing the optimized weight matrix into blocks, multiple optimized weight sub-matrices are determined. The optimized weight sub-matrices are then quantized in blocks to achieve quantization and update of the optimized weight matrix corresponding to the large-scale behavior recognition model. This method can effectively reduce error propagation, avoid the accumulation of errors between model layers, and enable the large-scale target behavior recognition model to have better performance.

[0088] In another embodiment, if Figure 8 As shown, a behavior recognition method is provided, and this behavior recognition large model generation method is applied to... Figure 10 Taking a computer device as an example, the explanation includes the following steps: S801: Acquire the behavior data to be identified; S802: Input the behavior data to be identified into the target behavior recognition big model generated by the behavior recognition big model generation method, perform behavior recognition, and determine the target behavior corresponding to the behavior data to be identified.

[0089] Among them, the behavioral data to be identified refers to the data that needs to be used for human behavior identification, including but not limited to images, videos, audio, and text.

[0090] As an example, in step S801, the computer device acquires behavioral data to be identified in various scenarios such as smart home, public safety, traffic management, medical care, and human-computer interaction.

[0091] Among them, target behavior refers to the human behavior identified by the target behavior recognition model.

[0092] As an example, in step S802, the computer device inputs the acquired behavior data to be identified into the target behavior recognition model to perform human behavior recognition and outputs the target behavior corresponding to the behavior data to be identified. For example, the computer device takes an image of a person falling down captured in a smart home scenario as the behavior data to be identified, inputs it into the target behavior recognition model to perform human behavior recognition, and outputs that the target behavior corresponding to the behavior data to be identified is a person falling down.

[0093] In this embodiment, a large target behavior recognition model that has been fine-tuned, optimized, trained, and quantized is used. This model can accurately identify human behavior in multimodal data in various scenarios and has strong scenario generalization ability and stability.

[0094] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0095] In one embodiment, a behavior recognition large model generation apparatus is provided, which corresponds one-to-one with the behavior recognition large model generation method described in the above embodiments. For example... Figure 9 As shown, the behavior recognition large model generation device includes a multimodal training sample acquisition module 901, a fine-tuning behavior recognition large model generation module 902, an optimized behavior recognition large model generation module 903, and a target behavior recognition large model generation module 904. Detailed descriptions of each functional module are as follows: The multimodal training sample acquisition module 901 is used to acquire multiple multimodal training samples, each of which includes multimodal behavioral data and corresponding labeled behavioral tags. The fine-tuning behavior recognition large model generation module 902 fine-tunes the pre-training weights in the pre-trained multimodal large model based on multimodal behavior data and labeled behavior tags from multiple multimodal training samples, thereby generating a fine-tuned behavior recognition large model. The optimized behavior recognition large model generation module 903 optimizes the fine-tuned behavior recognition large model based on multimodal behavior data and labeled behavior tags from multiple multimodal training samples, and generates an optimized behavior recognition large model. The target behavior recognition large model generation module 904 is used to quantize the optimized behavior recognition large model to generate the target behavior recognition large model.

[0096] In one embodiment, the multimodal training sample acquisition module 901 includes: The behavior dataset determination submodule is used to obtain the initial multimodal behavior dataset, perform data filtering processing on the initial multimodal behavior dataset according to multiple preset behaviors, and determine the filtering behavior dataset corresponding to each preset behavior. The clustering data subset determination submodule is used to perform data clustering analysis on each filtering behavior dataset to determine multiple clustering data subsets corresponding to each filtering behavior dataset; The multimodal behavior data determination submodule is used to perform data deduplication on each clustered data subset and determine the multimodal behavior data corresponding to each clustered data subset. The behavior label determination submodule is used to label multimodal behavior data and determine the corresponding behavior label for each multimodal behavior data. The multimodal training sample determination submodule determines multiple multimodal training samples based on each multimodal behavioral data and the corresponding labeled behavioral data.

[0097] In one embodiment, the fine-tuning behavior recognition large model generation module 902 includes: The target key layer determination submodule is used to identify key layers in a pre-trained multimodal large model and determine the target key layers. The submodule for determining the weight matrix to be updated is used to determine the weight matrix to be updated corresponding to the pre-trained weights based on the pre-trained weights of the target key layer. The fine-tuning behavior recognition large model generation submodule, based on multiple multimodal behavior data and labeled behavior tags, performs fine-tuning on the weight matrix to be updated, generating a fine-tuned behavior recognition large model.

[0098] In one embodiment, the fine-tuning behavior recognition large model generation submodule includes: The initial recognition result output unit is used to insert a preset low-rank matrix into the target key layer, input multimodal behavior data into the pre-trained multimodal large model, and output the initial recognition result; The first update matrix determination unit updates the preset low-rank matrix based on the initial recognition results and the labeled behavior tags corresponding to the multimodal behavior data, and determines the first update matrix. The current fine-tuning number determination unit, based on the first update matrix, fine-tunes the weight matrix to be updated, determines the fine-tuned weight matrix and the corresponding update behavior recognition model, and determines the current fine-tuning number; The first judgment unit is used to update the fine-tuning weight matrix to a preset low-rank matrix, update the behavior recognition big model to a pre-trained multimodal big model, repeatedly insert the preset low-rank matrix into the target key layer, input the multimodal behavior data into the pre-trained multimodal big model, and output the initial recognition result if the current fine-tuning number has not reached the preset fine-tuning number. The second judgment unit is used to determine the updated behavior recognition model as the fine-tuned behavior recognition model if the current number of fine-tunings reaches the preset number of fine-tunings.

[0099] In one embodiment, the optimized behavior recognition large model generation module 903 includes: The reference log probability determination submodule is used to input multimodal behavior data into the fine-tuned behavior recognition model and output the reference log probability corresponding to the fine-tuned behavior recognition model. The current log probability determination submodule is used to input multimodal behavior data into the current policy model and output the K current behavior recognition results corresponding to the multimodal behavior data and the current log probability corresponding to each current behavior recognition result, where K>1; The dominance function determination submodule determines K dominance functions based on K current behavior recognition results and labeled behavior tags; The KL divergence determination submodule determines K KL divergences based on the reference log probability and K current log probabilities. The target loss function value determination submodule determines the target loss function value based on the reference log probability, K dominance functions, K current log probabilities, and K KL divergences. The first preset convergence condition judgment submodule is used to repeatedly input multimodal behavior data into the fine-tuned behavior recognition large model if the target loss function value does not meet the preset convergence condition. The second preset convergence condition judgment submodule is used to determine the current strategy model as the optimized behavior recognition large model if the target loss function value meets the preset convergence condition.

[0100] In one embodiment, the target loss function value determination submodule includes: The first loss function value determination unit determines the first loss function values ​​corresponding to K current behavior recognition results based on the reference log probability, the advantage function corresponding to each current behavior recognition result, and the current log probability. The second loss function value determination unit is used to correct the KL divergence corresponding to each current behavior recognition result and determine the second loss function value corresponding to K current behavior recognition results. The intra-group loss function value determination unit determines the intra-group loss function values ​​corresponding to the current behavior recognition results based on the K first loss function values ​​and the K second loss function values; The target loss function value determination unit determines the target loss function value based on the intra-group loss function values ​​corresponding to the K current behavior recognition results.

[0101] In one embodiment, the target behavior recognition large model generation module 904 includes: The optimization weight submatrix determination submodule is used to obtain the optimization weight matrix corresponding to the large model of behavior recognition, and to divide the optimization weight matrix into blocks to determine multiple optimization weight submatrices. The scaling factor and zero point determination submodule is used to process each optimized weight submatrix using a preset quantization parameter algorithm to determine the scaling factor and zero point corresponding to each optimized weight submatrix. The quantization weight matrix determination submodule performs quantization processing on the optimized weight submatrix based on the scaling factor and zero point, determines the quantized weight submatrix corresponding to the optimized weight submatrix, and determines the quantization weight matrix based on multiple quantized weight submatrixes. The target behavior recognition large model determination submodule, based on the quantized weight matrix, performs quantization updates on the optimized behavior recognition large model to determine the target behavior recognition large model.

[0102] In one embodiment, a behavior recognition device is provided, which corresponds one-to-one with the behavior recognition methods described in the above embodiments. The behavior recognition device includes a behavior data acquisition module and a target behavior determination module. Detailed descriptions of each functional module are as follows: The module for acquiring behavior data to be identified is used to acquire behavior data to be identified. The target behavior determination module is used to input the behavior data to be identified into the target behavior recognition big model generated by the behavior recognition big model generation method, perform behavior recognition, and determine the target behavior corresponding to the behavior data to be identified.

[0103] Specific limitations regarding the behavior recognition large-scale model generation device can be found in the limitations regarding the behavior recognition large-scale model generation method above, and specific limitations regarding the behavior recognition device can be found in the limitations regarding the behavior recognition method above, and will not be repeated here. The aforementioned behavior recognition large-scale model generation device and each module within the behavior recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0104] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data used and generated during the execution of the behavior recognition large model generation method and the behavior recognition method. The network interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a behavior recognition large model generation method, or, when executed by the processor, implements a behavior recognition method.

[0105] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the behavior recognition large model generation method described in the above embodiments, for example... Figure 1 As shown in S101-S104, or Figures 2 to 7 As shown, to avoid repetition, it will not be described again here. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the behavior recognition large model generation device, for example... Figure 9 The functions of the multimodal training sample acquisition module 901, the fine-tuning behavior recognition large model generation module 902, the optimized behavior recognition large model generation module 903, and the target behavior recognition large model generation module 904 shown are not described in detail here to avoid repetition.

[0106] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the behavior recognition method described in the above embodiments, for example... Figure 8 S801-S802, as shown, will not be described again here to avoid repetition. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the behavior recognition device, such as the functions of the module for acquiring data on behavior to be recognized and the module for determining target behavior. To avoid repetition, these will not be described again here.

[0107] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the behavior recognition large model generation method described above, for example... Figure 1 As shown in S101-S104, or Figures 2 to 7As shown, to avoid repetition, it will not be described again here. Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in this embodiment of the behavior recognition large model generation device, for example... Figure 9 The functions of the multimodal training sample acquisition module 901, the fine-tuning behavior recognition large model generation module 902, the optimized behavior recognition large model generation module 903, and the target behavior recognition large model generation module 904 shown are not described in detail here to avoid repetition.

[0108] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the behavior recognition method described in the above embodiment, for example... Figure 8 S801-S802, as shown, will not be described again here to avoid repetition. Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in this embodiment of the behavior recognition device, such as the functions of the module for acquiring data on behavior to be recognized and the module for determining target behavior. To avoid repetition, these will not be described again here.

[0109] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0110] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0111] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for generating a large model for behavior recognition, characterized in that, include: Multiple multimodal training samples are obtained, each of which includes multimodal behavioral data and labeled behavioral tags corresponding to the multimodal behavioral data; Based on the multimodal behavior data and labeled behavior tags in multiple multimodal training samples, the pre-training weights in the pre-trained multimodal large model are fine-tuned to generate a fine-tuned behavior recognition large model. Based on the multimodal behavior data and labeled behavior tags in multiple multimodal training samples, the fine-tuned behavior recognition model is optimized and trained to generate an optimized behavior recognition model. The optimized behavior recognition model is quantized to generate a target behavior recognition model.

2. The method for generating a large model for behavior recognition according to claim 1, characterized in that, The acquisition of multiple multimodal training samples includes: Obtain an initial multimodal behavior dataset, and perform data filtering processing on the initial multimodal behavior dataset according to multiple preset behaviors to determine the filtered behavior dataset corresponding to each preset behavior; Perform data clustering analysis on each of the aforementioned filtering behavior datasets to determine multiple clustered data subsets corresponding to each of the aforementioned filtering behavior datasets; For each clustered data subset, perform data deduplication to determine the multimodal behavioral data corresponding to each clustered data subset; The multimodal behavior data is labeled to determine the labeled behavior label corresponding to each multimodal behavior data. Based on each of the multimodal behavioral data and the corresponding labeled behavioral tags, multiple multimodal training samples are determined.

3. The method for generating a large model for behavior recognition according to claim 1, characterized in that, The step of fine-tuning the pre-training weights in the pre-trained multimodal large model based on the multimodal behavior data and the labeled behavior tags from multiple multimodal training samples to generate a fine-tuned behavior recognition large model includes: Key layer identification is performed on the pre-trained multimodal large model to determine the target key layer; Based on the pre-trained weights of the target key layer, determine the weight matrix to be updated corresponding to the pre-trained weights; Based on the multiple multimodal behavior data and the labeled behavior tags, the weight matrix to be updated is fine-tuned to generate a fine-tuned behavior recognition model.

4. The method for generating a large model for behavior recognition according to claim 3, characterized in that, The step of fine-tuning the weight matrix to be updated based on multiple multimodal behavior data and the labeled behavior tags to generate a fine-tuned behavior recognition large model includes: A preset low-rank matrix is ​​inserted into the target key layer, the multimodal behavior data is input into the pre-trained multimodal large model, and the initial recognition result is output. Based on the initial recognition results and the labeled behavior tags corresponding to the multimodal behavior data, the preset low-rank matrix is ​​updated to determine the first update matrix; Based on the first update matrix, the weight matrix to be updated is fine-tuned to determine the fine-tuned weight matrix and the corresponding update behavior recognition model, and the current number of fine-tunings is determined. If the current number of fine-tunings has not reached the preset number of fine-tunings, then the fine-tuning weight matrix is ​​updated to the preset low-rank matrix, the updated behavior recognition big model is updated to the pre-trained multimodal big model, the process of inserting the preset low-rank matrix in the target key layer is repeated, the multimodal behavior data is input into the pre-trained multimodal big model, and the initial recognition result is output. If the current number of fine-tunings reaches the preset number of fine-tunings, then the updated behavior recognition model is determined as the fine-tuned behavior recognition model.

5. The method for generating a large model for behavior recognition according to claim 1, characterized in that, The step of optimizing and training the fine-tuned behavior recognition model based on the multimodal behavior data and the labeled behavior tags from multiple multimodal training samples to generate an optimized behavior recognition model includes: The multimodal behavior data is input into the fine-tuned behavior recognition model, and the reference log probability corresponding to the fine-tuned behavior recognition model is output. The multimodal behavior data is input into the current policy model, and the K current behavior recognition results corresponding to the multimodal behavior data and the current log probability corresponding to each current behavior recognition result are output, where K > 1; Based on the K current behavior recognition results and the labeled behavior tags, K dominant functions are determined; Based on the reference log probability and K current log probabilities, determine K KL divergences; The target loss function value is determined based on the reference log probability, K of the dominance functions, K of the current log probabilities, and K of the KL divergences. If the target loss function value does not meet the preset convergence condition, then the process of inputting the multimodal behavior data into the fine-tuned behavior recognition model is repeated. If the target loss function value satisfies the preset convergence condition, then the current strategy model is determined as the large-scale optimization behavior recognition model; The current strategy model is the model after each optimization during the optimization training process of the fine-tuned behavior recognition large model.

6. The method for generating a large model for behavior recognition according to claim 5, characterized in that, The step of determining the target loss function value based on the reference log probability, K of the dominance functions, K of the current log probabilities, and K of the KL divergences includes: Based on the reference log probability, the advantage function corresponding to each current behavior recognition result, and the current log probability, determine K first loss function values ​​corresponding to the current behavior recognition results; The KL divergence corresponding to each current behavior recognition result is corrected to determine K second loss function values ​​corresponding to the current behavior recognition results; Based on K first loss function values ​​and K second loss function values, determine K in-group loss function values ​​corresponding to the current behavior recognition results; The target loss function value is determined based on the K intra-group loss function values ​​corresponding to the current behavior recognition results.

7. A behavior recognition method, characterized in that, include: Acquire the behavior data to be identified; The behavior data to be identified is input into the target behavior recognition large model generated by any one of the behavior recognition large model generation methods of claims 1 to 6, and behavior recognition is performed to determine the target behavior corresponding to the behavior data to be identified.

8. A device for generating large-scale behavior recognition models, characterized in that, include: A multimodal training sample acquisition module is used to acquire multiple multimodal training samples, each of which includes multimodal behavior data and labeled behavior tags corresponding to the multimodal behavior data; The fine-tuning behavior recognition large model generation module fine-tunes the pre-training weights in the pre-trained multimodal large model based on the multimodal behavior data and the labeled behavior tags in multiple multimodal training samples, thereby generating a fine-tuned behavior recognition large model. The optimized behavior recognition large model generation module optimizes and trains the fine-tuned behavior recognition large model based on the multimodal behavior data and the labeled behavior tags in multiple multimodal training samples, thereby generating an optimized behavior recognition large model. The target behavior recognition large model generation module is used to quantize the optimized behavior recognition large model to generate a target behavior recognition large model.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the behavior recognition large model generation method according to any one of claims 1 to 6, or when the processor executes the computer program, it implements the behavior recognition method according to claim 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the behavior recognition large model generation method according to any one of claims 1 to 6, or, when the computer program is executed by the processor, it implements the behavior recognition method according to claim 7.