A Real-time Event Stream Recognition Method and System
By introducing model-independent meta-learning strategies into the human event detection model, integrating video and sensing data features and performing individual adaptive meta-learning, the problem of insufficient generalization ability of the model is solved, and individual adaptive event flow detection is realized, which improves the accuracy and universality of the detection.
Patent Information
- Application Number
- CN202111069905.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-09-13
AI Technical Summary
In the prior art, the human event detection model has poor generalization ability in data centers with uneven data, and cannot personalize the individual, resulting in bias in the detection results.
The event stream detection model is optimized by model-independent meta-learning strategy. By fusing video data and sensing data characteristics, the boundary matching network and classification recognition network are used to perform individual adaptive meta-learning, and an individual adaptive event stream detection model is constructed.
The generalization ability of the model is improved, making it more sensitive and adaptable in the event detection tasks of new users, reducing changes to model parameters, and improving the accuracy and universality of detection.
Smart Images

Figure CN113901880B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimedia technology, and in particular to a real-time event stream recognition method and system. Background Art
[0002] In the research on human event detection, an event stream detection model is often used, and this model often has the problem of poor model generalization ability. For example, in some data-imbalanced data sets, due to large data distribution deviations, the trained model often prefers certain specific event categories.
[0003] The task of real-time health event recognition generally includes two steps: time series detection and multi-modal event recognition. However, since the existing human event detection solutions treat all research objects equally when training the model without distinguishing the research objects, this will cause the problem of poor model generalization ability due to factors such as data distribution imbalance, and cannot meet the need for individual adaptation. Summary of the Invention
[0004] The present invention provides a real-time event stream recognition method and system to solve the defect in the prior art that it is impossible to make individual distinctions for the individuals to be detected for event recognition, resulting in easy deviation of the detection results.
[0005] In a first aspect, the present invention provides a real-time event stream recognition method, including:
[0006] Determine the individual event to be recognized;
[0007] Input the individual event to be recognized into a pre-trained event stream detection model to obtain an individual event recognition result; wherein, the event stream detection model is obtained by fusing different individual event sample data features to obtain a fusion feature, then performing boundary matching on the fusion feature, and performing classification and recognition based on individual adaptive meta-learning.
[0008] In one embodiment, the event stream detection model is obtained through the following steps:
[0009] Obtain the video data feature and the sensing data feature;
[0010] Connect and fuse the video data feature and the sensing data feature to obtain the fusion feature;
[0011] Input the fusion feature into a boundary matching network to obtain candidate time segment data;
[0012] Input the candidate time segment data into a classification and recognition network to obtain the individual event type and the individual event time information;
[0013] Determine the meta-learning tasks for different user individuals, optimize the fused features, the candidate time segment data, the individual event types, and the individual event time information to obtain the event stream detection model.
[0014] In one embodiment, the obtaining the video data features and the sensing data features includes:
[0015] Input a video segment of any length into a preset video feature extractor to obtain the video data features;
[0016] Input the sensor data corresponding to the video segment of any length into a preset sensing feature extractor to obtain the sensing data features.
[0017] In one embodiment, the concatenating and fusing the video data features and the sensing data features to obtain the fused features includes:
[0018] Concatenate the video data features and the sensing data features with the same feature dimension to obtain a concatenated feature;
[0019] Perform feature fusion on the concatenated feature by a linear network to obtain the fused features with a unified representation of multi-source heterogeneous data.
[0020] In one embodiment, the inputting the fused features into a boundary matching network to obtain candidate time segment data includes:
[0021] Determine the temporal feature sequence of the fused features;
[0022] Based on the temporal candidate initial boundary positions and the temporal candidate lengths of the temporal feature sequence, construct the temporal candidates into a boundary matching feature map, and obtain a boundary matching confidence map based on the boundary matching feature map;
[0023] Extract the feature representation of the fused features according to the confidence scores of the boundary matching confidence map to obtain the candidate time segment data.
[0024] In one embodiment, the based on the temporal candidate initial boundary positions and the temporal candidate lengths of the temporal feature sequence, constructing the temporal candidates into a boundary matching feature map, and obtaining a boundary matching confidence map based on the boundary matching feature map includes:
[0025] Obtain any temporal candidate segment on a boundary matching map with a preset size;
[0026] Obtain several first sampling points in the temporal feature sequence, and determine the several first sampling points as the first temporal feature sequence of the any temporal candidate segment;
[0027] Obtain the start time and end time of any of the time series candidate segments, determine a preset extended time series range based on the start time and the end time, and obtain a number of second sampling points within the preset extended time series range that is the same as the number of the first sampling points;
[0028] Determine the value of each sampling point among the number of second sampling points to obtain a second time series feature sequence;
[0029] Perform dot multiplication and linear interpolation on the first time series feature sequence and the second time series feature sequence to obtain a feature of any time series candidate segment;
[0030] Use a three-dimensional convolutional network to eliminate the number of sampling points in the feature of any time series candidate segment, and determine the context information of any time series candidate segment through multiple two-dimensional convolutional layers to obtain the boundary matching confidence map.
[0031] In one embodiment, the inputting the candidate time segment data into a classification and recognition network to obtain the individual event type and the individual event time information includes:
[0032] Input the data stream feature of each candidate time segment into the classification and recognition network, classify the corresponding data time period of each candidate time segment based on the boundary matching confidence map, and obtain the individual event category, the individual event start time, and the individual event end time in each candidate time segment.
[0033] In one embodiment, the determining the meta-learning tasks of different user individuals, optimizing the fusion feature, the candidate time segment data, the individual event type, and the individual event time information to obtain the event stream detection model includes:
[0034] Determine an initial meta-learning model, where the initial meta-learning model includes a first parameter corresponding to the fusion feature, a second parameter corresponding to the candidate time segment data, and a third parameter corresponding to the individual event type and the individual event time information;
[0035] Determine a model update step size hyperparameter, and obtain a meta-training task loss function composed of any user data in the initial meta-learning model;
[0036] Based on the first parameter, the second parameter, the third parameter, the model update step size hyperparameter, and the meta-training task loss function, determine a single gradient descent update to obtain an updated meta-learning model;
[0037] Adapt and update the updated meta-learning model with a number of different user data to obtain the event stream detection model.
[0038] In a second aspect, the present invention further provides a real-time event stream recognition system, including:
[0039] A determination module, configured to determine an individual event to be recognized;
[0040] A detection module, configured to input the individual event to be recognized into a pre-trained event stream detection model to obtain an individual event recognition result; wherein, the event stream detection model is obtained by fusing the data features of different individual event samples to obtain a fusion feature, then performing boundary matching on the fusion feature, and performing classification and recognition based on individual adaptive meta-learning.
[0041] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any one of the above-mentioned real-time event stream recognition methods are implemented.
[0042] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned real-time event stream recognition methods are implemented.
[0043] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above-mentioned real-time event stream recognition methods are implemented.
[0044] The real-time event stream recognition method and system provided by the present invention, when recognizing a health real-time event, optimize the event stream detection model by adopting a model-agnostic meta-learning strategy, realize obtaining a detection model with individual adaptability at the lowest cost, and greatly improve the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0046] Figure 1 is a flowchart of the real-time event stream recognition method provided by the present invention;
[0047] Figure 2 is a logical diagram of the real-time event stream recognition method provided by the present invention;
[0048] Figure 3 is a structural diagram of the real-time event stream recognition system provided by the present invention;
[0049] Figure 4 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments
[0050] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0051] Currently, in the research of human event detection, there are often problems with poor model generalization ability. For example, in some datasets with unbalanced data, due to large data distribution deviations, the trained models often tend to prefer certain specific event categories. The real-time health event recognition task generally includes two steps: time series detection and multi-modal event recognition. Similar to traditional methods, the present invention proposes a real-time event stream recognition method.
[0052] Figure 1 It is a schematic flow diagram of the real-time event stream recognition method provided by the present invention, as Figure 1 shown, including:
[0053] S1. Determine the individual event to be recognized;
[0054] S2. Input the individual event to be recognized into a pre-trained event stream detection model to obtain an individual event recognition result; wherein, the event stream detection model is obtained by fusing the data features of different individual event samples to obtain a fusion feature, then performing boundary matching on the fusion feature, and performing classification recognition based on individual adaptive meta-learning.
[0055] Specifically, aiming at the problems existing in the prior art, the present invention adopts a model-agnostic meta-learning strategy to optimize the proposed detection model, so that it can be optimized at the lowest cost to obtain an individual-adaptive model when dealing with new research objects. The proposed detection method mainly updates the parameters of the model by learning, so that it has good sensitivity in the event detection task of new users, and a small change in the model parameters can make it achieve good performance in any new user event detection task. This method has no restrictions on the form of the model, only requires that the model can be optimized and trained by the way of gradient backpropagation, so it can be well applied to multi-modal event detection including modules such as time series detection and event recognition.
[0056] Specifically, without changing the model structure, multiple subtasks are constructed, and the model is optimized on each subtask to update the parameters of each sub-module in the model. Finally, based on the constructed multiple subtasks, the original model parameters are updated to make it most conducive to adapting to all subtasks. In this way, the final event detection model with individual adaptability characteristics can greatly improve the generalization ability of the model.
[0057] Logically, the adaptive real-time health event stream recognition system proposed by the present invention is divided into two modules. The first is a multi-modal human event detection module based on feature fusion. First, the video data and sensor data are respectively subjected to feature extraction to obtain feature representations, and then the data stream after fusing the two feature representations is sampled with a fixed time window. Then, a deep neural network is used to obtain the start and end times of the data stream event, and then event detection is performed according to the candidate time segments. The second is an individual adaptive meta-learning module. First, the model parameters obtained from the first module are used as initialization parameters, and then meta-learning tasks are constructed according to the training samples of different users. The overall logical structure is as Figure 2 shown.
[0058] When the present invention recognizes real-time health events, it uses a model-agnostic meta-learning strategy to optimize the event stream detection model, realizing the detection model with individual adaptability at the lowest cost and greatly improving the generalization ability of the model.
[0059] Based on any of the above embodiments, the event stream detection model is obtained through the following steps:
[0060] Obtain the video data features and the sensor data features;
[0061] Connect and fuse the video data features and the sensor data features to obtain the fused features;
[0062] Input the fused features into the boundary matching network to obtain candidate time segment data;
[0063] Input the candidate time segment data into the classification and recognition network to obtain individual event types and individual event time information;
[0064] Determine the meta-learning tasks of different user individuals, and optimize the fused features, the candidate time segment data, the individual event types, and the individual event time information to obtain the event stream detection model.
[0065] Specifically, the event stream detection model proposed by the present invention specifically includes:
[0066] Video data and sensor data of any length are respectively passed through a feature extractor to obtain underlying feature representations with the same feature dimension;
[0067] Fuse the obtained video data features and sensor data features;
[0068] Input the fused features represented by the two output features into the BMN (Boundary-matching network) to obtain candidate time segment data;
[0069] Input the candidate time segment data generated by the boundary matching network into the classification and recognition network to obtain the final individual event type and the start and end times of the event;
[0070] Use the obtained basic model parameters as initialization parameters, construct a meta-learning task according to the training samples of different users, and update the corresponding model parameters in a targeted manner;
[0071] Finally, an event stream detection model with individual adaptive characteristics is obtained according to the model training and optimization in the above steps.
[0072] Through constructing, training and optimizing the event stream detection model, the present invention can identify and detect different individual events, and has strong generalization and universality.
[0073] Based on any of the above embodiments, the obtaining of the video data features and the sensor data features includes:
[0074] Input a video segment of any length into a preset video feature extractor to obtain the video data features;
[0075] Input the sensor data corresponding to the video segment of any length into a preset sensor feature extractor to obtain the sensor data features.
[0076] Specifically, the C3D neural network is used as the video feature extractor to obtain visual features:
[0077]
[0078] Wherein, l v is the length of the video, σ is the number of video frames input into the feature extractor each time, and t n is the nth (1 ≤ n ≤ l f ) sub-video segment in a video, where C is the feature dimension is the feature of the nth sub-video segment. At the same time, except for the different feature extractors selected, the sensor data is processed in the same way.
[0079] When the present invention processes data features corresponding to different individual events, the same feature extractor is adopted, which is convenient for unifying the data dimension and facilitating subsequent data fusion processing.
[0080] Based on any of the above embodiments, the connecting and fusing the video data features and the sensing data features to obtain the fused features includes:
[0081] Connecting the video data features and the sensing data features with the same feature dimension to obtain a connection feature;
[0082] Performing feature fusion on the connection feature by a linear network to obtain the fused feature with a unified representation of multi-source heterogeneous data.
[0083] Specifically, the video features and the sensing features with the same feature dimension are connected, and then feature fusion is performed through a linear network to obtain a unified representation of multi-source heterogeneous data.
[0084] Based on any of the above embodiments, the inputting the fused feature into a boundary matching network to obtain candidate time segment data includes:
[0085] Determining the timing feature sequence of the fused feature;
[0086] Based on the initial boundary position of the timing candidate and the timing candidate length of the timing feature sequence, constructing the timing candidate into a boundary matching feature map, and obtaining a boundary matching confidence map based on the boundary matching feature map;
[0087] Extracting the feature representation of the fused feature according to the confidence score of the boundary matching confidence map to obtain the candidate time segment data.
[0088] Wherein, the constructing the timing candidate into a boundary matching feature map based on the initial boundary position of the timing candidate and the timing candidate length of the timing feature sequence, and obtaining a boundary matching confidence map based on the boundary matching feature map includes:
[0089] Obtaining any timing candidate segment on a boundary matching map with a preset size;
[0090] Obtaining a plurality of first sampling points in the timing feature sequence, and determining the plurality of first sampling points as the first timing feature sequence of the any timing candidate segment;
[0091] Obtaining the start time and the end time of the any timing candidate segment, determining a preset extended timing range based on the start time and the end time, and obtaining a plurality of second sampling points with the same number as the plurality of first sampling points within the preset extended timing range;
[0092] Determine the value of each sampling point among the several second sampling points to obtain a second time series feature sequence;
[0093] Perform dot product and linear interpolation on the first time series feature sequence and the second time series feature sequence to obtain the feature of any time series candidate segment;
[0094] Use a three-dimensional convolutional network to eliminate the number of sampling points in the feature of any time series candidate segment, and determine the context information of any time series candidate segment through multiple two-dimensional convolutional layers to obtain the boundary matching confidence map.
[0095] Specifically, combine all possible time series candidates into a two-dimensional boundary matching map according to the position of the start boundary of the time series candidate and the length of the time series candidate. The time series candidates on each column of this two-dimensional boundary matching map have the same start time, while the time series candidates on each row have the same time series length. This two-dimensional boundary matching map is mainly used to represent all potential time series candidate segments, and the value of each point represents the confidence score of the corresponding time series candidate. Therefore, the boundary matching confidence map can be generated to generate confidence scores for all time series candidates simultaneously.
[0096] It should be noted that in order to generate a boundary matching feature map from the time series feature sequence S F ∈R C×T and then generate a boundary matching confidence map based on the boundary matching feature map. For any time series candidate segment on the boundary matching map of size D×T N points need to be sampled from the time series feature sequence within its time series range to form m i,j ∈R C×N as the feature of this candidate segment. To make the entire sampling process accurate and efficient, the feature sampling processes of all candidate segments can be carried out simultaneously. For a time series candidate segment with start and end times t s and t e sample N points within its extended time series range [t s -0.25d, t e +0.25d] to construct a sampling matrix w i,j ∈R N×T . The value w n corresponding to the nth sampling point t i,j,n ∈R T is defined as:
[0097]
[0098] where dec represents the decimal operation and floor represents the floor operation. Then perform operations on the time series dimension of the time series feature sequence S F ∈RC×T and w i,j ∈R N×T Perform a dot product to obtain m i,j ∈R C×N :
[0099]
[0100] The features of the candidate segments here can be calculated by linear interpolation in the form of dot products. Finally, the sampling matrix is extended from w i,j ∈R N×T to W ∈ R N×T×D×T , and through the same dot product operation, the boundary matching feature map M ∈ R of all temporal candidate segments can be obtained C×N×D×T . Through this matrix dot product method, accurate corresponding feature representations can be efficiently generated for all temporal candidate segments. Finally, a three-dimensional convolutional network is used to eliminate the number of sampling points N, and the context information around each temporal candidate segment is mined through multiple two-dimensional convolutional layers, and finally a boundary matching confidence map is generated.
[0101] Based on any of the above embodiments, inputting the candidate time segment data into the classification and recognition network to obtain the individual event type and individual event time information includes:
[0102] Input the data stream features of each candidate time segment into the classification and recognition network, classify based on the boundary matching confidence map to determine the data time period corresponding to each candidate time segment, and obtain the individual event category, individual event start time, and individual event end time in each candidate time segment.
[0103] Specifically, input the data stream features of each segment into the classification and recognition network, select the corresponding data time period for classification according to the generated boundary matching confidence map, obtain the final individual event category, and the final output detection results include the event category and the corresponding start and end times.
[0104] Based on any of the above embodiments, determining the meta-learning tasks of different user individuals, optimizing the fused features, the candidate time segment data, the individual event type, and the individual event time information to obtain the event stream detection model includes:
[0105] Determine an initial meta-learning model, where the initial meta-learning model includes a first parameter corresponding to the fused features, a second parameter corresponding to the candidate time segment data, and a third parameter corresponding to the individual event type and the individual event time information;
[0106] Determine the model update step size hyperparameter, and obtain the meta-training task loss function composed of any user data in the initial meta-learning model;
[0107] Determine a single gradient descent update based on the first parameter, the second parameter, the third parameter, the model update step size hyperparameter, and the meta-training task loss function to obtain an updated meta-learning model;
[0108] Adapt and update the updated meta-learning model with a number of different user data to obtain the event stream detection model.
[0109] Specifically, for the optimization of the model, it is completed through individual adaptive meta-learning.
[0110] Due to differences in their living habits, the event data distributions of different individuals also vary to some extent. Modeling such individual differences helps to construct a more effective event detection model. If a unique model is trained for each user, it will consume a large amount of computing resources, and in actual usage scenarios, it is difficult for the data volume of each user to reach an ideal scale to meet the training of a complete model. In order to optimize the performance (localization and classification accuracy) of the event recognition and detection model according to user individual characteristics, the present invention uses a model-agnostic meta-learning method to adaptively optimize the event detection model, so that it can be optimized at the lowest cost on new user data to obtain an individually adaptive model. The present invention introduces an individual adaptive mechanism from the model update stage, which uses the historical information of the user as the guidance for model adaptation learning.
[0111] The adopted individual adaptive event detection method mainly learns to update the parameters of the model, so that it has good sensitivity in the event detection tasks of new users, and a small change in the model parameters can enable it to achieve good performance in any new user event detection task. This method has no restrictions on the form of the model, only requiring that the model can be optimized and trained through the method of gradient backpropagation, so it can be well applied to multi-modal event detection including modules such as localization and recognition. For the sake of more formal description, first assume that the event detection method adopted in this project consists of a feature fusion network f θ , a boundary matching network and a classification and recognition These three sub-models are composed.
[0112] First, without considering user differences, fully train the entire detection network framework on the complete training set to obtain the initial model parameters In order to adaptively optimize the initialized model for new test users, construct a meta-training task according to the training samples of user n The corresponding model parameters will be updated by the method of gradient descent as:
[0113]
[0114]
[0115]
[0116] Here, α is a hyperparameter used to specify the step size for model adaptation and update. It represents the loss function in the meta-training task composed of the data of the nth user for initializing the model. in the meta-training task.
[0117] To reduce the training cost during the adaptation and update of the initial model in practical applications, this project will use a single gradient descent update to calculate Similarly, the initial model can be adapted and updated on the data of N different users to obtain N sets of adapted model parameters. Finally, the original model parameters are updated to optimize the model parameters to make them most conducive to adapting to N different meta-training tasks:
[0118]
[0119] The above objective function can be optimized and trained using meta-gradient descent, and finally an event detection model with individual self-adaptive characteristics is obtained.
[0120] Next, the real-time event stream recognition system provided by the present invention will be described. The real-time event stream recognition system described below can be mutually corresponding and referred to with the real-time event stream recognition method described above.
[0121] Figure 3 is a schematic structural diagram of the real-time event stream recognition system provided by the present invention. As Figure 3 shown, it includes: a determination module 31 and a detection module 32, where:
[0122] The determination module 31 is used to determine the individual event to be recognized; the detection module 32 is used to input the individual event to be recognized into a pre-trained event stream detection model to obtain an individual event recognition result; among them, the event stream detection model is obtained after fusing the feature data of different individual event samples to obtain a fused feature, performing boundary matching on the fused feature, and performing classification and recognition based on individual self-adaptive meta-learning.
[0123] The present invention optimizes the event stream detection model by adopting a model-agnostic meta-learning strategy, realizes obtaining a detection model with individual self-adaptation at the lowest cost, and greatly improves the generalization ability of the model.
[0124] Figure 4 Illustrates a schematic structural diagram of an electronic device. As Figure 4As shown in the figure, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute a real-time event stream recognition method, which includes: determining an individual event to be recognized; inputting the individual event to be recognized into a pre-trained event stream detection model to obtain an individual event recognition result; wherein, the event stream detection model is obtained by fusing the feature data of different individual event samples to obtain a fused feature, performing boundary matching on the fused feature, and performing classification recognition based on individual adaptive meta-learning.
[0125] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0126] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the real-time event stream recognition method provided by the above-mentioned methods, which includes: determining an individual event to be recognized; inputting the individual event to be recognized into a pre-trained event stream detection model to obtain an individual event recognition result; wherein, the event stream detection model is obtained by fusing the feature data of different individual event samples to obtain a fused feature, performing boundary matching on the fused feature, and performing classification recognition based on individual adaptive meta-learning.
[0127] In another aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the real-time event stream recognition method provided by the above-mentioned various methods. The method includes: determining an individual event to be recognized; inputting the individual event to be recognized into a pre-trained event stream detection model to obtain an individual event recognition result; wherein, the event stream detection model is obtained by fusing different individual event sample data features to obtain a fused feature, performing boundary matching on the fused feature, and performing classification recognition based on individual adaptive meta-learning.
[0128] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0129] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A real-time event stream recognition method, characterized in that, Including: Determine the individual event to be recognized; Input the individual event to be recognized into a pre-trained event stream detection model to obtain an individual event recognition result; wherein, the event stream detection model is obtained by fusing different individual event sample data features to obtain a fusion feature, performing boundary matching on the fusion feature, and performing classification and recognition based on individual adaptive meta-learning; The event stream detection model is obtained through the following steps: Obtain video data features and sensor data features; Connect and fuse the video data features and the sensor data features to obtain the fusion feature; Input the fusion feature into a boundary matching network to obtain candidate time segment data; Input the candidate time segment data into a classification and recognition network to obtain the individual event type and individual event time information; Determine the meta-learning tasks of different user individuals, optimize the fusion feature, the candidate time segment data, the individual event type, and the individual event time information to obtain the event stream detection model; The determining the meta-learning tasks of different user individuals, optimizing the fusion feature, the candidate time segment data, the individual event type, and the individual event time information to obtain the event stream detection model includes: Determine an initial meta-learning model, where the initial meta-learning model includes a first parameter corresponding to the fusion feature, a second parameter corresponding to the candidate time segment data, and a third parameter corresponding to the individual event type and individual event time information; Determine a model update step size hyperparameter, and obtain a meta-training task loss function composed of any user data in the initial meta-learning model; Based on the first parameter, the second parameter, the third parameter, the model update step size hyperparameter, and the meta-training task loss function, determine a single gradient descent update to obtain an updated meta-learning model; Adapt and update the updated meta-learning model with several different user data to obtain the event stream detection model.
2. The real-time event stream recognition method according to claim 1, wherein The obtaining the video data features and the sensor data features includes: Input a video segment of any length into a preset video feature extractor to obtain the video data features; Input the sensor data corresponding to the video segment of any length into a preset sensor feature extractor to obtain the sensor data features.
3. The real-time event stream recognition method according to claim 1, characterized in that The connecting and fusing the video data features and the sensor data features to obtain the fusion feature includes: Connect the video data features and the sensor data features with the same feature dimension to obtain a connection feature; Perform feature fusion on the connection feature by a linear network to obtain the fusion feature with a unified representation of multi-source heterogeneous data.
4. The real-time event stream recognition method according to claim 1, characterized in that The inputting the fusion feature into a boundary matching network to obtain candidate time segment data includes: Determine the temporal feature sequence of the fusion feature; Based on the temporal candidate initial boundary position and the temporal candidate length of the temporal feature sequence, construct the temporal candidate into a boundary matching feature map, and obtain a boundary matching confidence map based on the boundary matching feature map; According to the confidence scores of the confidence map of the boundary match, extract the feature representation of the fused features to obtain the candidate time segment data.
5. The real-time event stream recognition method according to claim 4, characterized in that Based on the temporal candidate initial boundary positions and the temporal candidate lengths of the temporal feature sequences, constructing the temporal candidates into a boundary matching feature map, and obtaining the boundary match confidence map based on the boundary matching feature map, includes: Obtain any temporal candidate segment on the boundary matching map with a preset size; Obtain several first sampling points in the temporal feature sequence, and determine the several first sampling points as the first temporal feature sequence of the any temporal candidate segment; Obtain the start time and end time of the any temporal candidate segment, determine a preset extended temporal range based on the start time and the end time, and obtain several second sampling points with the same number as the several first sampling points within the preset extended temporal range; Determine the numerical value of each sampling point in the several second sampling points to obtain a second temporal feature sequence; Perform dot multiplication and linear interpolation on the first temporal feature sequence and the second temporal feature sequence to obtain the feature of any temporal candidate segment; Use a three-dimensional convolutional network to eliminate the number of sampling points in the feature of any temporal candidate segment, and determine the context information of the any temporal candidate segment through multiple two-dimensional convolutional layers to obtain the boundary match confidence map.
6. The real-time event stream recognition method according to claim 1, characterized in that Inputting the candidate time segment data into a classification and recognition network to obtain the individual event type and the individual event time information, includes: Input the data stream features of each candidate time segment into the classification and recognition network, classify the data time periods corresponding to each candidate time segment based on the boundary match confidence map, and obtain the individual event category, the individual event start time, and the individual event end time in each candidate time segment.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the steps of the real-time event stream recognition method according to any one of claims 1 to 6.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the real-time event stream recognition method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image watermark processing method and device based on artificial intelligence and electronic equipment
CN111160335A
Intelligent investor implementation method and system
CN111353013A