Event information generation method and device, program product and storage medium
By extracting and fusing feature sequences from user commands and video data, event information is generated, solving the problem of information mismatch in traditional methods and achieving more accurate and targeted event detection.
Patent Information
- Application Number
- CN202511130989.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-28
AI Technical Summary
Traditional event information generation methods ignore users' specific detection needs, resulting in generated information that does not match the actual application scenario and reduces the system's usability.
By acquiring user commands and video data, a pre-trained target model is used to extract text features and image feature sequences, which are then fused to generate event information. The model is then trained iteratively and optimized using a loss function to ensure that it can comprehensively understand user intent and video content.
It improves the accuracy and relevance of event information, generates information that better meets user needs, and enhances the accuracy of event detection.
Smart Images

Figure CN121033728A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and more specifically, to a method and apparatus for generating event information, a program product, and a storage medium. Background Technology
[0002] Traditional event information generation methods mostly rely on single-pattern analysis, such as vision-based event detection. This typically involves directly processing video data, using computer vision technology to identify objects, behaviors, or anomalies in the video, and generating corresponding event reports. However, these methods often neglect the specific detection needs of users, resulting in generated information that may not fully match the actual application scenario, thus reducing the system's practicality. Summary of the Invention
[0003] This application provides a method, apparatus, program product, and storage medium for generating event information, in order to at least solve the problem of inaccurate event information generation in related technologies.
[0004] According to one embodiment of this application, a method for generating event information is provided, comprising: acquiring a user instruction, and acquiring video data within a target area based on the user instruction; performing an event information generation operation within the target area on the user instruction and the video data using a pre-trained target model to obtain target event information; wherein the generation operation comprises: determining a first feature sequence using the user instruction, the first feature sequence including a vector representation of the text features of the user instruction; determining a second feature sequence using the video data, the second feature sequence including a vector representation of the image features of multiple frames in the video data; fusing the first feature sequence and the second feature sequence to obtain a target feature sequence; and generating the target event information based on the target feature sequence.
[0005] In an exemplary embodiment, the target model is a model obtained by iteratively training an initial model using a set of sample information. The iterative training process includes: acquiring sample information, wherein the sample information includes sample user instructions and sample video data corresponding to the sample user instructions, and the sample video data includes sample labels, which are used to represent the sample event category corresponding to the sample video data; inputting the sample information into the initial model, and having the initial model generate sample labels based on the sample information; calculating the difference between the sample labels and the actual labels using a loss function to obtain the loss value of the initial model, wherein the actual labels are used to represent the actual event category of the sample video data; and updating the parameters of the initial model based on the loss value to obtain the target model, wherein the loss value of the target model satisfies a preset threshold.
[0006] In an exemplary embodiment, the above-mentioned sample information is input into the above-mentioned initial model, and the initial model generates sample labels based on the above-mentioned sample information, including: performing a first feature extraction operation on the above-mentioned sample user instruction and the above-mentioned sample video data respectively, and outputting a first sample feature sequence and a second sample feature sequence, wherein the above-mentioned first sample feature sequence includes a vector representation of the text features of the above-mentioned sample user instruction, and the above-mentioned second sample feature sequence includes a vector representation of the image features of multiple frames of sample images in the above-mentioned sample video data; performing a feature fusion operation on the above-mentioned first sample feature sequence and the above-mentioned second sample feature sequence to obtain a target sample feature sequence; and generating the above-mentioned sample labels based on the above-mentioned target sample feature sequence.
[0007] In an exemplary embodiment, the first feature extraction operation includes: deleting abnormal characters from the sample user instruction to obtain a first sample instruction; performing word segmentation on the instruction text in the first sample instruction to obtain a first sample instruction sequence; and converting multiple word segments in the first sample instruction sequence into vector representations to obtain the first sample feature sequence.
[0008] In an exemplary embodiment, the first feature extraction operation further includes: performing a serialization operation on the sample video data to obtain a first sample video sequence including vector representations of each frame of sample images in the sample video data; aggregating the spatial information of the vector representations of each frame of sample images to obtain a second sample video sequence; parsing the second sample video sequence to obtain the motion trajectory and sample events of the sample objects, and generating the second sample feature sequence based on the vector representations corresponding to the motion trajectories and the vector representations corresponding to the sample events.
[0009] In one exemplary embodiment, performing a serialization operation on the aforementioned sample video data to obtain a first sample video sequence including vector representations of each frame of sample image in the aforementioned sample video data includes: extracting M frames of sample images from the aforementioned sample video data, where M is a natural number greater than 1; dividing the M frames of the aforementioned sample images into P sample image blocks, where P is a natural number greater than M; performing vector transformation on the P sample image blocks to obtain P sample image block vectors; and sorting the P sample image block vectors based on a preset sorting method to obtain the aforementioned first sample video sequence.
[0010] In an exemplary embodiment, performing a feature fusion operation on the first sample feature sequence and the second sample feature sequence to obtain a target sample feature sequence includes: converting all elements in the first sample feature sequence and the second sample feature sequence into a preset format to obtain a third sample feature sequence and a fourth sample feature sequence; performing at least one of the following fusion operations on the third sample feature sequence and the fourth sample feature sequence to obtain the target sample feature sequence: concatenating the fourth sample feature sequence to the third sample feature sequence, concatenating the third sample feature sequence to the fourth sample feature sequence, and concatenating the third sample feature sequence and the fourth sample feature sequence element by element.
[0011] In an exemplary embodiment, generating the sample label based on the target sample feature sequence includes: generating a fifth sample feature sequence using the target sample feature sequence, wherein the fifth sample feature sequence includes a plurality of sample words, the sample words being used to represent event features in the sample video data; and generating the sample label based on the fifth sample feature sequence.
[0012] In an exemplary embodiment, generating a fifth sample feature sequence using the target sample feature sequence includes: selecting a set of sample words from the sixth sample feature sequence and the target sample feature sequence, wherein the sixth sample feature sequence is a sequence including historical words generated during the previous iteration of training of the initial model; and selecting multiple sample words from the set of sample words to generate the fifth sample feature sequence.
[0013] According to another embodiment of this application, an event information generation apparatus is provided, including a first memory, a first processor, and a first computer program stored in the first memory and executable on the first processor. When the first processor executes the first computer program, it performs the following operations: acquiring a user instruction and acquiring video data within a target area based on the user instruction; using a pre-trained target model to perform an event information generation operation within the target area based on the user instruction and the video data to obtain target event information; wherein the generation operation includes: determining a first feature sequence using the user instruction, the first feature sequence including a vector representation of the text features of the user instruction; determining a second feature sequence using the video data, the second feature sequence including a vector representation of the image features of multiple frames in the video data; fusing the first feature sequence and the second feature sequence to obtain a target feature sequence; and generating the target event information based on the target feature sequence.
[0014] According to yet another embodiment of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0015] According to yet another embodiment of this application, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and the computer program is configured to perform the steps in any of the above method embodiments when it is run.
[0016] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0017] This application utilizes a trained target model to determine a first feature sequence and a second feature sequence based on user commands and video data within the target area obtained from user commands, respectively. Then, it generates target event information using the target feature sequence obtained by fusing the first and second feature sequences. This enables real-time reception and processing of user commands, adjusting the analysis strategy for video data according to the user commands, prioritizing certain types of behaviors or regions. By employing a fusion mechanism to deeply interact with the feature sequences corresponding to user text commands and video data, the model ensures a comprehensive understanding of user intent and video content, improving the accuracy of event detection and effectively generating target event information. Therefore, it solves the problem of inaccurate event information generation in related technologies, resulting in more targeted event information and improved accuracy. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the hardware environment for an event information generation method according to an embodiment of this application;
[0019] Figure 2 This is a flowchart of an event information generation method according to an embodiment of this application;
[0020] Figure 3 This is a schematic diagram of the structure of an event information generation model according to an embodiment of this application;
[0021] Figure 4 This is a flowchart of a method for generating target event information according to an embodiment of this application;
[0022] Figure 5 This is a structural block diagram of an event information generation apparatus according to an embodiment of this application. Detailed Implementation
[0023] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0025] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a schematic diagram of the hardware environment for an event information generation method according to an embodiment of this application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0026] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to an event information generation method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the aforementioned method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0027] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0028] This embodiment provides a method for generating event information. Figure 2 This is a flowchart of an event information generation method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:
[0029] Step S202: Obtain user instructions and acquire video data within the target area based on the aforementioned user instructions;
[0030] Optionally, the solution in this embodiment can be applied to fields requiring real-time monitoring, event detection, and alarms. Application scenarios include, but are not limited to, scenarios requiring real-time monitoring of specific roads or intersections (e.g., monitoring congestion, violations, or traffic accidents), scenarios for security monitoring and analysis of public places to detect abnormal events (e.g., setting up cameras in schools, shopping malls, and parks, specifying areas of interest through user instructions, promptly detecting abnormal behavior and generating event reports), and scenarios requiring intelligent security alarms (e.g., identifying security events such as intrusion and theft in residential areas and corporate parks). The solution in this embodiment can also be applied to industrial production, event sites, hospital intensive care units or nursing homes, and wildlife reserves.
[0031] Optionally, in this embodiment, the user instructions refer to those provided by the user through natural language or structured queries, used to adjust the working mode of the target device acquiring video data to guide the focus on specific types of events or specific areas. User instructions include, but are not limited to, details such as event type (e.g., "traffic accident", "illegal parking"), monitoring area (e.g., "school gate", "highway entrance"), and time range.
[0032] Optionally, the video data within the target area in this embodiment is real-time or historical video data collected from a target device (e.g., a camera) in the designated area according to user instructions. The video data includes continuous multi-frame image information of the monitored area and can be used for object detection, behavior analysis, and event recognition.
[0033] Step S204: Using a pre-trained target model, perform an event information generation operation on the user instruction and the video data within the target area to obtain target event information; wherein, the generation operation includes: determining a first feature sequence using the user instruction, the first feature sequence including a vector representation of the text features of the user instruction; determining a second feature sequence using the video data, the second feature sequence including a vector representation of the image features of multiple frames in the video data; fusing the first feature sequence and the second feature sequence to obtain a target feature sequence; and generating the target event information based on the target feature sequence.
[0034] Optionally, the target model in this embodiment is a pre-trained neural network model capable of processing user commands and video data to generate event information. For example, the target model may combine the Transformer architecture and the Mamba module to efficiently extract and fuse visual and semantic features to generate structured event reports.
[0035] Optionally, in this embodiment, the first feature sequence is a sequence of text feature vectors extracted from the user instruction, containing the semantic and contextual information of the instruction. The first feature sequence may be converted into a vector using techniques such as embedding layers and positional encoding, facilitating the target model's subsequent understanding of the meaning and key points of the instruction.
[0036] Optionally, in this embodiment, the second feature sequence is a sequence of image feature vectors extracted from video data, reflecting dynamic information about objects, behaviors, and scenes in the video. The second feature sequence can be extracted using a hybrid architecture of the Swin Transformer and Mamba modules.
[0037] Optionally, the target feature sequence in this embodiment is the result of fusing the first feature sequence and the second feature sequence, including comprehensive information from user instructions and video data, which integrates the understanding of the visual scene and guidance for the user's specific needs.
[0038] Optionally, the target event information in this embodiment is generated based on the target feature sequence and is used to describe detailed information about a specific event in the video data. The target event information may include event type, timestamp, location description and related attributes, such as "traffic accident", "2023-04-15 14:30", "XX National Highway, mileage marker 123km", "involving two cars, no injuries".
[0039] Optionally, the target model in this embodiment may include a first feature extraction subnetwork, a second feature extraction subnetwork, a feature fusion subnetwork, and an information generation subnetwork. The corresponding generation operations may include: the first feature extraction subnetwork performing a first feature extraction operation on the user instruction and outputting a first feature sequence, which includes a vector representation of the text features of the user instruction; the second feature extraction subnetwork performing a second feature extraction operation on the video data and outputting a second feature sequence, which includes a vector representation of the image features of multiple frames in the video data; inputting the first feature sequence and the second feature sequence into the feature fusion subnetwork, which performs a feature fusion operation on the first feature sequence and the second feature sequence and outputs a target feature sequence; and finally, the information generation subnetwork generating target event information based on the target feature sequence.
[0040] In this embodiment, the entities executing the above steps can be widely distributed intelligent processing nodes, not limited to ground facilities, but also including devices in the air and even space. These can be intelligent processing devices installed in satellites, edge computing devices, ground control centers or data centers, or cloud computing platforms. For example, in applications requiring remote, large-scale monitoring such as forest fire early warning, marine monitoring, and rapid response to natural disasters, the executing entity for the above steps in this embodiment is an intelligent processing device installed in the satellite. This device can be a dedicated processor or controller responsible for receiving user instructions, processing video data, and generating event information. In scenarios requiring centralized processing of large amounts of data (e.g., security monitoring of large-scale events, public safety emergency response centers), the executing entity for the above steps in this embodiment is a control center or data center set up on the ground. This center receives video data and user instructions from multiple data sources and utilizes powerful computing resources for in-depth analysis and event information generation. In applications such as intelligent transportation, urban security monitoring, and industrial automation, the executing entity for the above steps in this embodiment is an edge computing node deployed in monitoring devices (such as cameras) or sensor networks near the target area. This node can directly process the collected video data and user instructions, perform feature extraction and event judgment, and then upload key information to the cloud or control center. Instead of uploading large amounts of raw data to the cloud, this node processes the data in real time at the edge, reducing data transmission volume and preventing the leakage of privacy information.
[0041] For example, in urban traffic management, management departments can use the method provided in this embodiment to monitor traffic violations at specific times and locations using an intelligent monitoring system. If the user instruction is "monitor illegal parking near the city center square from 8:00 AM to 10:00 AM on Monday," the system will acquire video data from cameras near the square from 8:00 AM to 10:00 AM, analyze the video content using a target model, and combine this with the semantic features of the user instruction to determine whether, when, and where an illegal parking incident occurred, and generate a detailed event report, such as "An illegal parking incident occurred at 08:45 on April 17, 2023, on the north side of the city center square, involving a blue truck with license plate number XXXXXXXX." Through the above method, the intelligent monitoring system can not only detect and describe events in real time, but also flexibly adjust its behavior according to the specific needs of the user, providing strong technical support for urban safety and traffic management.
[0042] Through the above steps, using a trained target model, a first feature sequence and a second feature sequence are determined based on user commands and video data within the target area obtained based on user commands, respectively. Then, the target feature sequence obtained by fusing the first and second feature sequences is used to generate target event information. This enables real-time reception and processing of user commands, adjusting the analysis strategy for video data according to the guidance of user commands, prioritizing certain types of behaviors or regions. By utilizing a fusion mechanism to deeply interact with the feature sequences corresponding to user text commands and video data, the model ensures a comprehensive understanding of user intent and video content, improving the accuracy of event detection and effectively generating target event information. Therefore, it solves the problem of inaccurate event information generation in related technologies, thereby achieving more targeted event information and improving the accuracy of generated event information.
[0043] In an exemplary embodiment, the target model is a model obtained by iteratively training an initial model using a set of sample information. The iterative training includes: acquiring sample information, wherein the sample information includes sample user instructions and sample video data corresponding to the sample user instructions, and the sample video data includes sample labels, which are used to represent the sample event category corresponding to the sample video data; inputting the sample information into the initial model, and having the initial model generate sample labels based on the sample information; calculating the difference between the sample labels and the actual labels using a loss function to obtain the loss value of the initial model, wherein the actual labels are used to represent the actual event category of the sample video data; and updating the parameters of the initial model based on the loss value to obtain the target model, wherein the loss value of the target model satisfies a preset threshold.
[0044] Optionally, the sample information in this embodiment is a specific dataset used to train the initial model, including sample user instructions, video data corresponding to the instructions, and sample labels used to represent sample event categories.
[0045] Optionally, the initial model in this embodiment consists of four sub-networks for processing text instructions and video data and generating event information. The four sub-networks of the initial model are: a first feature extraction sub-network, a second feature sub-network, a feature fusion sub-network, and an information generation sub-network.
[0046] The iterative training of the initial model specifically includes: collecting samples containing user commands, video data, and actual labels to provide examples for training; feature extraction, where user commands and video data are fed into the first and second feature extraction sub-networks respectively to extract text features and image features, and output their respective feature sequences; feature fusion, where the extracted text feature and image feature sequences are fed into the feature fusion sub-network to generate a fused feature sequence that integrates user intent and video content; event information generation, where the fused feature sequence is fed into the information generation sub-network to generate sample labels; and loss function and model update, where the model performance is evaluated by calculating the difference between the actual labels output by the model and the sample labels. If the loss value exceeds a preset threshold, the model parameters are updated until the model performance meets expectations.
[0047] For example, in an intelligent traffic monitoring scenario, the first feature extraction subnetwork uses a BERT encoder to process the user command "detect a rear-end collision on the highway," while the second feature extraction subnetwork uses a Swin Transformer to process the real-time video stream, and a Mamba module further extracts long-range temporal features. Subsequently, the feature fusion subnetwork deeply integrates the textual features of the user command with the visual features of the video through a multimodal attention mechanism. Finally, the information generation subnetwork uses a BART structure to generate an event report and corresponding sample labels. The event report content is as follows: "At 14:27 on March 15, 2023, a rear-end collision was discovered on the XX highway. The vehicles involved were two cars. No injuries were observed. The accident was located at 300km. Immediate dispatch of rescue is recommended." The sample label is "Rear-end collision occurred." This architecture not only efficiently processes video and text information but also ensures that the model can accurately understand and respond to user commands, generating event information closely related to the video content, providing strong technical support for intelligent monitoring and event response systems.
[0048] Through the above steps, the initial model is iteratively trained by inputting the sample information set into it to obtain sample labels. Then, the difference between the sample labels and the actual labels is calculated using a loss function to obtain the loss value of the initial model. Based on the determined loss value, the parameters of the initial model are updated to obtain the target model. This allows the trained target model to effectively extract and fuse features from user commands and video data, accurately generate sample labels, reduce errors during training, and ensure the high performance and generalization ability of the target model.
[0049] In an exemplary embodiment, the above-mentioned sample information is input into the above-mentioned initial model, and the initial model generates sample labels based on the above-mentioned sample information, including: performing a first feature extraction operation on the above-mentioned sample user instruction and the above-mentioned sample video data respectively, and outputting a first sample feature sequence and a second sample feature sequence, wherein the above-mentioned first sample feature sequence includes a vector representation of the text features of the above-mentioned sample user instruction, and the above-mentioned second sample feature sequence includes a vector representation of the image features of multiple frames of sample images in the above-mentioned sample video data; performing a feature fusion operation on the above-mentioned first sample feature sequence and the above-mentioned second sample feature sequence to obtain a target sample feature sequence; and generating the above-mentioned sample labels based on the above-mentioned target sample feature sequence.
[0050] Optionally, in practical use, the first feature extraction operation is performed on the sample user instructions and the sample video data respectively, and the output first sample feature sequence and second sample feature sequence can be performed by the first feature extraction subnetwork and the second feature subnetwork in the initial model. The first feature extraction subnetwork performs the first feature extraction operation on the sample user instructions and outputs the first sample feature sequence; the second feature extraction subnetwork performs the second feature extraction operation on the sample video data and outputs the second sample feature sequence.
[0051] Optionally, the first feature extraction sub-network in this embodiment is mainly used to process and extract text features from user instructions. The first feature extraction sub-network can be based on a Transformer architecture, such as a Transformer-based encoder, Bidirectional Encoder Representations from Transformers (BERT), or a robustly optimized BERT approach (RoBERTa).
[0052] Optionally, the second feature extraction subnetwork in this embodiment is responsible for extracting features from the video data. The first feature extraction sub-network can be a hybrid coding architecture of Swin Transformer + Mamba. Swin Transformer can perform multi-head self-attention computation within a local window and within a limited cross-window range, efficiently capturing fine spatial features within video frames and short-range inter-frame spatiotemporal changes. The output is a series of feature sequences that are spatially aggregated but still retain the frame temporal order. For each time step (frame), Swin Transformer outputs a feature vector sequence representing the spatial information of that frame. The Mamba module focuses on processing long-range temporal series features of the video. It uses the efficient sequence modeling capability of the Selective State Space Model (Selective SSM) to process long-range time series. The Mamba module captures long-range temporal dependencies spanning tens or even hundreds of frames with linear computational complexity. The Mamba module can capture the motion trajectory of objects over a long period of time, identify and analyze continuous events that occur over a long period of time (e.g., vehicle queues, traffic congestion, abnormal personnel lingering), understand the slow or rapid changes in scene background (e.g., lighting, weather, etc.) over time, and discover the correlation between events with long intervals.
[0053] Optionally, in practical applications, the feature fusion operation performed on the first and second sample feature sequences to obtain the target sample feature sequence can be performed by the feature fusion sub-network in the initial model. The feature fusion sub-network performs the feature fusion operation on the first and second sample feature sequences and outputs the target sample feature sequence. The feature fusion sub-network includes, but is not limited to, the following fusion mechanisms: multimodal attention mechanism, feature concatenation and mapping (concatenating text feature sequences with image feature sequences and then transforming them through linear mapping or fully connected layers to create a unified fused feature representation), Bidirectional Long Short-Term Memory Network (BiLSTM), or Recurrent Neural Networks (RNNs).
[0054] Optionally, in practical use, the generation of sample labels based on the target sample feature sequence can be performed by the information generation sub-network in the initial model. The information generation sub-network generates target sample event information based on the target sample feature sequence, wherein the target sample event information includes the actual label.
[0055] Optionally, the information generation sub-network in this embodiment includes, but is not limited to, the following architectures: Transformer decoder or bidirectional and auto-regressive transformers (BART), sequence to sequence (Seq2Seq).
[0056] Through the above steps, the initial model first performs a first feature extraction operation on the sample user commands and sample video data to obtain a first sample feature sequence and a second sample feature sequence. Then, it performs a feature fusion operation on the first and second sample feature sequences to obtain a target sample feature sequence. Finally, it generates sample labels based on the target sample feature sequence. Through the step-by-step processing of the initial model, sample labels can be determined simultaneously based on the sample user commands and sample video information included in the sample information. This ensures that the initial model can comprehensively understand user intent and video content, improve the accuracy of event detection, and effectively generate sample labels.
[0057] In an exemplary embodiment, the first feature extraction operation includes: deleting abnormal characters from the sample user instruction to obtain a first sample instruction; performing word segmentation on the instruction text in the first sample instruction to obtain a first sample instruction sequence; and converting multiple word segments in the first sample instruction sequence into vector representations to obtain the first sample feature sequence.
[0058] Optionally, the first feature extraction operation in this embodiment may be performed by the first feature extraction sub-network in the initial model that is used to process text instructions, converting natural language instructions into vector sequences that can be understood and processed by computers. The first feature extraction sub-network includes a preprocessing layer, a word segmentation layer, and a vector transformation layer, with the aim of extracting key information from user instructions and converting them into feature sequences that can be operated by the model subsequently.
[0059] Optionally, in this embodiment, in addition to deleting abnormal characters (such as special symbols and non-printable characters) in the sample user instructions, operations such as unifying text format (such as converting uppercase and lowercase, removing extra spaces) and correcting spelling errors can also be performed on the sample user instructions.
[0060] Optionally, the first sample instruction in this embodiment is a clean, uniformly formatted instruction text obtained after the deletion operation. For example, if the sample user instruction is "Detect all vehicles running red lights, #&@", after the above processing, the first sample instruction will become "Detect all vehicles running red lights".
[0061] Optionally, the word segmentation operation in this embodiment is the process of decomposing the first sample instruction into a series of meaningful phrase sequences for subsequent feature extraction and encoding.
[0062] Optionally, in this embodiment, converting multiple word segments in the first sample instruction sequence into vector representations can be done through vector transformation operations. Vector transformation operations include, but are not limited to, converting each phrase sequence into a fixed-size vector through a word embedding model (such as Word2Vec, GloVe, FastText, etc.), which contains the semantic information of the phrase sequence, or encoding through a vector transformation layer to map the phrase sequence to a continuous vector space and combining positional encoding to generate a vector sequence containing semantic content and word order information, thus obtaining the first sample feature sequence.
[0063] Optionally, in this embodiment, the first sample feature sequence is a sequence composed of word embedding vectors in the first sample instruction sequence after vector transformation. The first sample feature sequence is a semantic feature vector representing the semantics of the user instruction, which accurately expresses the user's specific requirements and focus of attention for video content analysis.
[0064] Through the above steps, the first sample instruction is obtained by performing a deletion operation on the sample user instruction. Then, the first sample instruction is subjected to word segmentation and vector transformation operations to obtain the first sample feature sequence. The deletion and word segmentation operations improve the clarity and quality of the subsequently obtained first sample feature sequence, and the vector transformation makes it easier for the model to understand the meaning of the instruction, making feature extraction more accurate and obtaining a more accurate first sample feature sequence.
[0065] In an exemplary embodiment, the first feature extraction operation further includes: performing a serialization operation on the sample video data to obtain a first sample video sequence including vector representations of each frame of sample images in the sample video data; aggregating the spatial information of the vector representations of each frame of sample images to obtain a second sample video sequence; parsing the second sample video sequence to obtain the motion trajectory and sample events of the sample objects, and generating the second sample feature sequence based on the vector representations corresponding to the motion trajectories and the vector representations corresponding to the sample events.
[0066] Optionally, the first feature extraction operation in this embodiment may be performed by the second feature extraction sub-network used for video data processing in the initial model. The second feature extraction sub-network may be a hybrid coding architecture of Swin Transformer + Mamba, that is, Swin Transformer is responsible for local and global spatial feature extraction, while the Mamba module handles long-range time series modeling, ensuring that the second feature extraction sub-network can efficiently process image patch sequences and capture the dynamic changes of objects and the long-term evolution of events in the video.
[0067] Optionally, the sample video data in this embodiment includes continuous multi-frame video data of traffic accidents or other events that need to be detected. The sample video data covers different event occurrence scenarios and different environmental conditions, such as daytime, nighttime, rainy or snowy weather, so that the model can learn and adapt to various situations comprehensively.
[0068] Optionally, the serialization operation in this embodiment refers to the operation of decomposing the input video data into a series of independent image block vector sequences to obtain the first sample video sequence.
[0069] Optionally, the aggregation operation in this embodiment is used to integrate or summarize information from adjacent image blocks or consecutive video frames to form a higher-level feature representation. The aggregation operation can be performed in the Swing Transformer part through a multi-head self-attention mechanism within a local window and a limited cross-window range. It is used to extract and fuse fine spatial features within the frame and spatiotemporal changes between short-range frames to obtain a series of feature sequences that have been spatially aggregated but still retain the frame temporal order, i.e., a feature vector sequence representing the spatial information of the frame.
[0070] Optionally, the parsing operation in this embodiment refers to the process of performing in-depth analysis on the aggregated second sample video sequence to identify and classify the motion trajectories of targets (such as vehicles and pedestrians) and their related events (i.e., the aforementioned sample events) in the video. The parsing operation can be performed by the second feature extraction subnetwork in the initial model. In practical use, it can be a parsing operation performed using the Mamba module to model the second sample video sequence, analyze the motion trajectories of objects in the video over a period of time, and identify and classify the behavior or events of the objects. For example, in the traffic accident detection example, the target processing operation includes analyzing the motion trajectory of vehicles in multiple video frames to determine whether a collision or emergency braking event has occurred. If a car suddenly decelerates in several consecutive frames and stops in the next frame, the second feature extraction subnetwork will identify this behavior as a potential accident precursor and further analyze the surrounding environment and the reactions of other vehicles.
[0071] Optionally, in this embodiment, the second sample feature sequence is a high-dimensional feature vector sequence that integrates fine spatial details and long-range temporal correlations. The second sample feature sequence is a high-level semantic representation of the video content, containing key information such as objects, events, and scene dynamics in the video, and is used for subsequent event detection.
[0072] For example, suppose there is a video clip where a car suddenly brakes while driving normally, causing a rear-end collision as the vehicle behind cannot react in time. In this case, the second feature extraction subnetwork located at the edge first performs a serialization operation, segmenting the video data into a series of image patch vector sequences (corresponding to the first sample video sequence mentioned above). Next, the Swin Transformer captures the vehicle's speed changes and positional offset information through aggregation operations, obtaining a feature vector sequence containing spatial information for each frame of the video (corresponding to the second sample video sequence mentioned above). Then, the Mamba module analyzes the feature vector sequence containing spatial information for each frame of the video, determining the motion trajectory of the vehicle (corresponding to the sample object mentioned above) and the events that occur, resulting in a high-dimensional feature sequence that integrates fine spatial details and long-range temporal correlations (corresponding to the second sample feature sequence mentioned above). Through this process, the Swin Transformer efficiently processes intra-frame spatial structure and short-range spatiotemporal correlations, providing feature representations with rich spatial context at each time step. Mamba, based on this, efficiently models the long-range dependencies of these feature sequences along the time axis, thereby comprehensively understanding the dynamic evolution of the video. This decoupling and synergy between spatial feature extraction and efficient time series modeling enables the second feature extraction sub-network to maintain high performance while reducing the computational and memory complexity of traditional global self-attention Transformers when processing high-resolution, long-time-series videos. This allows for real-time and efficient video analysis on edge devices.
[0073] Through the above steps, the sample video data is first serialized to obtain the first sample video sequence. Then, the first sample video sequence is aggregated to obtain the second sample video sequence, which is spatially aggregated but still retains the frame temporal order. Finally, the second sample video sequence is parsed to obtain a high-dimensional second sample feature sequence that integrates fine spatial details and long-range temporal correlations. The serialization operation ensures the orderliness and efficiency of video data processing, while the aggregation and parsing operations help capture the motion trajectory and event details of objects in the video data, improving the accuracy of video content analysis and resulting in a second sample feature sequence with richer and more compact features.
[0074] In one exemplary embodiment, performing a serialization operation on the aforementioned sample video data to obtain a first sample video sequence including vector representations of each frame of sample image in the aforementioned sample video data includes: extracting M frames of sample images from the aforementioned sample video data, where M is a natural number greater than 1; dividing the M frames of the aforementioned sample images into P sample image blocks, where P is a natural number greater than M; performing vector transformation on the P sample image blocks to obtain P sample image block vectors; and sorting the P sample image block vectors based on a preset sorting method to obtain the aforementioned first sample video sequence.
[0075] Optionally, in this embodiment, the P sample image blocks are fixed-size, non-overlapping image blocks segmented from M frame sample images. The segmentation operation can be performed using the Vision Transformer method.
[0076] Optionally, in this embodiment, the vector transformation of P sample image blocks to obtain P sample image block vectors can be: the P sample image blocks are transformed into initial embedding vectors through a linear projection layer to obtain P sample image block vectors.
[0077] Optionally, in this embodiment, the sample image block vector is a set of high-dimensional vectors that arrange the sample image blocks along the time dimension through linear projection, and each vector represents the feature representation of an image block.
[0078] Optionally, in this embodiment, the first sample video sequence is an ordered vector sequence formed after vector transformation and sorting of the sample image blocks. The preset sorting method refers to organizing the image blocks according to their position and time order in the video frame. First, the P sample image block vectors are arranged into an intra-frame sequence in the spatial dimension, and then arranged across frames along the temporal dimension to form a unified first sample video sequence containing both spatial and temporal information.
[0079] Through the above steps, the sample video data is decomposed to obtain M frame sample images. Then, the M frame sample images are divided into P non-overlapping sample image blocks of fixed size. The P sample image blocks are sorted according to a preset sorting method to obtain the first sample video sequence. The integrity of the video data is ensured through a more detailed video frame processing flow. The temporal and spatial structure of the sequence is maintained by sorting, which helps to learn and utilize the temporal features of the video more effectively in the future.
[0080] In an exemplary embodiment, performing a feature fusion operation on the first sample feature sequence and the second sample feature sequence to obtain a target sample feature sequence includes: converting all elements in the first sample feature sequence and the second sample feature sequence into a preset format to obtain a third sample feature sequence and a fourth sample feature sequence; performing at least one of the following fusion operations on the third sample feature sequence and the fourth sample feature sequence to obtain the target sample feature sequence: concatenating the fourth sample feature sequence to the third sample feature sequence, concatenating the third sample feature sequence to the fourth sample feature sequence, and concatenating the third sample feature sequence and the fourth sample feature sequence element by element.
[0081] Optionally, in this embodiment, after converting the elements in the first sample feature sequence and the elements in the second sample feature sequence into a preset format, the two feature sequences can be synchronized according to a certain rule to ensure that they are uniform in format, thus obtaining the third sample feature sequence and the fourth sample feature sequence.
[0082] Optionally, the fusion operation in this embodiment refers to merging the third sample feature sequence and the fourth sample feature sequence into a unified feature representation, so that the model can simultaneously process information from text instructions and video data. The fusion operation can be an operation that concatenates the third sample feature sequence and the fourth sample feature sequence in different ways to obtain the target sample feature sequence. For example, the feature sequence of the user request to "detect all vehicle violations" (corresponding to the aforementioned third sample feature sequence) is concatenated with the feature sequence of vehicle trajectories in the video (corresponding to the aforementioned fourth sample feature sequence) to obtain the target sample feature sequence, which is used for further event detection and recognition.
[0083] For example, suppose a user sends a command to a traffic monitoring system via a mobile application: "Detect all rear-end collisions on a certain road section between 10:00 AM and 11:00 AM on October 10th." The system first converts this semantic command into a first sample feature sequence using embedding and location encoding, capturing the time, location, and event type. Simultaneously, the system extracts multiple consecutive frames of video data from the monitoring video during that time period. It then uses a Swing Transformer to extract local and global features from each frame, and processes the long-range time series using the Mamba module to obtain a second sample feature sequence, which includes the dynamic changes of vehicles and pedestrians in the video and the long-term evolution features of the events. Next, the system performs an alignment operation, unifying the format of the first sample feature sequence (i.e., the feature representation of the user command) and the second sample feature sequence (i.e., the spatiotemporal feature representation of the video content) to ensure they can be processed in a correlated manner. The aligned sequences are then labeled as the third and fourth sample feature sequences, respectively. Finally, the system performs a fusion operation. For example, the system might choose to concatenate the third sample feature sequence (user instruction features) after the fourth sample feature sequence (video content features) to obtain a target sample feature sequence. This sequence contains the spatiotemporal information of the video data and the user's focus on specific time, location, and event type. This allows the model to focus on detecting rear-end collisions on a certain road segment between 10:00 AM and 11:00 AM on October 10th, based on the user's specific requirements, thereby improving the accuracy and efficiency of detection. Through the above process, the system can not only intelligently analyze video data according to user instructions but also provide a foundation for subsequently generating detailed event information including timestamps, geographical locations, and event types. This provides traffic management departments with immediate accident reports, helping them to respond quickly and take appropriate rescue or traffic control measures.
[0084] Through the above steps, a fusion operation is performed on the third and fourth sample feature sequences, which have been converted to a preset format, to obtain the target sample feature sequence. The conversion operation ensures that the first and second sample feature sequences are identical in format. Feature fusion ensures that user instructions are closely integrated with video content, thereby enhancing the accuracy and flexibility of event detection.
[0085] In an exemplary embodiment, generating the sample label based on the target sample feature sequence includes: generating a fifth sample feature sequence using the target sample feature sequence, wherein the fifth sample feature sequence includes a plurality of sample words, the sample words being used to represent event features in the sample video data; and generating the sample label based on the fifth sample feature sequence.
[0086] Optionally, in practical use, the target sample feature sequence can be input into the information generation subnetwork of the initial model, and the information generation subnetwork can generate target sample event information including sample labels (i.e. sample event types) based on the target sample feature sequence. The target sample event information includes at least one of the following: sample event occurrence time, sample event occurrence location, and sample event attributes.
[0087] Optionally, the target sample feature sequence in this embodiment is a comprehensive sequence that integrates visual and semantic features, including a deep understanding of the video content and explicit requirements of user instructions.
[0088] Optionally, the fifth sample feature sequence in this embodiment is a sequence generated after decoding by the information generation sub-network, including a series of sample words representing event features.
[0089] Optionally, the sample words in this embodiment are text elements predicted by the information generation subnetwork that can describe the events occurring in the video. They can be words, phrases, or tags. The selection and generation of sample words reflect the information generation subnetwork's understanding of the video content and its response to user commands.
[0090] Optionally, in this embodiment, while generating sample labels based on the fifth sample feature sequence, structured data, namely target sample event information, is also generated based on the fifth sample feature sequence. The target sample event information includes key information such as the type, time, location, and attributes of the event. It is a summary and description of the event recognition results requested by the user after intelligent analysis of the video content. For example, in an intelligent traffic monitoring system, if the model detects a traffic accident involving pedestrians that occurred at 10:30 AM on October 10th on a main road in the city center, the target sample event information might be in the form shown in Table 1. Event information generated in this way not only provides basic information about the event but also incorporates aspects of particular concern in the user's instructions (such as whether pedestrians were involved), facilitating relevant departments to quickly understand the details of the event and take timely measures. The target sample event information can be a text sequence that follows a specific JavaScript Object Notation (JSON) pattern. This pattern defines the structure and fields of the output JSON object to ensure that the generated event information is standardized. In the event information shown in Table 1, eventType (event type identifier) is used to represent the event type, which is consistent with the event type that the user command is interested in. Timestamp represents the timestamp or time range of the event. Location represents the spatial location description of the event. Attributes is used to represent other attribute information related to a specific event type, presented in the form of nested JSON objects or arrays.
[0091] Table 1:
[0092]
[0093] For example, suppose a traffic monitoring system receives a user instruction: "Detect and report all rear-end collisions that occurred at the highway entrance between 9:00 AM and 10:00 AM today." The system first generates a target sample feature sequence based on the user instruction and video content. This sequence contains the feature representation of the user instruction and the spatiotemporal feature representation of the video footage at the highway entrance during the corresponding time period. Then, the system uses the target sample feature sequence to decode and generate a fifth sample feature sequence, which may contain lexical units such as "HighwayEntrance," "Tailgating," "Collision," "Morning," and "9AM to 10AM," all of which are event features identified from the video content and associated with the user instruction. Finally, the system generates structured target sample event information and sample labels based on the fifth sample feature sequence. The target sample event information details a rear-end collision that occurred at Highway Entrance No. 1, specifically at 09:45:00 on October 10, 2023, involving a blue sedan and a red truck. The system notes that traffic was heavy and the weather was clear that day, and the sample label indicates a rear-end collision occurred. Such incident information clearly conveys the key details of traffic accidents, facilitating timely response, investigation, and handling by traffic management departments.
[0094] Through the above steps, a fifth sample feature sequence is generated using the target sample feature sequence. Then, sample labels are generated based on the fifth sample feature sequence. The target sample feature sequence provides concise sample labels that describe the event, which is beneficial for rapid response and handling of the event, and also facilitates data analysis and later backtracking.
[0095] In an exemplary embodiment, generating a fifth sample feature sequence using the target sample feature sequence includes: selecting a set of sample words from the sixth sample feature sequence and the target sample feature sequence, wherein the sixth sample feature sequence is a sequence including historical words generated during the previous iteration of training of the initial model; and selecting multiple sample words from the set of sample words to generate the fifth sample feature sequence.
[0096] Optionally, the sixth sample feature sequence in this embodiment is a sequence generated by the information generation subnetwork based on the sample words generated in the previous round during the multi-round autoregressive decoding process. It reflects the model's further understanding of the current prediction result (sample word set) and is used to guide the prediction of words in the next round. For example, the Qth sample word is generated based on the sequence composed of the first Q-1 sample words and the target sample feature sequence, wherein the sequence composed of the first Q-1 sample words and the Qth sample word refer to the same video sequence.
[0097] Optionally, in this embodiment, the sample word set is a collection of all words generated during the decoding process, which gradually becomes richer as the decoding process progresses, ultimately forming a description of the entire event. The sample word set is a dynamically constructed lexicon used to subsequently generate structured event reports and sample tags.
[0098] Optionally, in this embodiment, multiple sample words can be selected from the sample word set based on preset conditions. The preset conditions may be filtering logic based on event type, time, location or other specific attributes. The preset conditions may be the sample words with the highest probability of selection.
[0099] Optionally, the sample words in this embodiment are selected from the sample word set. They are most relevant to the user's query command or event detection needs and are used for the final event description and report. For example, the fifth sample feature sequence is generated using an autoregressive approach. That is, the decoder predicts the target sample words in the fifth sample feature sequence one by one. The generation process begins with a special starting word as input. At each decoding time step, the decoder predicts the next sample word based on the currently generated sequence (i.e., the word sequence predicted from the starting word to the previous time step) and the target sample feature sequence to obtain the sample word set. Then, it is transformed by a feedforward neural network to determine the probability of the sample words in the sample word set. The sample word with the highest probability is determined as the final selected sample word. The prediction of the next target sample word to obtain the sample word set is completed by multiple mechanisms within the BART decoder layer: the decoder performs self-attention calculation on its own generated sequence to maintain the internal consistency and structure of the output text. At the same time, the decoder performs cross-attention calculation, using the self-attention output of the previous step as the query and the target sample feature sequence as the key and value. Through the cross-attention mechanism, the decoder can selectively focus on the part of the target sample feature sequence that is most relevant to the user's instructions, effectively injecting information from video data and user instructions into the prediction of the current sample word.
[0100] For example, suppose the system needs to detect and report traffic accidents occurring in a specific area, explicitly defined as "CityCenter" in the user command. During decoding, the information generation subnetwork generates a sixth sample feature sequence based on the sample tokens output from the previous round (e.g., "TrafficAccident", "Car", "12:34PM") and the target sample feature sequence (a sequence that fuses the user command "CityCenter" and video spatiotemporal features). This sequence contains more detailed event description features. As decoding progresses, the system generates a set of sample tokens, including tokens such as "TrafficAccident", "12:34PM", "CityCenter", "Crossroad", "SpeedingCar", and "StoppedVehicle". Next, the system selects multiple sample tokens from the sample token set according to preset conditions (e.g., the event must occur in CityCenter and be related to a traffic accident). In this example, the selected sample keywords might be "Traffic Accident," "12:34 PM," "CityCenter," "Crossroad," and "SpeedingCar," as these keywords are most closely related to the traffic accident information that users are interested in. Ultimately, the system organizes these target sample keywords into structured event information, generating a detailed event report that includes sample tags. The event report includes the type of accident (Traffic Accident, i.e., the aforementioned sample tags), the specific time (12:34 PM), the location (City Center Crossroad), and relevant attributes (vehicle traveling at high speed, cause of the accident being a speeding vehicle), thus providing traffic management departments with timely and accurate accident information to facilitate appropriate measures.
[0101] Through the above steps, the sixth sample feature sequence, which is composed of the target sample feature sequence and the previously generated sample words, is continuously generated to generate sample words, and finally the fifth sample feature sequence is obtained. The autoregressive word generation ensures the coherence and completeness of the event information description and improves the performance and stability of the model when generating multiple words or long sentences.
[0102] The above method will be illustrated with specific examples below:
[0103] An intelligent traffic management system, deployed on edge devices, aims to detect and report traffic accidents in the city center in real time. The system includes features such as... Figure 3 The event information generation model shown is as follows: Figure 3As shown, the event information generation model mainly consists of two parts: an encoder and a decoder. The encoder includes a first feature extraction subnetwork, a second feature extraction subnetwork, and a feature fusion subnetwork. The decoder includes an information generation subnetwork. Figure 4 This is a flowchart of a method for generating target event information according to an embodiment of this application. The system utilizes an event information generation model through... Figure 4 The following steps are shown to generate target event information:
[0104] Step S402: Obtain user instructions and acquire video data within the target area based on user instructions. The system automatically acquires video data from multiple high-resolution cameras deployed in the target area according to the received instructions from the user.
[0105] Step S404: The user instruction is input into the first feature extraction sub-network (such as a BERT-based text encoder). The first feature extraction sub-network performs a first feature extraction operation on the user instruction and outputs a first feature sequence. The first feature extraction operation is used to perform feature extraction operations on the user instruction. The first feature extraction sub-network first performs word segmentation on the user instruction that has undergone preprocessing to obtain an instruction sequence, and then performs vector transformation operation on the instruction sequence to obtain the first feature sequence.
[0106] Step S406: Input the video data into the second feature extraction sub-network (such as the Swing Transformer+Mamba module), and the second feature extraction sub-network performs the second feature extraction operation on the video data and outputs the second feature sequence. The second feature extraction operation is used to perform feature extraction operation on the video data, and the second feature sequence includes vector representations of image features of multiple frames in the video data.
[0107] Step S408: Input the first feature sequence and the second feature sequence into the feature fusion sub-network, and the feature fusion sub-network performs feature fusion operation on the first feature sequence and the second feature sequence to output the target feature sequence;
[0108] Step S410: Input the target feature sequence into the information generation sub-network. The information generation sub-network generates target event information based on the target feature sequence. The target event information may include event type (also known as event label), specific time, location and detailed attributes, such as "2023-10-10 14:23:45, urban center area, a blue car rear-ended a red van, involving pedestrian injuries".
[0109] The execution order of steps S404 and S406 can be interchanged; that is, step S404 can be executed first and then step S406, or step S406 can be executed first and then step S404.
[0110] In deploying the event information generation model on edge devices, a refined model lightweighting strategy was adopted. Quantization techniques were used to reduce the computational and storage overhead of the model, and innovative activation channel group smoothing calibration (ACGS) and post-quantization output calibration steps were incorporated to minimize the accuracy loss of the quantized model and ensure that the lightweight event information generation model maintains high performance at the edge. For each layer in each sub-network of the event information generation model, the input channels are divided into G groups, and the activation statistics for each group (the maximum value across calibration samples and spatial dimensions) are calculated. Then, for each group g, an optimal alpha_g value is searched to minimize the difference between the contribution of that channel group to the layer output and the FP32 output after applying the smoothing transformation of that group. After ACGS and smoothing quantization, for each layer, the quantized output Y_q and FP32 output Y_fp are calculated using calibration data. The optimal beta_c and bias_c parameters are learned using the least squares method, such that ||Y|| = ||Y||. fp -(Y qbeta +bias)|| 2 Minimum. This deployment strategy takes into account the heterogeneity of activation distribution in different channel groups, making smoothing more refined and targeted. The alpha parameter is no longer layer-wide but channel-local, which can more flexibly cope with activation differences in different attention heads or different parts in complex models (such as Transformer). This makes the event information generation model more suitable for deployment on edge devices, effectively compensating for accumulated quantization errors and improving the accuracy of the event information generation model.
[0111] Through the above steps, the system can immediately process real-time video data using an event information generation model on the edge device according to user instructions. This model can not only detect traffic accidents but also identify specific details of the event, such as the type and color of the vehicles involved, and whether pedestrians are involved, thereby generating accurate target event information. By deploying the event information generation model at the edge, low-latency event detection and information generation are achieved, improving the system's real-time response capability. Edge-processed data only sends key event information to the cloud, reducing data transmission volume, avoiding the uploading of raw video data, and enhancing privacy protection. The event information generation model uses user instructions to focus on target areas and event types, improving the accuracy and relevance of event detection.
[0112] It should be noted that, through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0113] This embodiment also provides an event information generation apparatus, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0114] Figure 5 This is a structural block diagram of an event information generation apparatus according to an embodiment of this application, such as... Figure 5 As shown, the device includes a first memory 52, a first processor 54, and a first computer program 5202 stored in the first memory 52 and executable on the first processor 54. When the first processor 54 executes the first computer program 5202, it performs the following operations: acquiring user instructions and acquiring video data within a target area based on the user instructions; using a pre-trained target model to perform an event information generation operation on the user instructions and the video data within the target area to obtain target event information; wherein the generation operation includes: determining a first feature sequence using the user instructions, the first feature sequence including a vector representation of the text features of the user instructions; determining a second feature sequence using the video data, the second feature sequence including a vector representation of the image features of multiple frames in the video data; fusing the first feature sequence and the second feature sequence to obtain a target feature sequence; and generating the target event information based on the target feature sequence.
[0115] The target model is obtained by iteratively learning and training an initial model using a set of sample information. When the first processor 54 executes the first computer program 5202, it also performs the following operations: acquiring sample information, wherein the sample information includes sample user instructions and sample video data corresponding to the sample user instructions, and the sample video data includes sample labels, which are used to represent the sample event category corresponding to the sample video data; inputting the sample information into the initial model, and having the initial model generate sample labels based on the sample information; calculating the difference between the sample labels and the actual labels using a loss function to obtain the loss value of the initial model, wherein the actual labels are used to represent the actual event category of the sample video data; updating the parameters of the initial model based on the loss value to obtain the target model, wherein the loss value of the target model satisfies a preset threshold.
[0116] When the first processor 54 executes the first computer program 5202, it also performs the following operations: performs a first feature extraction operation on the sample user instruction and the sample video data respectively, and outputs a first sample feature sequence and a second sample feature sequence, wherein the first sample feature sequence includes a vector representation of the text features of the sample user instruction, and the second sample feature sequence includes a vector representation of the image features of multiple frames of sample images in the sample video data; performs a feature fusion operation on the first sample feature sequence and the second sample feature sequence to obtain a target sample feature sequence; and generates the sample label based on the target sample feature sequence.
[0117] When the first processor 54 executes the first computer program 5202, it also performs the following operations: deleting abnormal characters in the sample user instructions to obtain the first sample instructions; performing word segmentation on the instruction text in the first sample instructions to obtain the first sample instruction sequence; and converting multiple word segments in the first sample instruction sequence into vector representations to obtain the first sample feature sequence.
[0118] When the first processor 54 executes the first computer program 5202, it also performs the following operations: performs a serialization operation on the sample video data to obtain a first sample video sequence including the vector representation of each frame of sample image in the sample video data; aggregates the spatial information of the vector representation of each frame of sample image to obtain a second sample video sequence; parses the second sample video sequence to obtain the motion trajectory and sample events of the sample objects, and generates the second sample feature sequence based on the vector representation corresponding to the motion trajectory and the vector representation corresponding to the sample events.
[0119] When the first processor 54 executes the first computer program 5202, it also performs the following operations: extracting M sample images from the sample video data, where M is a natural number greater than 1; dividing the M sample images into P sample image blocks, where P is a natural number greater than M; performing vector transformation on the P sample image blocks to obtain P sample image block vectors; and sorting the P sample image block vectors according to a preset sorting method to obtain the first sample video sequence.
[0120] When the first processor 54 executes the first computer program 5202, it also performs the following operations: converting the elements in the first sample feature sequence and the elements in the second sample feature sequence into a preset format to obtain a third sample feature sequence and a fourth sample feature sequence; performing at least one of the following fusion operations on the third sample feature sequence and the fourth sample feature sequence to obtain the target sample feature sequence: concatenating the fourth sample feature sequence to the third sample feature sequence, concatenating the third sample feature sequence to the fourth sample feature sequence, and concatenating the third sample feature sequence and the fourth sample feature sequence element by element.
[0121] When the first processor 54 executes the first computer program 5202, it also performs the following operations: generating a fifth sample feature sequence using the target sample feature sequence, wherein the fifth sample feature sequence includes multiple sample words, which are used to represent event features in the sample video data; and generating the sample label based on the fifth sample feature sequence.
[0122] When the first processor 54 executes the first computer program 5202, it also performs the following operations: selecting a set of sample words from the sixth sample feature sequence and the target sample feature sequence, wherein the sixth sample feature sequence is a sequence including historical words generated based on the initial model during the previous iteration of training; selecting multiple sample words from the set of sample words to generate the fifth sample feature sequence.
[0123] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when run.
[0124] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0125] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0126] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0127] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0128] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0129] The embodiments described herein also provide a computer program that includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the above method embodiments.
[0130] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0131] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0132] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A method for generating event information, characterized in that, include: Obtain user instructions, and acquire video data within the target area based on the user instructions; Using a pre-trained target model, the event information within the target area is generated based on the user command and the video data to obtain target event information; The generation operation includes: determining a first feature sequence using the user instruction, wherein the first feature sequence includes a vector representation of the text features of the user instruction; determining a second feature sequence using the video data, wherein the second feature sequence includes a vector representation of the image features of multiple frames in the video data; fusing the first feature sequence and the second feature sequence to obtain a target feature sequence; and generating the target event information based on the target feature sequence.
2. The method according to claim 1, characterized in that, The target model is obtained by iteratively training an initial model using a set of sample information. This iterative training includes: Obtain sample information, wherein the sample information includes sample user instructions and sample video data corresponding to the sample user instructions, and the sample video data includes sample tags, the sample tags being used to represent the sample event category corresponding to the sample video data; The sample information is input into the initial model, and the initial model generates sample labels based on the sample information; The difference between the sample labels and the actual labels is calculated using a loss function to obtain the loss value of the initial model, wherein the actual labels are used to represent the actual event categories of the sample video data; The parameters of the initial model are updated based on the loss value to obtain the target model, wherein the loss value of the target model satisfies a preset threshold.
3. The method according to claim 2, characterized in that, The sample information is input into the initial model, and the initial model generates sample labels based on the sample information, including: The first feature extraction operation is performed on the sample user command and the sample video data respectively, and the first sample feature sequence and the second sample feature sequence are output. The first sample feature sequence includes a vector representation of the text features of the sample user command, and the second sample feature sequence includes a vector representation of the image features of multiple frames of sample images in the sample video data. Perform a feature fusion operation on the first sample feature sequence and the second sample feature sequence to obtain the target sample feature sequence; The sample label is generated based on the feature sequence of the target sample.
4. The method according to claim 3, characterized in that, The first feature extraction operation includes: Delete the abnormal characters in the sample user command to obtain the first sample command; The instruction text in the first sample instruction is segmented into words to obtain the first sample instruction sequence; The first sample feature sequence is obtained by converting multiple word segments in the first sample instruction sequence into vector representations.
5. The method according to claim 3, characterized in that, The first feature extraction operation further includes: Perform a serialization operation on the sample video data to obtain a first sample video sequence that includes vector representations of each frame of sample images in the sample video data; By aggregating the spatial information of the vector representation of each sample image frame, a second sample video sequence is obtained; The second sample video sequence is parsed to obtain the motion trajectory and sample events of the sample objects, and the second sample feature sequence is generated based on the vector representations corresponding to the motion trajectory and the vector representations corresponding to the sample events.
6. The method according to claim 5, characterized in that, Performing a serialization operation on the sample video data yields a first sample video sequence comprising vector representations of each frame of sample images in the sample video data, including: M sample images are extracted from the sample video data, where M is a natural number greater than 1; The M frames of sample images are divided into P sample image blocks, where P is a natural number greater than M; Perform vector transformation on the P sample image blocks to obtain P sample image block vectors; The P sample image block vectors are sorted according to a preset sorting method to obtain the first sample video sequence.
7. The method according to claim 3, characterized in that, Perform a feature fusion operation on the first sample feature sequence and the second sample feature sequence to obtain the target sample feature sequence, including: The elements in the first sample feature sequence and the elements in the second sample feature sequence are converted into a preset format to obtain the third sample feature sequence and the fourth sample feature sequence; Perform at least one of the following fusion operations on the third sample feature sequence and the fourth sample feature sequence to obtain the target sample feature sequence: concatenate the fourth sample feature sequence to the third sample feature sequence, concatenate the third sample feature sequence to the fourth sample feature sequence, and concatenate the third sample feature sequence and the fourth sample feature sequence element by element.
8. The method according to claim 3, characterized in that, Generating the sample label based on the target sample feature sequence includes: A fifth sample feature sequence is generated using the target sample feature sequence, wherein the fifth sample feature sequence includes multiple sample words, which are used to represent event features in the sample video data; The sample label is generated based on the feature sequence of the fifth sample.
9. The method according to claim 8, characterized in that, Generating a fifth sample feature sequence using the target sample feature sequence includes: A set of sample words is selected from the sixth sample feature sequence and the target sample feature sequence, wherein the sixth sample feature sequence is a sequence including historical words generated by the initial model during the previous iteration of training; Multiple sample words are selected from the set of sample words to generate the fifth sample feature sequence.
10. An event information generation device, characterized in that, It includes a first memory, a first processor, and a first computer program stored in the first memory and executable on the first processor. When the first processor executes the first computer program, it performs the following operations: Obtain user instructions, and acquire video data within the target area based on the user instructions; Using a pre-trained target model, the event information within the target area is generated based on the user command and the video data to obtain target event information; The generation operation includes: determining a first feature sequence using the user instruction, wherein the first feature sequence includes a vector representation of the text features of the user instruction; determining a second feature sequence using the video data, wherein the second feature sequence includes a vector representation of the image features of multiple frames in the video data; fusing the first feature sequence and the second feature sequence to obtain a target feature sequence; and generating the target event information based on the target feature sequence.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of claims 1 to 9.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Attribute information generation method and device, storage medium, equipment and program product
CN118470719A
Multi-modal model-based intention recognition method, apparatus and device, and storage medium
CN118916531A