A method and device for unsupervised event extraction for spatial operations
By generating semantic relationship diagrams of video sequences and changing states of cluster topology structures, the event extraction network is trained, and the problem of event extraction in videos in the prior art is solved, unsupervised video event extraction is realized, and the ability to understand videos is improved.
Patent Information
- Application Number
- CN202311345148.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-10-17
AI Technical Summary
The prior art is difficult to unsupervisedly extract events in space operations from videos. Event extraction is mainly limited to the text category and lacks effective methods in the field of video comprehension.
By obtaining multiple sample video sequences for spatial operations, a semantic relationship diagram sequence of each video sequence is generated, topological structure change states are determined, and these states are clustered to generate training sample pairs for training event extraction networks, thereby realizing the extraction of unsupervised events.
It realizes unsupervised extraction of events from video sequences, breaks through the limitations of text extraction, and can more comprehensively and accurately understand and analyze event information in videos.
Smart Images

Figure CN117315546B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to an unsupervised event extraction method and device for spatial operations. Background Art
[0002] Space operations are the on-orbit activities that space robots perform to complete prescribed actions or tasks in space, such as on-orbit maintenance and assembly. Events refer to specific events involving one or more participants that occur at a specific time and place, and are the main way to describe changes in environmental states. Therefore, how robots autonomously identify the occurrence of various events is a key technology in space scene understanding.
[0003] Since events are abstract and complex concepts, event extraction is currently limited to the text category, using keyword matching. In the field of video understanding, how to extract unsupervised events from videos has become an urgent problem to be solved. Summary of the invention
[0004] The embodiments of the present invention provide a method and device for unsupervised event extraction oriented to spatial operation, which can realize the extraction of unsupervised events from a video.
[0005] In a first aspect, an embodiment of the present invention provides an unsupervised event extraction method for spatial operations, comprising:
[0006] Acquire multiple sample video sequences for spatial operations, each of which has the same time length;
[0007] For each sample video sequence, the following steps are performed: generating multiple semantic relationship graphs corresponding to multiple frame images in the sample video sequence one by one to obtain a semantic relationship graph sequence of the sample video sequence; the semantic relationship graph includes a topological structure formed by multiple objects in the corresponding image; and determining a topological structure change state corresponding to the sample video sequence based on the semantic relationship graph sequence of the sample video sequence;
[0008] Clustering the topological structure change states respectively corresponding to a plurality of sample video sequences, and obtaining a plurality of training sample pairs according to the clustering results; wherein the training sample pairs include the topological structure change state as input and the event category of the class to which the topological structure change state belongs as output;
[0009] The event extraction network is trained using the multiple training samples, so as to extract events using the trained event extraction network.
[0010] In a possible design, the acquiring of a plurality of sample video sequences for spatial operation includes:
[0011] Obtaining a temporally continuous initial video sequence for spatial operations;
[0012] The initial video sequence is divided into a plurality of sample video sequences of the same time length; the plurality of sample video sequences are continuous.
[0013] In a possible design, the event extraction using the trained event extraction network includes:
[0014] Acquire a video sequence to be processed; the time length of the video sequence to be processed is the same as the time length of the sample video sequence;
[0015] Generate semantic relationship graphs one by one for multiple frames of images in the video sequence to be processed, and obtain a semantic relationship graph sequence of the video sequence to be processed;
[0016] Based on the semantic relationship graph sequence of the video sequence to be processed, determining a topological structure change state corresponding to the video sequence to be processed;
[0017] The topological structure change state corresponding to the video sequence to be processed is input into the event extraction network to output the event category of the video sequence to be processed.
[0018] In a possible design, the semantic relationship graph is generated in the following manner:
[0019] Use the target detection and recognition algorithm to identify the object category of each object in the current frame image;
[0020] Calculate the pose of each object in the current frame image using a pose estimation algorithm;
[0021] Determine the relative position between any two objects based on the position of each object;
[0022] The objects in the current frame image are taken as nodes, and the topological structure formed by connecting the nodes in pairs is taken as the semantic relationship graph of the current frame image; the edges in the topological structure are the relative postures between the corresponding two objects, and the nodes in the topological structure are marked with corresponding object categories.
[0023] In a possible design, a method for determining the topology structure change state includes:
[0024] Determine, according to the topological structure of each semantic relationship graph in the semantic relationship graph sequence, a static scene feature of each semantic relationship graph, wherein the static scene feature includes at least one of the number of objects, the category of objects, and the relative position between any two objects;
[0025] According to the time sequence in the semantic relationship graph sequence, the change characteristics of the static scene features in the spatiotemporal relationship are determined to obtain the change state of the topological structure.
[0026] In one possible design, the topological structure change state is represented by a spatiotemporal feature vector extracted by a spatiotemporal graph convolutional neural network.
[0027] In a possible design, clustering the topological structure change states corresponding to the plurality of sample video sequences respectively includes:
[0028] Based on the spatiotemporal feature vectors of multiple sample video sequences, similar spatiotemporal feature vectors are clustered into one category.
[0029] In a second aspect, an embodiment of the present invention further provides an unsupervised event extraction device for spatial operations, comprising:
[0030] An acquisition module, used for acquiring a plurality of sample video sequences for spatial operation, each of which has the same time length;
[0031] The processing module is used to perform, for each sample video sequence, the following steps: generating multiple semantic relationship graphs corresponding to multiple frames of images in the sample video sequence, to obtain a semantic relationship graph sequence of the sample video sequence; the semantic relationship graph includes a topological structure formed by multiple objects in the corresponding image; and determining a topological structure change state corresponding to the sample video sequence based on the semantic relationship graph sequence of the sample video sequence;
[0032] A clustering module clusters the topological structure change states corresponding to the plurality of sample video sequences, and obtains a plurality of training sample pairs according to the clustering results; wherein the training sample pairs include the topological structure change state as input and the event category of the class to which the topological structure change state belongs as output;
[0033] A training module, used for training the event extraction network using the multiple training samples;
[0034] The extraction module is used to extract events using the trained event extraction network.
[0035] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method described in any embodiment of this specification is implemented.
[0036] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, enables the computer to execute the method described in any embodiment of this specification.
[0037] The embodiment of the present invention provides an unsupervised event extraction method and device for spatial operations, which generates a semantic relationship graph for multiple frame images in a sample video sequence in a one-to-one correspondence. Since the semantic relationship graph includes a topological structure formed by multiple objects in the corresponding image, the topological structure change state can be determined for the semantic relationship graph sequence corresponding to the sample video sequence, and the occurrence of an event can be regarded as a change in the topological structure. By clustering multiple topological structure change states, the topological structure change states in the same class can be regarded as the same event. Therefore, clustering can be used to obtain multiple training sample pairs in an unsupervised manner. The multiple training sample pairs can be used to train the event extraction network, and then the trained event extraction network can be used to extract unsupervised events from the video sequence. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0039] Figure 1 It is a flow chart of an unsupervised event extraction method for spatial operation provided by one embodiment of the present invention;
[0040] Figure 2 is a semantic relationship diagram provided by an embodiment of the present invention;
[0041] Figure 3 This is a hardware architecture diagram of an electronic device provided by an embodiment of the present invention;
[0042] Figure 4 It is a structural diagram of an unsupervised event extraction device for spatial operations provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0044] Please refer to Figure 1 The embodiment of the present invention provides an unsupervised event extraction method for spatial operation, the method comprising:
[0045] Step 100, obtaining multiple sample video sequences for spatial operation, each sample video sequence having the same time length;
[0046] Step 102, for each sample video sequence, the following steps are performed: generating multiple semantic relationship graphs corresponding to multiple frames of images in the sample video sequence, to obtain a semantic relationship graph sequence of the sample video sequence; the semantic relationship graph includes a topological structure formed by multiple objects in the corresponding image; and determining a topological structure change state corresponding to the sample video sequence based on the semantic relationship graph sequence of the sample video sequence;
[0047] Step 104, clustering the topological structure change states corresponding to the plurality of sample video sequences, and obtaining a plurality of training sample pairs according to the clustering results; wherein the training sample pairs include the topological structure change state as input and the event category of the class to which the topological structure change state belongs as output;
[0048] Step 106: train the event extraction network using the multiple training samples, so as to extract events using the trained event extraction network.
[0049] The embodiment of the present invention provides an unsupervised event extraction method and device for spatial operations, which generates a semantic relationship graph for multiple frame images in a sample video sequence in a one-to-one correspondence. Since the semantic relationship graph includes a topological structure formed by multiple objects in the corresponding image, the topological structure change state can be determined for the semantic relationship graph sequence corresponding to the sample video sequence, and the occurrence of an event can be regarded as a change in the topological structure. By clustering multiple topological structure change states, the topological structure change states in the same class can be regarded as the same event. Therefore, clustering can be used to obtain multiple training sample pairs in an unsupervised manner. The multiple training sample pairs can be used to train the event extraction network, and then the trained event extraction network can be used to extract unsupervised events from the video sequence.
[0050] Described below Figure 1 How the various steps are performed.
[0051] First, with respect to step 100, a plurality of sample video sequences for spatial operation are obtained, and each sample video sequence has the same time length.
[0052] When a space robot performs space operations in a space environment, it is always faced with a complex environment that is constantly changing. In order to improve the autonomy of the space robot and to be able to take corresponding response strategies for environmental changes, the space robot needs to be able to understand and identify changes in environmental states, and environmental state changes can be described by events. In an embodiment of the present invention, the space operation process can be regarded as the changes in the relative positions of various objects in the space environment generated by the space robot.
[0053] In order to obtain the changes in the relative positions of various objects in the space environment, it is first necessary to obtain multiple sample video sequences for space operations. In order to ensure the uniformity and accuracy of subsequent event analysis, the time length of the multiple sample video sequences is the same.
[0054] In one embodiment of the present invention, the method for acquiring multiple sample video sequences can be to acquire video clips for different spatial operation types, or to acquire video clips of the same length for different time periods of spatial operations, and determine the sample video sequence based on the video clips. In one implementation, all video frame images contained in the video clips can be directly used as sample video sequences; in another implementation, video frame images can be sampled at the same sampling interval, and the sampled video frame images can be used as sample video sequences. The number of video frame images in multiple sample video sequences is the same.
[0055] In another embodiment of the present invention, taking into account that the space robot will repeat the same execution action when performing space operation, in order to improve the relevant characteristics of the same event, the multiple sample video sequences are obtained in the following way: obtain a time-continuous initial video sequence for space operation; divide the initial video sequence into multiple sample video sequences with the same time length; the multiple sample video sequences are continuous.
[0056] For example, the time length of a time-continuous initial video sequence is T, so the initial video sequence corresponding to the time length of T is divided. Assuming that the time length of the division is ΔT, k=T / ΔT video segments can be obtained. The k video segments are used as k sample video sequences. The k sample video sequences can be connected end to end in sequence to obtain a continuous video sequence.
[0057] In the embodiment of the present invention, since the multiple sample video sequences are obtained by segmenting the temporally continuous initial video sequence, the uniformity of the multiple sample video sequences in spatial density can be ensured, and key information can be ensured not to be lost, further ensuring the integrity of event-related features.
[0058] Next, for step 102, for each sample video sequence, the following are performed: generating multiple semantic relationship graphs corresponding to multiple frame images in the sample video sequence one by one to obtain a semantic relationship graph sequence of the sample video sequence; the semantic relationship graph includes a topological structure formed by multiple objects in the corresponding image; based on the semantic relationship graph sequence of the sample video sequence, determining the topological structure change state corresponding to the sample video sequence.
[0059] In an embodiment of the present invention, the spatial operation process needs to be regarded as the change of the relative posture of each object in the spatial environment generated by the action of the space robot, and the sample video sequence includes multiple frames of images, each frame of the image includes the relative posture of the object, so that the change of the relative posture of the object can be obtained based on the multiple frames of images. Among them, the topological structure formed by multiple objects in the semantic relationship graph can be used to represent the relative posture of each object in the spatial environment. By generating multiple semantic relationship graphs for the multiple frames of images in the video sequence in a one-to-one correspondence, a sequence of semantic relationship graphs of the video sequence can be obtained. For example, a video sequence includes image 1, image 2, ..., image m (m is an integer not less than 2), then the sequence of semantic relationship graphs may include semantic relationship Figure 1 , semantic relationship Figure 2 ,……,semantic relationship graph m.
[0060] Specifically, the generation method of the semantic relationship graph can be implemented by at least the following steps A1-A4:
[0061] A1. Use the target detection and recognition algorithm to identify the object category of each object in the current frame image;
[0062] A2. Calculate the pose of each object in the current frame image using a pose estimation algorithm;
[0063] A3. Determine the relative position between any two objects based on the position of each object;
[0064] A4. The objects in the current frame image are taken as nodes, and the topological structure formed by connecting the nodes in pairs is taken as the semantic relationship graph of the current frame image; the edges in the topological structure are the relative postures between the corresponding two objects, and the nodes in the topological structure are marked with corresponding object categories.
[0065] In the embodiment of the present invention, the target detection algorithm can be implemented by YOLO, MaskRCNN, DeepLab and other algorithms. Specifically, the number of objects in the current frame image and the object category of each object can be identified.
[0066] In the embodiment of the present invention, the pose estimation algorithm can at least be implemented by PoseCNN.
[0067] For example, suppose there are three objects identified in the image, and the object categories are nozzle, filling gun, and filling port. Each object is used as a node, and the nodes are connected in pairs to form a topological structure. Please refer to Figure 2 , which is a schematic diagram of a semantic relationship graph. Since the edges in the topological structure are the relative positions between the corresponding two objects, the topological structure can represent the relative positions of the objects in the image.
[0068] Furthermore, after obtaining the semantic relationship graph, in order to be able to perform event representation, it is necessary to determine the topological structure change state corresponding to the sample video sequence based on the semantic relationship graph sequence of the sample video sequence, so as to use the topological structure change state to represent the corresponding event.
[0069] Specifically, the method for determining the topology structure change state may include:
[0070] Determine, according to the topological structure of each semantic relationship graph in the semantic relationship graph sequence, a static scene feature of each semantic relationship graph, wherein the static scene feature includes at least one of the number of objects, the category of objects, and the relative position between any two objects;
[0071] According to the time sequence in the semantic relationship graph sequence, the change characteristics of the static scene features in the spatiotemporal relationship are determined to obtain the change state of the topological structure.
[0072] Among them, the static scene features of each semantic relationship graph can be extracted through the Gconv network. The static scene features include at least one of the number of objects, the category of objects, and the relative posture between any two objects. When extracting static scene features, the attention module of the corresponding static scene features can be used to perform attention enhancement processing on the attention features. For example, an attention module for object category and an attention module for relative posture are added to the Gconv network to perform attention enhancement processing on the object category and relative posture respectively, so as to focus on the object category and relative posture features. Preferably, the static scene features include the number of objects, the category of objects, and the relative posture between any two objects.
[0073] Since multiple semantic relationship graphs exist in a time sequence, the CNN network can be used to determine the changing characteristics of static scene features in the spatiotemporal relationship, that is, the topological structure change state corresponding to the video sequence can be obtained.
[0074] The above topological structure change state can be obtained by processing the semantic relationship graph sequence using a spatiotemporal graph convolutional neural network. For example, the spatiotemporal graph convolutional neural network is a STGNNs network, which can simultaneously capture the spatial and temporal dependencies of an image and realize feature extraction through compression. It can be seen that the topological structure change state can be represented by the spatiotemporal feature vector extracted by the spatiotemporal graph convolutional neural network.
[0075] Then, for step 104, the topological structure change states corresponding to the multiple sample video sequences are clustered, and multiple training sample pairs are obtained according to the clustering results; wherein the training sample pairs include the topological structure change state as input and the event category of the class to which the topological structure change state belongs as output.
[0076] In the embodiment of the present invention, the topological structure change states corresponding to the multiple sample video sequences are clustered to obtain multiple classes, each class contains a set of similar topological structure change states, and each class corresponds to an event category. According to the clustering results, multiple training sample pairs can be generated. These training sample pairs are obtained based on an unsupervised method, and these training sample pairs can be used to train the model of the event extraction network. By performing deep learning of the features of the topological structure change states between the sample video sequences, the event extraction network can be used to quickly realize event extraction.
[0077] In one implementation, when clustering the topological structure change states corresponding to the multiple sample video sequences, the other topological structure change states can be traversed for each current topological structure change state to determine whether there are other topological structure change states similar to the current topological structure change state. In an embodiment of the present invention, it can be determined whether two topological structure change states are similar according to a preset strategy. For example, the preset strategy is that the difference in relative posture between the same two objects in the corresponding two frames of images does not exceed a set threshold.
[0078] For example, if two sample video sequences each include two frames of images, then in the first frame of the corresponding two frames, the objects are object 1, object 2, and object 3, the relative pose of object 1 and object 2 in the first frame of the first sample video sequence is P11, and the relative pose of object 1 and object 2 in the first frame of the second sample video sequence is P21. The relative pose difference between object 1 and object 2 is P1. Similarly, the relative pose difference between object 1 and object 3 is P2, and the relative pose difference between object 2 and object 3 is P3. If P1, P2, and P3 do not exceed the set threshold, it is determined that the topological structure change states corresponding to the two sample video sequences are similar.
[0079] In another implementation, when the topological structure change state is represented by a spatiotemporal feature vector extracted by a spatiotemporal graph convolutional neural network, clustering the topological structure change states corresponding to a plurality of sample video sequences may include clustering similar spatiotemporal feature vectors into one category based on the spatiotemporal feature vectors of the plurality of sample video sequences using a clustering algorithm. The clustering algorithm may at least be implemented by K-means.
[0080] For example, assuming that there are k space-time feature vectors corresponding to k sample video sequences, the k space-time feature vectors are clustered using a clustering algorithm. Specifically, the similarity of the space-time feature vectors is used to implement the entire clustering process. The similarity can be calculated using calculation formulas such as Euclidean distance and Manhattan distance.
[0081] In the embodiment of the present invention, by clustering k spatiotemporal feature vectors using a clustering algorithm, a clustering result can be quickly obtained, and then a sample set can be obtained in an unsupervised manner.
[0082] Regardless of which of the above implementation methods is used, after the multiple topological structure change states corresponding to the multiple sample video sequences are divided into multiple classes, each class includes at least one topological structure change state, each class can be regarded as an event category, and each topological structure change state included in the class can be characterized as the event category corresponding to the class.
[0083] Furthermore, if k sample video sequences are obtained by segmenting a time-continuous initial video sequence, then based on the time order of the k sample video sequences, the event categories of the classes to which the k topological structure change states corresponding to the k sample video sequences belong are determined, and the event development process of the initial video sequence can be obtained.
[0084] In the embodiment of the present invention, after obtaining the clustering result, a sample set of the event extraction network can be obtained, and the sample set includes multiple training sample pairs. If the number of sample video sequences is k, the number of training sample pairs can also be k. Specifically, the topological structure change state corresponding to each sample video sequence is used as input, and the event category of the class to which the topological structure change state belongs is used as output, so that k training sample pairs can be obtained.
[0085] For example, the topological structure change states of k sample video sequences are divided into 7 categories, corresponding to event category 1, event category 2... event category 7; then a topological structure change state can be selected from the first class without replacement as input, event category 1 as output, to obtain a training sample pair, and continue to select from the first class without replacement until all the topological structure change states in the first class are selected, to obtain multiple training sample pairs; then, a topological structure change state can be selected from the second class without replacement as input, event category 2 as output, and continue to select from the second class without replacement until all the topological structure change states in the second class are selected, to obtain multiple training sample pairs; .... In this way, k training sample pairs can be obtained.
[0086] Among them, each training sample pair can be formalized as an ordered pair <V ij ,e i >, V ij is the jth topological structure change state in the i-th class, e i is the event category corresponding to the i-th class. Among them, the topological structure change state can be characterized by a spatiotemporal feature vector. In an embodiment of the present invention, the event category can be characterized by one-hot encoding. Assuming that there are 7 event categories, the one-hot encoding of event category e1 can be [1, 0, 0, 0, 0, 0], and the i-th bit of the one-hot encoding of event category ei (i = 1, 2, ..., 7) is 1, and the other bits are 0.
[0087] Finally, with respect to step 106, the event extraction network is trained using the plurality of training samples, so as to perform event extraction using the trained event extraction network.
[0088] In an embodiment of the present invention, the multiple training samples obtained in the above steps can be used to train the event extraction network, wherein the structure of the event extraction network can be designed by using a fully connected MLP network or an attention network.
[0089] After the event extraction network training is completed, the method of using the trained event extraction network to extract events may include the following steps S1-S4:
[0090] S1. Obtain a video sequence to be processed; the time length of the video sequence to be processed is the same as the time length of the sample video sequence;
[0091] S2, generating a semantic relationship graph for multiple frames of images in the video sequence to be processed in a one-to-one correspondence, and obtaining a semantic relationship graph sequence of the video sequence to be processed;
[0092] S3, based on the semantic relationship graph sequence of the video sequence to be processed, determining the topological structure change state corresponding to the video sequence to be processed;
[0093] S4. Inputting the topological structure change state corresponding to the video sequence to be processed into the event extraction network to output the event category of the video sequence to be processed.
[0094] In an embodiment of the present invention, the method of generating a semantic relationship graph in the above step S2 is the same as the method of generating a semantic relationship graph for a video frame image in a sample video sequence in step 102, and the method of determining a topological structure change state corresponding to a video sequence to be processed in the above step S3 is the same as the method of determining a topological structure change state corresponding to a sample video sequence in step 102. For details, please refer to the description of the above embodiment, which will not be repeated in this embodiment. In step S4, the topological structure change state of the video sequence to be processed is input into the event extraction network, and the event extraction network can use the learned association between the topological structure change state and the event category to output the event category corresponding to the video sequence to be processed.
[0095] In an embodiment of the present invention, the topological structure change state corresponding to the sample video sequence is paired with the corresponding event category. By training the sample pairs, the event extraction network can learn the relationship between the spatiotemporal feature vector and its corresponding event category, and further extract event-related high-dimensional information from the video sequence, rather than just being limited to text information. The event extraction network obtained through training can map multiple sample video sequences to corresponding event categories to achieve event extraction in the field of video understanding. This innovation breaks through the limitation that the current event extraction method is limited to text, and brings broader application prospects to the fields of video understanding and content understanding. By associating video change data with event categories, event information in the video can be understood and analyzed more comprehensively and accurately.
[0096] like Figure 3 , Figure 4 As shown, the embodiment of the present invention provides an unsupervised event extraction device for spatial operations. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. From the hardware level, Figure 3 As shown, a hardware architecture diagram of an electronic device in which an unsupervised event extraction device for space operation provided by an embodiment of the present invention is located, except Figure 3 In addition to the processor, memory, network interface, and non-volatile memory shown, the electronic device in the embodiment may also include other hardware, such as a forwarding chip responsible for processing messages, etc. Taking software implementation as an example, Figure 4As shown, as a device in a logical sense, the CPU of the electronic device in which it is located reads the corresponding computer program in the non-volatile memory into the memory and runs it. This embodiment provides an unsupervised event extraction device for spatial operation, including:
[0097] An acquisition module 400 is used to acquire a plurality of sample video sequences for spatial operation, each of which has the same time length;
[0098] The processing module 402 is used to perform, for each sample video sequence, the following steps: generating multiple semantic relationship graphs corresponding to multiple frames of images in the sample video sequence, to obtain a semantic relationship graph sequence of the sample video sequence; the semantic relationship graph includes a topological structure formed by multiple objects in the corresponding image; and determining a topological structure change state corresponding to the sample video sequence based on the semantic relationship graph sequence of the sample video sequence;
[0099] A clustering module 404 clusters the topological structure change states corresponding to the plurality of sample video sequences, and obtains a plurality of training sample pairs according to the clustering results; wherein the training sample pairs include the topological structure change state as input and the event category of the class to which the topological structure change state belongs as output;
[0100] A training module 406, configured to train an event extraction network using the plurality of training samples;
[0101] The extraction module 408 is used to extract events using the trained event extraction network.
[0102] In one embodiment of the present invention, the acquisition module is specifically used to: acquire a temporally continuous initial video sequence for spatial operations; divide the initial video sequence into a plurality of sample video sequences of the same temporal length; and the plurality of sample video sequences are continuous.
[0103] In one embodiment of the present invention, the extraction module is specifically used to: obtain a video sequence to be processed; the time length of the video sequence to be processed is the same as the time length of the sample video sequence; generate a semantic relationship graph for multiple frame images in the video sequence to be processed in a one-to-one correspondence to obtain a semantic relationship graph sequence of the video sequence to be processed; based on the semantic relationship graph sequence of the video sequence to be processed, determine the topological structure change state corresponding to the video sequence to be processed; input the topological structure change state corresponding to the video sequence to be processed into the event extraction network to output the event category of the video sequence to be processed.
[0104] In one embodiment of the present invention, the processing module specifically generates a semantic relationship graph in the following manner: using a target detection and recognition algorithm to identify the object category of each object in the current frame image; using a pose estimation algorithm to calculate the pose of each object in the current frame image; determining the relative pose between any two objects based on the pose of each object; using the objects in the current frame image as nodes, and using a topological structure formed by connecting the nodes in pairs as the semantic relationship graph of the current frame image; the edges in the topological structure are the relative poses between the corresponding two objects, and the nodes in the topological structure are marked with corresponding object categories.
[0105] In one embodiment of the present invention, the processing module specifically determines the topological structure change state in the following manner: according to the topological structure of each semantic relationship graph in the semantic relationship graph sequence, the static scene features of each semantic relationship graph are determined, and the static scene features include at least one of the number of objects, object categories, and the relative posture between any two objects; according to the time sequence in the semantic relationship graph sequence, the change characteristics of the static scene features in the spatiotemporal relationship are determined to obtain the topological structure change state.
[0106] In one embodiment of the present invention, the topological structure change state is represented by a spatiotemporal feature vector extracted by a spatiotemporal graph convolutional neural network.
[0107] In one embodiment of the present invention, the clustering module is specifically used to: cluster similar spatiotemporal feature vectors into one category based on the spatiotemporal feature vectors of multiple sample video sequences.
[0108] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on an unsupervised event extraction device for space operations. In other embodiments of the present invention, an unsupervised event extraction device for space operations may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0109] The information interaction, execution process and other contents between the modules in the above-mentioned device are based on the same concept as the embodiment of the method of the present invention. For the specific contents, please refer to the description in the embodiment of the method of the present invention, and no further description is given here.
[0110] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, an unsupervised event extraction method for spatial operations in any embodiment of the present invention is implemented.
[0111] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor executes an unsupervised event extraction method for spatial operations in any embodiment of the present invention.
[0112] Specifically, a system or device equipped with a storage medium can be provided, on which software program code that implements the functions of any of the above-mentioned embodiments is stored, and a computer (or CPU or MPU) of the system or device can be enabled to read and execute the program code stored in the storage medium.
[0113] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.
[0114] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer by a communication network.
[0115] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.
[0116] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or to a memory provided in an expansion module connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or expansion module is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.
[0117] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical factors in the process, method, article or device including the elements.
[0118] A person of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiments; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk, etc., various media that can store program codes.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An unsupervised event extraction method for spatial operations, It is characterized in that include: Acquire multiple sample video sequences for spatial operations, each of which has the same time length; For each sample video sequence, the following steps are performed: generating multiple semantic relationship graphs corresponding to multiple frame images in the sample video sequence one by one to obtain a semantic relationship graph sequence of the sample video sequence; the semantic relationship graph includes a topological structure formed by multiple objects in the corresponding image; and determining a topological structure change state corresponding to the sample video sequence based on the semantic relationship graph sequence of the sample video sequence; Clustering the topological structure change states respectively corresponding to a plurality of sample video sequences, and obtaining a plurality of training sample pairs according to the clustering results; wherein the training sample pairs include the topological structure change state as input and the event category of the class to which the topological structure change state belongs as output; Using the multiple training samples to train an event extraction network, so as to extract events using the trained event extraction network; The method for generating the semantic relationship graph includes: using a target detection and recognition algorithm to identify the object category of each object in the current frame image; using a posture estimation algorithm to calculate the posture of each object in the current frame image; determining the relative posture between any two objects based on the posture of each object; using the objects in the current frame image as nodes, and using a topological structure formed by connecting the nodes in pairs as the semantic relationship graph of the current frame image; the edges in the topological structure are the relative postures between the corresponding two objects, and the nodes in the topological structure are marked with corresponding object categories; The method for determining the topological structure change state includes: determining the static scene features of each semantic relationship graph according to the topological structure of each semantic relationship graph in the semantic relationship graph sequence, wherein the static scene features include at least one of the number of objects, the category of objects, and the relative posture between any two objects; determining the change features of the static scene features in the spatiotemporal relationship according to the time sequence in the semantic relationship graph sequence, and obtaining the topological structure change state.
2. The method according to claim 1, It is characterized in that The step of acquiring a plurality of sample video sequences for spatial operation comprises: Obtaining a temporally continuous initial video sequence for spatial operations; The initial video sequence is divided into a plurality of sample video sequences of the same time length; the plurality of sample video sequences are continuous.
3. The method according to claim 1, It is characterized in that The event extraction using the trained event extraction network includes: Acquire a video sequence to be processed; the time length of the video sequence to be processed is the same as the time length of the sample video sequence; Generate semantic relationship graphs one by one for multiple frames of images in the video sequence to be processed, and obtain a semantic relationship graph sequence of the video sequence to be processed; Based on the semantic relationship graph sequence of the video sequence to be processed, determining a topological structure change state corresponding to the video sequence to be processed; The topological structure change state corresponding to the video sequence to be processed is input into the event extraction network to output the event category of the video sequence to be processed.
4. The method according to claim 1, It is characterized in that The topological structure change state is represented by the spatiotemporal feature vector extracted by the spatiotemporal graph convolutional neural network.
5. The method according to claim 4, It is characterized in that The clustering of topological structure change states respectively corresponding to a plurality of sample video sequences comprises: Based on the spatiotemporal feature vectors of multiple sample video sequences, similar spatiotemporal feature vectors are clustered into one category.
6. An unsupervised event extraction device for spatial operations, It is characterized in that include: An acquisition module, used for acquiring a plurality of sample video sequences for spatial operation, each of which has the same time length; The processing module is used to perform, for each sample video sequence, the following steps: generating multiple semantic relationship graphs corresponding to multiple frames of images in the sample video sequence, to obtain a semantic relationship graph sequence of the sample video sequence; the semantic relationship graph includes a topological structure formed by multiple objects in the corresponding image; and determining a topological structure change state corresponding to the sample video sequence based on the semantic relationship graph sequence of the sample video sequence; A clustering module clusters the topological structure change states corresponding to the plurality of sample video sequences, and obtains a plurality of training sample pairs according to the clustering results; wherein the training sample pairs include the topological structure change state as input and the event category of the class to which the topological structure change state belongs as output; A training module, used for training the event extraction network using the multiple training samples; The extraction module is used to extract events using the trained event extraction network; The method for generating the semantic relationship graph includes: using a target detection and recognition algorithm to identify the object category of each object in the current frame image; using a posture estimation algorithm to calculate the posture of each object in the current frame image; determining the relative posture between any two objects based on the posture of each object; using the objects in the current frame image as nodes, and using a topological structure formed by connecting the nodes in pairs as the semantic relationship graph of the current frame image; the edges in the topological structure are the relative postures between the corresponding two objects, and the nodes in the topological structure are marked with corresponding object categories; The method for determining the topological structure change state includes: determining the static scene features of each semantic relationship graph according to the topological structure of each semantic relationship graph in the semantic relationship graph sequence, wherein the static scene features include at least one of the number of objects, the category of objects, and the relative posture between any two objects; determining the change features of the static scene features in the spatiotemporal relationship according to the time sequence in the semantic relationship graph sequence, and obtaining the topological structure change state.
7. An electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-video event blind area change process deduction method based on geographical semantic association constraints
CN112214642A
Video clustering method and device thereof
CN113515668A