Complex scene understanding method and system based on mixed attention dynamic feedback adjustment

Through the mixed attention mechanism and dynamic feedback adjustment mechanism of self-supervised training, the problem of user personalized recognition in complex scenario understanding is solved, real-time update of object importance scores and personalized service quality improvement is achieved.

CN120472460APending Publication Date: 2025-08-12TSINGHUA UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510483568.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art cannot provide different users with personalized and complex scenario understanding, and it is difficult to identify the importance of objects in the scenario and meet user needs.

Method used

A hybrid attention mechanism of self-supervised training is used to score complex scene semantic graph sequences in multiple dimensions, and the weight ratio is adjusted according to user preference data through a dynamic feedback adjustment mechanism, and the object importance score is updated in real time.

Benefits of technology

It realizes dynamic adjustment of object importance sorting according to user needs, and improves the personalization and quality of downstream application services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472460A_ABST
    Figure CN120472460A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of complex scene understanding, in particular to a complex scene understanding method, system and device based on mixed attention dynamic feedback adjustment and a computer storage medium. The complex scene understanding method comprises the following steps: performing multi-dimensional scoring on each class of objects in a complex scene semantic graph sequence based on a self-supervised training mixed attention mechanism, and performing weighted calculation on a plurality of scores to obtain an importance score of each class of objects; on the basis of a dynamic feedback adjustment mechanism, the weight ratio is adjusted according to preference data fed back by a user online, and the importance score of each class of objects is updated in real time; according to the invention, through a large number of complex scene videos, self-supervised learning is carried out on model parameters based on a mixed attention mechanism. Through personalized interaction feedback, the personalized demand of each user is dynamically modeled, and the quality of downstream application services is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of complex scene understanding technology, and in particular to a complex scene understanding method, system, device and computer storage medium based on hybrid attention dynamic feedback regulation. Background Art

[0002] In applications such as scene rendering and visually impaired perception, as scene complexity increases, there is an urgent need to understand scene semantics, analyze the importance of different objects in complex scenes to different users, and align user needs and values to provide personalized, high-quality services. For example, in complex scene rendering applications, if we can model that users are more interested in historical relics, we can increase the rendering priority of objects related to historical relics, thereby improving the user experience. In complex scene visually impaired perception applications, if we can model that users are more interested in natural scenery, we can prioritize the delivery of information related to natural scenery to visually impaired users, helping them understand the scene semantics.

[0003] Existing mainstream technologies use statistical methods to rank the importance of objects in a scene. Cutting-edge technologies use deep neural network-based methods that rely on large amounts of labeled data pairs to learn the importance of objects in a scene.

[0004] However, in real-world scenarios, object importance labels are difficult to obtain, making supervised learning methods difficult to train. Furthermore, different users have different cognitive patterns regarding objects in a scene, resulting in the same object having different importance for different users. Consequently, traditional technologies are unable to provide personalized recognition for each user. Summary of the Invention

[0005] To this end, the technical problem to be solved by the present invention is to overcome the problem that the complex scene understanding method of the prior art cannot perform personalized recognition for users.

[0006] To solve the above technical problems, the present invention provides a complex scene understanding method, comprising:

[0007] Obtain video frame image sequences of complex scenes;

[0008] Perform object detection and relationship recognition on each video frame to obtain a complex scene semantic graph sequence;

[0009] A hybrid attention mechanism based on self-supervised training is used to perform multi-dimensional scoring on each type of object in the complex scene semantic graph sequence, and multiple scores are weighted to obtain the importance score of each type of object;

[0010] Based on the dynamic feedback adjustment mechanism, the weight ratio is adjusted according to the preference data of users' online feedback, and the importance score of each type of object is updated in real time.

[0011] Preferably, the hybrid attention mechanism based on self-supervised training performs multi-dimensional scoring on each type of object in the complex scene semantic graph sequence, and performs weighted calculation on multiple scores to obtain the importance score of each type of object, including:

[0012] Based on the attractiveness baseline attention mechanism and the background attention mechanism, the attractiveness score of each type of object in the complex scene semantic graph sequence is performed;

[0013] Based on the freshness attention mechanism, a freshness score is given to each type of object in the complex scene semantic graph sequence.

[0014] Preferably, the attractiveness scoring of each type of object in the complex scene semantic graph sequence based on the attractiveness baseline attention mechanism and the background attention mechanism includes:

[0015] Based on the attractiveness benchmark attention mechanism, each type of object in the complex scene semantic graph sequence is scored with an attractiveness benchmark score:

[0016] Obtain a complex scene semantic graph sequence and original frame sequence information arranged in chronological order, and use graph embedding technology to embed the nodes and edges in each complex scene semantic graph into a high-dimensional space, add a position code to each node, and obtain the first target complex scene semantic graph sequence;

[0017] Reconstructing nodes in the next single-frame scene semantic graph based on the single-frame graph in the first target complex scene semantic graph sequence;

[0018] Calculating a first reconstruction accuracy rate for each node, and calculating an attractiveness benchmark score for each node based on the first reconstruction accuracy rate;

[0019] Scoring the background score of each type of object in the complex scene semantic graph sequence based on the background attention mechanism;

[0020] Obtaining an attractiveness score for each node based on the attractiveness benchmark score and the background score;

[0021] Based on the attractiveness score of each node, the attractiveness score of each class of objects is integrated.

[0022] Preferably, the background scoring of each type of object in the complex scene semantic graph sequence based on the background attention mechanism includes:

[0023] Randomly mask a single frame in a complex scene semantic graph sequence and reconstruct all node categories in the masked graph with equal probability;

[0024] The second reconstruction accuracy of each node is calculated, and the background score of each node is calculated according to the second reconstruction accuracy.

[0025] Preferably, the freshness scoring of each type of object in the complex scene semantic graph sequence based on the freshness attention mechanism includes:

[0026] Constructing a temporal relationship representation for the complex scene semantic graph sequence by using a position encoding technology to obtain a second target complex scene semantic graph sequence;

[0027] Reconstructing the nodes in the next single-frame scene semantic graph based on the second target complex scene semantic graph sequence;

[0028] Calculating the third reconstruction accuracy of each node, and calculating the freshness score of each node according to the third reconstruction accuracy;

[0029] According to the freshness score of each node, the freshness score of each type of object is integrated.

[0030] Preferably, the hybrid attention mechanism based on self-supervised training performs multi-dimensional scoring on each type of object in the complex scene semantic graph sequence, and performs weighted calculation on multiple scores to obtain the importance score of each type of object, further comprising:

[0031] According to the degree of user demand for each type of object, a specific demand score is given to each type of object in the complex scene semantic graph sequence.

[0032] Preferably, the dynamic feedback adjustment mechanism based on which the weight ratio is adjusted according to the preference data of the user's online feedback includes:

[0033] According to the preference data of users' online feedback, the weight ratio is adjusted using the maximum likelihood estimation method.

[0034] The present invention also provides a complex scene semantic understanding system, comprising:

[0035] Video frame image acquisition module, used to acquire video frame image sequences of complex scenes;

[0036] The scene semantic graph acquisition module is used to perform target detection and relationship recognition on each video frame to obtain a complex scene semantic graph sequence;

[0037] A multi-dimensional scoring module is used to perform multi-dimensional scoring on each type of object in the complex scene semantic graph sequence based on a hybrid attention mechanism trained through self-supervision, and perform weighted calculation on multiple scores to obtain an importance score for each type of object;

[0038] The importance score determination module is used to adjust the weight ratio based on the preference data of users' online feedback based on the dynamic feedback adjustment mechanism, and update the importance score of each type of object in real time.

[0039] The present invention also provides a complex scene semantic understanding device, comprising:

[0040] memory for storing computer programs;

[0041] A processor is used to implement the above-mentioned steps of the complex scene understanding method when executing the computer program.

[0042] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned complex scene understanding method are implemented.

[0043] The above technical solution of the present invention has the following advantages over the prior art:

[0044] The complex scene understanding method described in this invention uses a self-supervised hybrid attention mechanism to perform multi-dimensional scoring on each object class in the complex scene semantic graph sequence. The multiple scores are weighted to calculate the importance score for each class of object. A dynamic feedback adjustment mechanism adjusts the weighting ratio based on user preference data provided online, updating the importance score of each class of object in real time. The method uses a large number of complex scene videos to perform self-supervised learning of the model parameters based on the hybrid attention mechanism. Through personalized interactive feedback, the method dynamically models each user's individual needs, improving the quality of downstream application services. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:

[0046] Figure 1 This is a flowchart of an implementation method for complex scene understanding provided by the present invention;

[0047] Figure 2 This is a specific architecture diagram of a complex scene semantic understanding system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The core of the present invention is to provide a complex scene understanding method, system, device and computer storage medium that can effectively adapt to the personalized needs of each user.

[0049] In order to enable those skilled in the art to better understand the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0050] Please refer to Figure 1 , Figure 1 This is a flowchart of a complex scene understanding method provided by the present invention; the specific steps are as follows:

[0051] S101: Acquire a video frame image sequence of a complex scene;

[0052] S102: Performing target detection and relationship recognition on each video frame image to obtain a complex scene semantic graph sequence;

[0053] S103: Performing multi-dimensional scoring on each type of object in the complex scene semantic graph sequence based on a hybrid attention mechanism trained by self-supervisory training, and performing weighted calculation on the multiple scores to obtain an importance score for each type of object;

[0054] S104: Based on the dynamic feedback adjustment mechanism, the weight ratio is adjusted according to the preference data of the user's online feedback, and the importance score of each type of object is updated in real time.

[0055] Based on the above embodiment, this embodiment describes step S101 in detail:

[0056] Read the video captured by the camera as input and process it as a sequential sequence of image frames.

[0057] Based on the above embodiment, this embodiment describes step S102 in detail:

[0058] Object detection methods can adopt, but are not limited to, mainstream methods such as Fast RCNN and the YOLO series. Relationship recognition can be based on existing scene graph generation methods. Based on object detection and relationship recognition algorithms, a scene semantic graph corresponding to the complex scene is generated. In this graph, objects in the scene are nodes, represented by their semantic labels and visual latent vectors. Relationships between objects are edges, represented by their semantic labels and visual latent vectors.

[0059] Based on the above embodiment, this embodiment describes step S103 in detail:

[0060] The sequence of scene semantic graphs corresponding to the video is used as the input of the attention mechanism. Through different attention mechanisms, the model divides the complex scene semantic graph into layers, resulting in a multi-layered graph sorted by importance, which is convenient for downstream applications.

[0061] User needs are often multi-dimensional, so the present invention designs a hybrid attention mechanism to enable it to learn different importance ranking patterns in a large amount of complex scene video data.

[0062] Specifically, the hybrid attention mechanism is composed of a combination of multiple different attention mechanisms. Through weighted calculation, it obtains the recommendation scores of objects in complex scenes in different dimensions, including attractiveness scores and freshness scores.

[0063] In one embodiment, the attention mechanism extracts features based on attractiveness and freshness, scores each object based on these dimensions, and weights them. Objects with higher weights are given higher importance.

[0064] Based on the Attraction Attention mechanism and the Background Attention mechanism, each type of object in the complex scene semantic graph sequence is scored for its attractiveness;

[0065] Based on the freshness attention mechanism, a freshness score is given to each type of object in the complex scene semantic graph sequence.

[0066] Based on the above embodiment, the attractiveness scoring of each type of object in the complex scene semantic graph sequence based on the attractiveness baseline attention mechanism and the background attention mechanism includes:

[0067] Based on the attractiveness benchmark attention mechanism, each type of object in the complex scene semantic graph sequence is scored with an attractiveness benchmark score:

[0068] Obtain a complex scene semantic graph sequence and original frame sequence information arranged in chronological order, and use graph embedding technology to embed the nodes and edges in each complex scene semantic graph into a high-dimensional space, add a position code to each node, and obtain the first target complex scene semantic graph sequence;

[0069] Reconstructing nodes in the next single-frame scene semantic graph based on the single-frame graph in the first target complex scene semantic graph sequence;

[0070] Calculating a first reconstruction accuracy rate for each node, and calculating an attractiveness benchmark score for each node based on the first reconstruction accuracy rate;

[0071] Scoring the background score of each type of object in the complex scene semantic graph sequence based on the background attention mechanism;

[0072] Obtaining an attractiveness score for each node based on the attractiveness benchmark score and the background score;

[0073] Based on the attractiveness score of each node, the attractiveness score of each class of objects is integrated.

[0074] The attractiveness score refers to the score of objects that are more attractive to users in the general sense. This design is based on an important assumption: when users stop to watch (that is, generate "focusing" behavior), their visual attention will remain stable in a local time period, thereby showing structural consistency in the short-term scene graph. On the contrary, if an attractive object (that is, a relatively stable object in a video clip) is not captured in a certain video, it means that the scene switches quickly and the tourists have left without generating "focus". The attractiveness baseline attention mechanism network receives a scene graph sequence input arranged in chronological order while retaining the original frame sequence information. The learning goal of the model is to reconstruct the target node in the next single-frame scene graph based on the current input sequence. The attractiveness baseline attention mechanism designed by the present invention has the following characteristics: only a sequence s is used each time. l The single-frame scene graph in the image is used to predict the nodes of the target scene graph (combined with its position encoding). As training progresses, nodes that appear frequently between consecutive frames (i.e., attractive objects) will show higher reconstruction accuracy in subsequent scene node predictions - because these attractive patterns are most likely to appear repeatedly in local consecutive frames. By combining graph embedding and position encoding techniques, the attractiveness benchmark attention mechanism can calculate the attractiveness benchmark score S in a position-sensitive manner. a (When predicting the target scene, the attractiveness baseline attention mechanism only acts on the current single-frame scene itself).

[0075] Based on the above embodiment, scoring the background score of each type of object in the complex scene semantic graph sequence based on the background attention mechanism includes:

[0076] Randomly mask a single frame in a complex scene semantic graph sequence and reconstruct all node categories in the masked graph with equal probability;

[0077] The second reconstruction accuracy of each node is calculated, and the background score of each node is calculated according to the second reconstruction accuracy.

[0078] A higher background score means that the object is likely to be evenly distributed in each segment of the tour video. Such objects cannot provide effective information for tour points of interest or freshness judgment. i The background score of this paper adopts the algorithm idea of mask modeling: given an input sequence s l When the attractive baseline attention mechanism model is used to randomly mask a graph input G [mask] , the training goal is to reconstruct the masked graph G with equal probability [mask] All node categories c∈[1,…C] in the masked graph G. This process can be transformed into the cross entropy loss calculation of the standard classification task, where the "true label" is the masked graph G. [mask]The core logic is that when the network receives a scene graph input with randomly shuffled frame order (from the same tour video) and needs to reconstruct all node categories of another randomly selected scene graph in the same video, the background objects are the easiest to be accurately reconstructed. Such nodes will show higher reconstruction accuracy because the background elements are highly consistent between any randomly selected frames. This method is essentially to apply an attention mechanism to randomly paired scene graphs in the entire dataset to achieve node reconstruction across scene graphs. Define the background score as S b , then the final attractiveness score S A = Attractiveness benchmark score S a -αS b , where α is a predefined hyperparameter.

[0079] Based on the above embodiment, the freshness scoring of each type of object in the complex scene semantic graph sequence based on the freshness attention mechanism includes:

[0080] Constructing a temporal relationship representation for the complex scene semantic graph sequence by using a position encoding technology to obtain a second target complex scene semantic graph sequence;

[0081] Reconstructing the nodes in the next single-frame scene semantic graph based on the second target complex scene semantic graph sequence;

[0082] Calculating the third reconstruction accuracy of each node, and calculating the freshness score of each node according to the third reconstruction accuracy;

[0083] According to the freshness score of each node, the freshness score of each type of object is integrated.

[0084] In the freshness detection task, position encoding plays a key role in the characterization of sequence temporal relationships. This network substructure aims to predict target objects in subsequent frames by mining the temporal causal relationship between frames. The present invention relies on position encoding technology to construct a temporal relationship representation of video frames, and its training goal is to predict objects in subsequent target scenes based on the input sequence. Objects with a low probability of reconstruction during the training process will be judged as novel objects. For the feature extraction of freshness objects, a second attention model is used to sample video clips in sequence, and ensure that under the correct constraints of position encoding, the background attention mechanism only acts on paired image embeddings within this continuous temporal session. The introduction of position encoding achieves dual protection: on the one hand, the attention calculation is strictly limited to the scene graph pairs within the temporal sequence, and on the other hand, the feature extraction process is forced to follow the established temporal relationship. Define the freshness score as S F .

[0085] Based on the above embodiment, the hybrid attention mechanism based on self-supervised training performs multi-dimensional scoring on each type of object in the complex scene semantic graph sequence, and performs weighted calculation on multiple scores to obtain the importance score of each type of object, which also includes:

[0086] According to the degree of user demand for each type of object, a specific demand score is given to each type of object in the complex scene semantic graph sequence.

[0087] The specific needs score refers to the importance of dimensions that require special attention in different application scenarios. In one embodiment, using a cognitive task for assisting the blind as an example, this score quantifies the degree of need for each category of objects by visually impaired participants based on questionnaire results. The specific calculation process is to first accumulate the total score of each category of objects in the questionnaire, and then normalize the scores of all categories.

[0088] The specific category c in the scene is in the sequence s l The final reasoning score is calculated as follows:

[0089] L c,l =λ·S c,p,l +β·S c,n,l +γ·S c,f,l

[0090] Based on the above embodiments, the present invention describes in detail the self-supervised training process of the hybrid attention mechanism:

[0091] For each training video, the present invention first uniformly samples n frames of images per minute from the first-person perspective video sequence. The total number of these sampled frames N tot =n×minutes will be used as training images. Then, scene graph generation technology is used to convert each training image into a scene graph. Specifically, each edge E in the generated scene graph ij Represents the connected node n i With n j The specific relationship type between (such as "in...", "on...", "lean on", etc.), where n i Represents the i-th node in the graph. Each node essentially corresponds to a detection object in the scene, such as "fish", "lake", "boy", etc. The scene graph generated by the k-th image is denoted as G k , composed of several {n i -E ij -n j} tuple. We press s l ={G l ,…,G l+m ,…,G l+M} form to construct the training sequence, where the lth sequence s l Contains M scene graphs. Each training batch contains Nb An input training sequence.

[0092] This paper proposes a sequence graph-level mask modeling paradigm. Object selection is based on ranking score, where the ranking score P l x ∈R C As output (C represents the total number of categories in the training data). l x Assign a score to each type of object in the scene, its i-th element Indicates the score of the corresponding category. The superscript x is used as a placeholder to replace the scores of various categories (such as S a ,S b ,S p ,S f ,S n ), each Based on sequence diagrams l ={G1,…,G k ,…,G M}Calculated, that is Here the function f x represents the feature extractor in a deep neural network (x identifies the subnetwork from which the feature originated: x = b for the background subnetwork (background attention mechanism), x = f for the freshness subnetwork (freshness attention mechanism), and x = a for the attractiveness subnetwork (attractive baseline attention mechanism)). Function g is the classifier responsible for mapping the embedded features of the scene graph sequence to C categories. Finally, the three scoring sub-sets (background / freshness / attractiveness) are combined to produce the final ranked list. The present invention sometimes adds additional subscripts to the scores, such as S c,a,l represents the attractiveness score of the c-th object in the l-th image sequence.

[0093] During each batch training, the sub-network (attention mechanism) receives a continuous graph input sequence s of length M l ={G l ,…,G l+m ,…,G l+M Each training batch contains N such sequences, and its objective function is defined as:

[0094]

[0095] Here, θ (j) k∈Y represents the learnable classifier parameters corresponding to the true label of the jth class. (l) Indicates that in the (l+M+1)th scene graph (i.e., the next frame scene graph) after the input sequence, there is at least one node belonging to the kth class. Represents the feature vector of the lth graph sequence, which is extracted by the attraction sub-network.

[0096] Based on the above embodiment, this embodiment describes step S104 in detail:

[0097] The present invention adjusts the weighting of multiple dimensions through interaction with users based on a dynamic feedback adjustment mechanism, so that the output results are close to the importance expected by each user for different objects.

[0098] In one embodiment, the present invention adjusts the weight ratio based on the maximum likelihood estimation method according to the preference data of the user's online feedback.

[0099] In a specific embodiment, the present invention gradually learns and adapts to the personalized needs and preferences of users through the interaction process with them; through a dynamic feedback adjustment mechanism, the multi-dimensional demand parameters of visually impaired people are continuously updated and recorded during the tour;

[0100] The present invention continuously learns and processes the preference data of participants through the user behavior recorded by the log reader. These data are used to train the online module in real time, thereby dynamically weighing the priority weights of the following three. Users can click "like" or "dislike" on the recommended items in their field of view. The present invention adjusts the weight parameters of the three branches in real time by reading these feedback signals. Based on the maximum likelihood estimation (MLE) method, dynamic feedback adjustment optimizes the weight ratio of the three recommendation sources, generates an updated navigation plan and presents it in the next frame. This process can be regarded as the participant implicitly defining the priority screening parameters of the scene graph nodes. The present invention continuously reads the feedback data of the visually impaired and dynamically updates the recommendation parameterized model accordingly.

[0101] The specific implementation is as follows:

[0102] For the i-th object in the recommendation list, we define:

[0103] Weight vector: ω T =[λ,β,γ] T ∈R 3

[0104] Rating vector:

[0105] Feedback from visually impaired users is modeled as binary values The judgment rules are:

[0106] When the user observes the i-th recommendation (e.g. "trees"):

[0107] Click

[0108] Click

[0109] Based on this, the log-likelihood function of user feedback is modeled as:

[0110]

[0111] It can be seen that the above equation is essentially a learnable parameter ω∈R 3 , input is A logistic regression model of . Represents all objects s in the training data i In other words: when the feedback (i.e., the user "likes" the current object such as a tree), this feedback will significantly affect the system's parameterized learning of the "like" label; when ("dislike"), the system will optimize parameters based on negative feedback. τ is a preset hyperparameter (the temperature coefficient of softmax). Based on this logistic regression model, the present invention uses maximum likelihood estimation (MLE) technology to learn the weight vector ω by maximizing the log-likelihood function log p(m_i^H(cl)│z,ω). This unit performs MLE optimization updates using gradient descent:

[0112]

[0113] Based on the above embodiment, by continuously repeating steps S101-S104, the correspondence between the system and the user's value requirements is continuously improved, and the importance of different objects in the scene is output.

[0114] The embodiment of the present invention further provides a complex scene semantic understanding system; the specific system may include:

[0115] Video frame image acquisition module, used to acquire video frame image sequences of complex scenes;

[0116] The scene semantic graph acquisition module is used to perform target detection and relationship recognition on each video frame to obtain a complex scene semantic graph sequence;

[0117] A multi-dimensional scoring module is used to perform multi-dimensional scoring on each type of object in the complex scene semantic graph sequence based on a hybrid attention mechanism trained through self-supervision, and perform weighted calculation on multiple scores to obtain an importance score for each type of object;

[0118] The importance score determination module is used to adjust the weight ratio based on the preference data of users' online feedback based on the dynamic feedback adjustment mechanism, and update the importance score of each type of object in real time.

[0119] The complex scene semantic understanding system of this embodiment is used to implement the aforementioned complex scene understanding method. Therefore, the specific implementation methods of the complex scene semantic understanding system can be seen in the embodiment part of the complex scene understanding method above. For example, the video frame image acquisition module, the scene semantic graph acquisition module, the multi-dimensional scoring module, and the importance score determination module are respectively used to implement steps S101, S102, S103, and S104 in the aforementioned complex scene understanding method. Therefore, its specific implementation methods can refer to the descriptions of the corresponding embodiments of each part and will not be repeated here.

[0120] like Figure 2 , Figure 2 A specific architecture diagram of a complex scene semantic understanding system provided for another embodiment of the present invention.

[0121] A specific embodiment of the present invention further provides a complex scene semantic understanding device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the above-mentioned complex scene understanding method when executing the computer program.

[0122] A specific embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned complex scene understanding method are implemented.

[0123] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0124] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A system that specifies the functions of a box or boxes.

[0125] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction system that is implemented in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0126] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0127] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A complex scene understanding method, characterized in that: include: Obtain video frame image sequences of complex scenes; Perform object detection and relationship recognition on each video frame to obtain a complex scene semantic graph sequence; A hybrid attention mechanism based on self-supervised training is used to perform multi-dimensional scoring on each type of object in the complex scene semantic graph sequence, and multiple scores are weighted to obtain the importance score of each type of object; Based on the dynamic feedback adjustment mechanism, the weight ratio is adjusted according to the preference data of users' online feedback, and the importance score of each type of object is updated in real time.

2. The complex scene understanding method according to claim 1, characterized in that: The hybrid attention mechanism based on self-supervised training performs multi-dimensional scoring on each type of object in the complex scene semantic graph sequence, and performs weighted calculation on multiple scores to obtain the importance score of each type of object, including: Based on the attractiveness baseline attention mechanism and the background attention mechanism, the attractiveness score of each type of object in the complex scene semantic graph sequence is performed; Based on the freshness attention mechanism, a freshness score is given to each type of object in the complex scene semantic graph sequence.

3. The complex scene understanding method according to claim 2, characterized in that: The method of scoring the attractiveness of each type of object in the complex scene semantic graph sequence based on the attractiveness baseline attention mechanism and the background attention mechanism includes: Based on the attractiveness benchmark attention mechanism, each type of object in the complex scene semantic graph sequence is scored with an attractiveness benchmark score: Obtain a complex scene semantic graph sequence and original frame sequence information arranged in chronological order, and use graph embedding technology to embed the nodes and edges in each complex scene semantic graph into a high-dimensional space, add a position code to each node, and obtain the first target complex scene semantic graph sequence; Reconstructing nodes in the next single-frame scene semantic graph based on the single-frame graph in the first target complex scene semantic graph sequence; Calculating a first reconstruction accuracy rate for each node, and calculating an attractiveness benchmark score for each node based on the first reconstruction accuracy rate; Scoring the background score of each type of object in the complex scene semantic graph sequence based on the background attention mechanism; Obtaining an attractiveness score for each node based on the attractiveness benchmark score and the background score; Based on the attractiveness score of each node, the attractiveness score of each class of objects is integrated.

4. The complex scene understanding method according to claim 3, characterized in that: Scoring the background score of each type of object in the complex scene semantic graph sequence based on the background attention mechanism includes: Randomly mask a single frame in a complex scene semantic graph sequence and reconstruct all node categories in the masked graph with equal probability; The second reconstruction accuracy of each node is calculated, and the background score of each node is calculated according to the second reconstruction accuracy.

5. The complex scene understanding method according to claim 2, characterized in that: The freshness scoring of each type of object in the complex scene semantic graph sequence based on the freshness attention mechanism includes: Constructing a temporal relationship representation for the complex scene semantic graph sequence by using a position encoding technology to obtain a second target complex scene semantic graph sequence; Reconstructing the nodes in the next single-frame scene semantic graph based on the second target complex scene semantic graph sequence; Calculating the third reconstruction accuracy of each node, and calculating the freshness score of each node according to the third reconstruction accuracy; According to the freshness score of each node, the freshness score of each type of object is integrated.

6. The complex scene understanding method according to claim 2, characterized in that: The hybrid attention mechanism based on self-supervised training performs multi-dimensional scoring on each type of object in the complex scene semantic graph sequence, and performs weighted calculation on multiple scores to obtain the importance score of each type of object, which also includes: According to the degree of user demand for each type of object, a specific demand score is given to each type of object in the complex scene semantic graph sequence.

7. The complex scene understanding method according to claim 1 or 6, characterized in that: The dynamic feedback adjustment mechanism, which adjusts the weight ratio according to the preference data of the user's online feedback, includes: According to the preference data of users' online feedback, the weight ratio is adjusted using the maximum likelihood estimation method.

8. A complex scene semantic understanding system, characterized by: include: Video frame image acquisition module, used to acquire video frame image sequences of complex scenes; The scene semantic graph acquisition module is used to perform target detection and relationship recognition on each video frame to obtain a complex scene semantic graph sequence; A multi-dimensional scoring module is used to perform multi-dimensional scoring on each type of object in the complex scene semantic graph sequence based on a hybrid attention mechanism trained through self-supervision, and perform weighted calculation on multiple scores to obtain an importance score for each type of object; The importance score determination module is used to adjust the weight ratio based on the preference data of users' online feedback based on the dynamic feedback adjustment mechanism, and update the importance score of each type of object in real time.

9. A complex scene semantic understanding device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of a complex scene understanding method as claimed in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a complex scene understanding method as claimed in any one of claims 1 to 7.

Citation Information

Cited By

  • Video data coding method and device, electronic equipment and storage medium

    CN121125992A