Complex scene visual identification rapid construction method based on natural language description

By obtaining natural language description information, performing language parsing and feature extraction, generating a scene semantic table and decoupling it into atomic conditions, using a visual encoder to extract image features, and combining a cross-modal alignment engine and a relational constraint matrix for joint reasoning, the problems of insufficient semantic understanding depth, limited understanding of complex scenes, and lack of adaptability in existing technologies are solved, achieving deep understanding and dynamic optimization of visual recognition.

CN120808107APending Publication Date: 2025-10-17SHANGHAI DISHI WANXIANG TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510905189.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies have problems such as insufficient depth of semantic understanding, limited understanding of complex scenarios, and lack of adaptive capabilities, and are unable to effectively parse the logical structure and multi-entity interaction scenarios in language descriptions.

Method used

By obtaining natural language description information, performing language parsing and feature extraction, generating a scene semantic table, decoupling it into atomic conditions and encoding it into an instruction format, using a visual encoder to extract image features, combining a cross-modal alignment engine and a relational constraint matrix for joint reasoning, and executing global logical verification operations.

Benefits of technology

It achieves a deep understanding of the semantic logic described in natural language, accurately models complex multi-entity interaction scenarios, dynamically optimizes visual recognition strategies, and improves the depth and adaptability of semantic understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808107A_ABST
    Figure CN120808107A_ABST
Patent Text Reader

Abstract

The invention provides a complex scene visual recognition rapid construction method based on natural language description, and relates to the technical field of computer vision and natural language processing, and the method comprises the steps: obtaining natural language description information inputted by a user; performing language analysis and element extraction on the natural language description information to obtain a scene semantic table; decoupling the scene semantic table into M atomic conditions, and obtaining an atomic instruction set; performing visual feature extraction on the to-be-recognized image by using a visual encoder to generate an image feature vector; converting the atomic instruction set into an atomic instruction vector set, and outputting an atomic instruction verification set; generating a relation verification matrix; and executing global logic expression verification operation based on the atomic instruction verification set and the relationship verification matrix, and outputting comprehensive confidence. The technical problems that in the prior art, semantic understanding depth is insufficient, complex scene understanding is limited, and self-adaptive capacity is insufficient are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision and natural language processing, and in particular to a complex scene visual recognition rapid construction method based on natural language description. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, the cross-application of natural language description and image processing is showing unprecedented importance and universality. From intelligent security monitoring to automatic driving environment perception, from medical image analysis to industrial quality inspection, visual understanding technology based on natural language description is profoundly changing the traditional human-computer interaction mode. Especially in the field of complex scene understanding, converting human intuitive language description into machine executable visual recognition task not only greatly reduces the technical use threshold, but also opens up new ways for the development of multi-modal intelligent systems. The application prospect of this technology is broad, covering intelligent customer service, auxiliary diagnosis and treatment, robot navigation and other popular fields, and has become one of the most valuable topics in the cross-research of computer vision and natural language processing.

[0003] However, the existing technical solutions still have several key defects, which seriously restrict the actual application effect. First, the existing technology has the problem of shallow semantic understanding, which cannot effectively analyze the logical structure implied in the language description, and lacks the ability to infer the implied common sense in the language description. Second, the scene understanding ability of multiple entities is limited, and it is difficult to accurately understand the description involving three or more entity interactions. In addition, the dynamic adaptation ability is lacking, and it cannot automatically adjust the recognition strategy according to the contradictory results in the verification process.

[0004] In summary, the existing technology has the technical problems of insufficient semantic understanding depth, limited complex scene understanding, and lack of self-adaptation ability. Therefore, a method is needed to solve the above problems. SUMMARY

[0005] The present disclosure provides a complex scene visual recognition rapid construction method based on natural language description, which solves the technical problems of insufficient semantic understanding depth, limited complex scene understanding, and lack of self-adaptation ability in the prior art.

[0006] According to a first aspect of the present disclosure, a complex scene visual recognition rapid construction method based on natural language description is provided, comprising:

[0007] acquiring natural language description information input by a user, the natural language description information including object entities, attribute descriptions, dynamic behaviors, and relationship topologies;

[0008] language description information is subjected to language analysis and element extraction to obtain a scene semantic table, the scene semantic table including an entity mapping set, a behavior chain list, a relationship constraint matrix, and a global logic expression;

[0009] The scene semantic table is decoupled into M atomic conditions, the atomic conditions are coded into an instruction format, and an atomic instruction set is obtained;

[0010] A visual encoder is used to extract visual features from an image to be recognized, image information is converted into a feature vector, and an image feature vector is generated;

[0011] The atomic instruction set is converted into an atomic instruction vector set, a semantic similarity entropy value of the image feature vector and the atomic instruction vector set is calculated by a cross-modal alignment engine, a dynamic threshold is used to determine the establishment state of the M atomic instructions, and an atomic instruction verification set is output;

[0012] A relationship verification matrix is generated by performing spatial behavior joint reasoning based on the atomic instruction verification set and the relationship constraint matrix;

[0013] A global logic expression verification operation is performed based on the atomic instruction verification set and the relationship verification matrix, and a comprehensive confidence is output.

[0014] One or more technical solutions provided in the present disclosure have at least the following technical effects or advantages: natural language description information input by a user is obtained, the natural language description information including object entities, attribute descriptions, dynamic behaviors, and relationship topologies; language analysis and element extraction are performed on the natural language description information to obtain a scene semantic table, the scene semantic table including an entity mapping set, a behavior chain list, a relationship constraint matrix, and a global logic expression; the scene semantic table is decoupled into M atomic conditions, the atomic conditions are coded into an instruction format, and an atomic instruction set is obtained; a visual encoder is used to extract visual features from an image to be recognized, image information is converted into a feature vector, and an image feature vector is generated; the atomic instruction set is converted into an atomic instruction vector set, a semantic similarity entropy value of the image feature vector and the atomic instruction vector set is calculated by a cross-modal alignment engine, a dynamic threshold is used to determine the establishment state of the M atomic instructions, and an atomic instruction verification set is output; a relationship verification matrix is generated by performing spatial behavior joint reasoning based on the atomic instruction verification set and the relationship constraint matrix; and a global logic expression verification operation is performed based on the atomic instruction verification set and the relationship verification matrix, and a comprehensive confidence is output. The technical problems of insufficient semantic understanding depth, limited complex scene understanding, and lack of self-adaptation in the prior art are solved. The technical effects of deeply understanding the semantic logic of natural language description, accurately modeling a multi-entity complex interaction scene, and dynamically optimizing a visual recognition strategy are achieved.

[0015] The above description is only a summary of the technical solutions of the present application. In order to enable the technical means of the present application to be more clearly understood, and to be implemented according to the content of the description, and in order to enable the above and other purposes, characteristics and advantages of the present application to be more apparent and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present disclosure or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only exemplary, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.

[0017] Figure 1 A flowchart of a natural language description-based complex scene visual recognition rapid construction method provided by the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION

[0018] The exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, including various details in order to facilitate understanding. They should be considered as merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, the description below omits the description of well-known functions and structures.

[0019] Embodiment one, the natural language description-based complex scene visual recognition rapid construction method provided by the embodiments of the present disclosure is described with reference to Figure 1 The method comprises the following steps:

[0020] S1: Obtain the natural language description information input by the user, wherein the natural language description information comprises object entities, attribute descriptions, dynamic behaviors and relationship topologies.

[0021] Specifically, the natural language description information input by the user is obtained, and the natural language description information mainly includes four core elements: object entities, attribute descriptions, dynamic behaviors and relationship topologies. These elements together constitute a complete description of the complex scene.

[0022] Among them, the object entity refers to the specific object or organism explicitly mentioned in the description as the basic constituent unit of the scene. These entities not only include specific categories, but also distinguish instances through determiners, and are often accompanied by existence statements. Attribute description is a detailed description of the static characteristics of the object, mainly including three types of characteristics: visual attributes, state attributes, and quantitative attributes. Dynamic behavior element describes the actions and state changes between entities or the entity itself, including three key characteristics: explicit behavior subject, specific action type, and temporal characteristics. Relationship topology element expresses the spatial and logical relationship between entities, which can be divided into spatial state relationship and behavior logical relationship. This structured element decomposition provides clear guidance for subsequent visual scene understanding: object entities correspond to specific targets of visual detection, attribute descriptions serve as filtering conditions for object recognition, dynamic behaviors indicate actions that need to be analyzed in detail, and relationship topology constitutes the constraint conditions for scene understanding. Through such detailed analysis, natural language descriptions can be accurately converted into structured representations that machines can understand, laying a solid foundation for complex scene recognition tasks.

[0023] S2: language analysis and element extraction on the natural language description information, obtaining a scene semantic table, the scene semantic table including an entity mapping set, a behavior chain list, a relationship constraint matrix, and a global logic expression;

[0024] Further, the step S2 further includes:

[0025] The pre-processing of the natural language description information includes basic cleaning, expression standardization, and complex sentence segmentation.

[0026] Based on the pre-processed natural language description information, the syntax structure is parsed, the dominance relationship mapping between syntax components is established, and a syntax dependency tree is constructed.

[0027] Based on the syntax dependency tree, a four-dimensional element extraction operation is performed, including entity extraction, attribute binding, behavior analysis, and relationship modeling.

[0028] According to the four-dimensional element extraction result, a scene semantic table is constructed, including an entity mapping set, a behavior chain list, a relationship constraint matrix, and a global logic expression.

[0029] Specifically, the pre-processing operation is performed on the natural language description information input by the user, including basic cleaning, expression standardization, and complex sentence segmentation. Among them, the basic cleaning operation removes stop words without actual semantics, such as "ah", "um", "ah", etc. The expression standardization operation unifies the expressions of synonyms, eliminates dialect and colloquial differences, such as standardizing "bicycle" to "bicycle". The complex sentence segmentation operation splits long sentences into independent simple sentences according to conjunctions.

[0030] The pre-processed natural language description information is parsed, the sentence syntax structure is analyzed, the core subject, predicate, object and modifying component of each simple sentence are recognized, the tense and voice features of the action verbs are labeled, the domination mapping relationship between the syntax components is established, and a syntax dependency tree is constructed.

[0031] Based on the syntax dependency tree, four-dimensional element extraction is performed on the object entity, attribute description, dynamic behavior and relationship topology. Entity extraction is performed, all noun phrases are recognized as candidate entities, a unique identifier is assigned, coreference resolution is performed, and pronouns are associated with specific entities referred to in the previous text; attribute binding is implemented, adjective and preposition phrases are extracted, such as "blue" and "hat-wearing", and attributes are associated with corresponding entities to form key-value pairs; behavior analysis is performed, core action verbs and their tense labels are labeled, and behavior subjects and objects are bound; relationship modeling is completed, spatial relationships are analyzed, and logical connection words are recognized to construct conditional constraints.

[0032] According to the four-dimensional element extraction result, a scene semantic table is constructed. A unique identifier is assigned to each entity, entity types and attribute key-value pairs are stored, an entity attribute mapping dictionary is constructed, and an entity mapping set is generated; the behavior identifier, subject and object references, and verb tense labels are recorded in time sequence, and a behavior chain list is compiled; a three-dimensional relationship table is established to store spatial and behavioral logical relationships, a relationship constraint matrix is constructed, the first dimension of the three-dimensional relationship table is the relationship type, including spatial and behavioral relationships, the second dimension is the subject identifier and the object identifier, and the third dimension is the spatial relationship parameter or the behavioral logical relationship description; using first-order logic rules, entities, behaviors and relationship conditions are combined to generate a global logical expression. The entity mapping set, behavior chain list, relationship constraint matrix and global logical expression are integrated into a scene semantic table, and the generated scene semantic table is checked for semantic consistency, including implicit relationship derivation, conflict detection and operation integrity check. Specifically, first, necessary relationships are supplemented based on preset rules, such as automatically adding a "contact" relationship to the "ride" behavior. Second, entity position contradictions and behavioral logic conflicts are scanned, and key element missing conditions are verified, such as no main behavior and unbound attributes. Finally, according to the conflict type, the coreference resolution engine is activated for secondary analysis, and user supplement requests are initiated according to missing elements to complete the overall optimization of the scene semantic table.

[0033] S3: decouple the scene semantic table into M atomic conditions, encode the atomic conditions into an instruction format, and obtain an atomic instruction set;

[0034] Further, the step S3 further includes:

[0035] The scene semantic table is disassembled into non-divisible basic verification conditions to obtain M atomic conditions, and the atomic conditions only verify a single fact and do not contain logical connection words.

[0036] Establish conditional dependencies between M atomic conditions, mark and display dependencies, and generate a dependency graph;

[0037] Generate an atomic instruction based on the atomic condition and the dependency graph, wherein the atomic instruction includes a verification type, a target object, verification parameters, and a dependency list;

[0038] Topologically sort M atomic instructions and output a structured atomic instruction set.

[0039] Specifically, the entity mapping set and behavior chain list in the scene semantic table are traversed to identify the minimum verification unit. Each component is broken down into indivisible basic verification conditions to obtain M atomic conditions. Each atomic condition must satisfy the following conditions: only verify a single fact, such as the existence of an object, the value of an attribute, or the occurrence of an action; and contain no logical conjunctions, such as "and" or "or." For example, valid atomic conditions include "a cat exists," "the cat is white," and "the cat is running," while invalid atomic conditions include "a cat exists and the cat is white."

[0040] Establish conditional dependencies and mark explicit dependencies. The dependency rules are as follows: behavioral conditions depend on the entity existence conditions of their subject and object; attribute conditions depend on the entity existence conditions of their entities. When an atomic condition requires the results of other conditions, annotate the dependency chain. For example, the condition "Verify that the dog is chasing the person" depends on "the dog exists" and "the person exists." Dependencies are represented using a directed graph, with nodes representing atomic conditions and directed edges A→B indicating that B depends on A. Finally, a dependency graph is generated based on the dependencies of all atomic conditions.

[0041] Structure M atomic conditions, encoding each as a machine-readable atomic instruction, specifically including the verification type, target object, verification parameters, and dependency list. Redundant checks are then merged for the generated M atomic instructions, combining multiple attribute checks for the same entity into a single instruction. The M atomic instructions are topologically sorted according to the dependency graph, with instructions without dependencies placed first and those that depend on other instructions placed last. Finally, each instruction is assigned a unique identifier, and the structured atomic instruction set is output.

[0042] S4: Using a visual encoder to extract visual features of the image to be identified, converting image information into a feature vector, and generating an image feature vector;

[0043] Furthermore, this step S4 also includes:

[0044] Preprocessing the image to be recognized, wherein the preprocessing includes size standardization and pixel normalization;

[0045] Through a multi-layer convolutional neural network, the low-level texture features, mid-level component features, and high-level semantic features of the image to be identified are extracted step by step to obtain the final feature tensor;

[0046] The final feature tensor is converted into the original feature vector by adopting a dual strategy of considering global information and spatial reservation, performing dimension reduction and normalization operations on the original feature vector, optimizing the representation density and matching efficiency, and outputting the image feature vector.

[0047] Specifically, the image to be identified is preprocessed. First, the size standardization operation is performed, the input image is scaled to the preset resolution, the bilinear interpolation algorithm is used to maintain the geometric proportion, and the edge adaptive padding is performed on the part exceeding the part. Then, pixel normalization is performed, the original RGB channel is linearly mapped from the integer range of 0-255 to the floating point interval of [-1, 1] or [0, 1], channel-level standardization is performed, and each pixel point is respectively subtracted from the mean vector and divided by the standard deviation vector.

[0048] Visual features are extracted step by step through a multi-layer convolutional neural network. In the primary convolution stage, a small size convolution kernel is used to scan the image in a sliding window manner, and the ReLU activation function is used to extract local edge, color block and texture, etc. bottom features, output the initial feature tensor H1xW1xC1, where H1 represents the height, W1 represents the width, and C1 represents the channel value. After entering the intermediate convolution stage, the feature is gradually sampled through convolution or maximum pooling operation with a step of 2, the receptive field of a single neuron is expanded, so that the feature map at each position can cover a larger area of the original image, at this time the feature starts to represent the object component level information such as wheels, window frames, etc. The intermediate feature tensor H2xW2xC2 is finally output after further optimization, at this time the size of the feature tensor is halved and the number of channels is multiplied. In the deep convolution stage, the network continues to capture global semantic features such as "vehicle as a whole" and "building outline" through stacked residual modules or attention mechanisms, at this time the spatial size of the feature tensor is further compressed, but the number of channels is expanded, and finally the final feature tensor H3xW3xC3 is output.

[0049] In order to preserve the spatial structure information of the feature tensor, a position enhancement coding strategy is adopted. The normalized coordinate value of each feature vector in the final feature tensor is directly spliced, for example, the relative coordinates of the center point of each position in the 7x7 feature map in the original image are added. This explicit position marking method can utilize both global statistical features and trace the original spatial distribution of each feature during subsequent processing.

[0050] The final feature tensor is converted into a vector form. A dual strategy that takes into account global information and spatial preservation is adopted. On the one hand, the global average pooling is used to calculate the mean value of each channel along the spatial dimension, compressing the multi-dimensional feature tensor of HxWxC into a vector of 1x1xC, preserving the statistical information of the channel dimension, and generating an overall representation vector. On the other hand, the position-enhanced feature map is flattened in spatial order to generate a composite feature vector that contains both visual semantics and position clues. The overall representation vector and the composite feature vector are integrated into the original feature vector, which is then fine-tuned. First, the high-dimensional features are projected into a compact embedding space through a fully connected layer, reducing computational redundancy and improving feature density. Then, L2 normalization is performed, dividing the feature vector by its Euclidean norm, so that all vectors are constrained to the unit hypersphere. At this point, the cosine similarity between vectors is equivalent to the dot product operation, optimizing the efficiency of subsequent matching. Finally, the image feature vector is output. This processing method makes the generated image feature vector suitable for both traditional classifier processing and tasks that require spatial positioning, such as identifying object categories and determining their approximate orientation in image retrieval.

[0051] It should be particularly noted that feature extraction based on a single frame of image cannot achieve accurate identification of behavior features. Since behavior feature extraction relies on motion information provided by consecutive multiple frames of images, the behavior features extracted in the current step can only serve as a preliminary reference and cannot be directly used for accurate verification of subsequent atomic instructions. Specifically, although some motion cues such as body posture changes and limb displacements can be detected, it is impossible to distinguish similar actions based on single-frame information, such as determining whether "both feet off the ground" is the jump-off action of "jumping" or the jump-off phase of "running". These preliminary extracted behavior features mainly play a pre-screening function, quickly excluding scenes that obviously do not meet the instructions, such as detecting a completely stationary human body, which can directly deny the "running" instruction. The final behavior verification needs to be completed in the subsequent relationship constraint matrix verification process combined with physical common sense rules and spatial relationships.

[0052] S5: converting the set of atomic instructions into a set of atomic instruction vectors, calculating the semantic similarity entropy value of the image feature vector and the set of atomic instruction vectors through the cross-modal alignment engine, determining the establishment state of the M atomic instructions based on a dynamic threshold, and outputting a set of atomic instruction verification results;

[0053] Further, the step S5 further comprises:

[0054] converting the set of atomic instructions into a set of atomic instruction vectors through a text encoder;

[0055] calculating the cosine similarity between the image feature vector and the set of atomic instruction vectors, and the specific formula is:

[0056]

[0057] wherein, similarity is cosine similarity, value range is [-1, 1], V image is image feature vector, V text is atomic instruction vector, molecule V image ·V text denotes dot product operation between vectors, denominator ||V image || and ||V text || respectively represent vector module length of image feature vector and atomic instruction vector;

[0058] According to the cosine similarity, a semantic similarity entropy value is calculated, and a specific formula is as follows:

[0059]

[0060] wherein, S is semantic similarity entropy value, value is normalized to [0, 1] interval, and similarity is cosine similarity;

[0061] The calculated semantic similarity entropy value is subjected to confidence calibration, so as to ensure that the semantic similarity entropy value is accurate and objective, and a specific formula is as follows:

[0062] S' = αS + β;

[0063] wherein, S' is calibrated semantic similarity entropy value, S is semantic similarity entropy value, α represents slope, and β represents intercept;

[0064] Based on the set initial threshold Q and the semantic similarity entropy value S', the atomic instruction state is judged, if S' ≥ Q, it is determined that the atomic instruction is established, if S' < Q-0.2, it is determined that the atomic instruction is not established, and if S' is between [Q-0.2, Q], it is determined that whether the atomic instruction is established cannot be confirmed.

[0065] According to the judgment result of the M atomic instructions, an atomic instruction verification set is generated.

[0066] Specifically, the atomic instruction set is converted into an atomic instruction vector set by a text encoder. The atomic instruction set is received, and the semantic integrity of each atomic instruction is checked. Since the instruction is a complete description sentence, it directly enters the encoding link. For the text information contained in each atomic instruction, it is input into the text encoder CLIP-Text-Transformer, performs word segmentation operation, uses a subword tokenizer to split the sentence into meaningful semantic units, and converts the word segmentation into a digital ID sequence through a preset word table. The digital ID sequence is obtained, and deep semantic understanding is performed. The Transformer architecture dynamically analyzes the dependency relationship between words by taking the dependency list in the atomic instruction as a reference object and accurately captures the modifier and limitation structure through multiple layers of self-attention mechanism. After multiple rounds of context perception calculation, the system condenses the semantic information of the entire sentence into a fixed-dimensional feature vector through a special flag or mean pooling strategy, and outputs a fixed-dimensional atomic instruction vector. The atomic instruction vector set is generated by integrating all atomic instruction vectors. The finally generated atomic instruction vector has three important characteristics: dimension uniformity ensures the comparability of instructions of different lengths; semantic encoding characteristics make the vector numerical distribution accurately reflect the text connotation; and cross-modal alignment characteristics ensure that the text vector and the image vector are in the same semantic space.

[0067] The cosine similarity is calculated according to the obtained image feature vector and atomic instruction vector set, and the specific formula is:

[0068]

[0069] Wherein, similarity is the cosine similarity, which mainly acts on the similarity through the angle between vectors, V image is the image feature vector, V text is the atomic instruction vector, the numerator V image ·V text is the dot product operation between vectors, and the denominator ||V image || and ||V text || respectively represent the vector lengths of the image feature vector and the atomic instruction vector. Since the vectors are normalized to unit length during actual calculation, ||V image || and ||V text || are both 1.

[0070] The semantic similarity entropy value is further calculated according to the obtained cosine similarity, and the specific formula is:

[0071]

[0072] Wherein, S is the semantic similarity entropy value, which is normalized to the interval [0, 1], and similarity is the cosine similarity.

[0073] The semantic similarity entropy value is calibrated for confidence to ensure the accuracy and objectivity of the semantic similarity entropy value. The specific formula is:

[0074] S' = aS + b;

[0075] wherein S' is the calibrated semantic similarity entropy value, S is the semantic similarity entropy value, a represents the slope, which is set to 0.8 in this value, for adjusting the semantic similarity entropy value and controlling the confidence compression degree to prevent the model from being overly confident and to suppress high confidence, and b represents the intercept, which is set to 0.1 in this value, for avoiding zero probability problems and preserving uncertainty.

[0076] According to the calibrated semantic similarity entropy value, the establishment state of the M atomic instructions is determined. The initial threshold Q is set to 0.7. If S' ≥ 0.7, the atomic instruction is determined to be established. If S' < 0.5, the atomic instruction is determined to be not established. If S' is between 0.5 and 0.7, it is determined that the establishment of the atomic instruction cannot be confirmed. Specifically, when S' is in the interval [0.9, 1], it represents that the atomic instruction has a very high matching degree. When S' is in the interval [0.7, 0.9], it represents that the atomic instruction has a high matching degree. When S' is in the interval [0.5, 0.7], it represents that the atomic instruction has a medium matching degree. When S' is in the interval [0, 0.5], it represents that the atomic instruction has a low matching degree.

[0077] According to the above verification process, in the storage format of the atomic instruction, state, entropy value and position data are added. The state includes three states of establishment, non-establishment and judgment, the entropy value is the calibrated semantic similarity entropy value obtained by calculation, and the position data is added only for the established atomic instruction and obtained through the position data in the image feature set. At the same time, the dependency list is rewritten according to the state of the atomic instruction. If the state of the other condition dependent on the atomic condition is not established or cannot be judged, the data in the dependency list is deleted. The atomic instruction verification set is obtained by integrating all the atomic instruction verification results.

[0078] S6: performing spatial behavior joint reasoning according to the atomic instruction verification set and the relationship constraint matrix to generate a relationship verification matrix;

[0079] Further, the step S6 further includes:

[0080] The atomic instruction verification set and the relationship constraint matrix in the scene semantic table are obtained. The relationship constraint matrix is filtered according to the atomic instruction verification set to eliminate non-existent relationships.

[0081] Geometric feature extraction and dynamic constraint solving are performed on each pair of entities with spatial constraints, including basic geometric parameters and high-level derived features, and the dynamic constraints include orientation relationship, distance relationship and occlusion relationship, to complete spatial relationship verification.

[0082] Action chain analysis and physical common sense rationality check are performed on each pair of entities with behavior constraints, including time sequence verification and motion consistency verification, to complete behavior relationship verification.

[0083] Based on the verification results of spatial relationship and behavior relationship, a hierarchical filling mechanism is used to construct a relationship verification matrix, and the matrix is subjected to sparse compression, symmetrization processing and confidence fusion.

[0084] Specifically, an atomic instruction verification set is received, which is stored in a structured form, and each entry contains a Boolean conclusion recording the "existence / non-existence / uncertainty" of the atomic instruction, and also carries position information and calibrated semantic similarity entropy value. At the same time, the relationship constraint matrix in the scene semantic table is obtained, which is stored in a three-dimensional form, the first dimension distinguishes the types of spatial relationship and behavior relationship, the second dimension corresponds to the identifier index of the subject and object, and the third dimension parameter domain saves specific constraint conditions such as maximum allowed distance and orientation angle threshold. According to the atomic instruction verification set, the relationship constraint matrix is screened and processed to eliminate non-existent relationships, for example, if the atomic instruction "there is a cat" is not established, all relationship constraints involving "cat" in the relationship constraint matrix are deleted, and the remaining relationship constraints are supplemented with position information, linking the position information in the atomic instruction verification set to the subject and object information in the second dimension of the relationship constraint matrix, to establish the mapping between entities and regions.

[0085] The spatial relationship verification engine is started. For each pair of entities with preset spatial constraints, information in the third dimension parameter domain is extracted, and geometric features of the entities are extracted from the image feature set, including basic geometric features such as center point coordinates, width and height dimensions, and derived advanced features such as direction vectors and region overlap. Among them, the coordinates of the center are obtained based on the relative position of the center points of the two entity bounding boxes, the size is obtained by calculating the bounding box information, the direction vector is obtained according to the angle of the subject center pointing to the object center, and the region overlap is obtained by the intersection over union IoU to quantify the spatial overlap of the two entities. According to the different spatial relationships, different calculation methods are used for dynamic constraint solving, and the information in the parameter domain is compared to obtain the confidence score of the relationship, which is used for subsequent construction of the relationship verification matrix: the orientation relationship verification adopts the comparison strategy of the direction vector and the preset angle interval, for example, the "rear" relationship requires the direction angle to fall within the range of 135° to 225°, the "left side" relationship requires the direction angle to fall within the range of 225 to 315 degrees, and the "upper side" relationship requires the direction angle to fall within the range of 45 to 135 degrees, and the angle tolerance is dynamically adjusted according to the object distance; the distance relationship is verified by calculating the normalized distance, for example, the "near" relationship may require a distance less than 0.2, and the "far" relationship may require a distance greater than 0.5; the occlusion relationship verification adopts multi-level analysis, first filters the possible occluded entity pairs through the fast intersection over union IoU calculation, then performs pixel-level segmentation mask analysis on the possible overlapping entities, and finally combines optical features such as edge continuity and shadow distribution to comprehensively judge the occlusion situation.

[0086] The behavior logic reasoning stage needs to integrate timing analysis, kinematics calculation and common sense verification. First, the behavior timing diagram is constructed, and for the entity pairs with logical constraints in the relationship constraint matrix, the action chain is strictly checked for the sequence, and in the static image, this verification is realized by analyzing the human posture features. For behaviors involving movement, check the consistency of the motion direction, requiring the angle between the motion vectors of the subject and the object to be less than 45 degrees, and estimate the reasonable relative speed through the size change of the bounding box. Integrate the lightweight physical common sense library to verify the physical reasonableness of key behaviors, such as calculating the distance of "frisbee in the air" from the possible support surface and checking whether the area below meets the support feature. Finally, the verification result is converted into a confidence score between 0 and 1, providing a data basis for the subsequent relationship verification matrix construction.

[0087] After all the verifications are completed, a relationship verification matrix is constructed. The matrix adopts a hierarchical filling strategy: for high-confidence verification results, 1.0 is directly written, representing that the verification relationship is established; for medium-confidence results, the spatial relationship retains the original value, and the behavior relationship is conservatively downgraded; for low-confidence results, 0.0 is directly written, representing that the verification relationship is not established. When a verification contradiction occurs, an arbitration mechanism is automatically started, and the conclusion with more sufficient visual evidence is preferentially adopted. After the relationship verification matrix is constructed, the unconstrained empty items are removed through sparsification optimization, and the symmetry of the bidirectional relationship is forced to be maintained. Each matrix element is attached with complete evidence traceability information, including geometric parameters, visual feature positions, etc. These metadata are encapsulated in a structured format for subsequent analysis and debugging.

[0088] S7: performing a global logical expression verification operation based on the atomic instruction verification set and the relationship verification matrix, and outputting a comprehensive confidence.

[0089] Further, the step S7 further includes:

[0090] transforming the global logical expression in the scene semantic table into a tree structure to generate a logical expression tree, wherein a leaf node of the logical expression tree is an atomic condition, and an intermediate node is a logical operator;

[0091] traversing the logical expression tree to calculate the confidence of the leaf node and the intermediate node;

[0092] obtaining a comprehensive confidence of a root node based on the confidence of the leaf node and the intermediate node, comparing the comprehensive confidence with a dynamic threshold, and judging the existence of the scene.

[0093] Specifically, a logical expression tree is constructed according to the global logical expression in the scene semantic table. The expression is transformed into a tree structure, wherein a leaf node is bound to a specific atomic instruction, and an intermediate node represents a logical operator. A hierarchical verification calculation is performed on the abstract syntax tree constructed. The operation process traverses the logical expression tree from bottom to top. For each leaf node, three key data are extracted from the verification set: verification state, confidence score, and associated visual evidence description. When the atomic instruction involves the interaction relationship of multiple entities, the corresponding relationship confidence in the relationship verification matrix is additionally queried, and the independent verification result and the relationship verification result are fused and calculated.

[0094] For intermediate nodes, the strict operation rules are followed. For AND nodes, the minimum value of all child confidence is taken as the base value, while dynamic decay is applied considering the number of sub-conditions. For OR nodes, the maximum value of child confidence is taken, but a penalizing decay is applied to highly conflicting sub-results. For NOT nodes, the confidence of the original condition is subtracted by 1, and a smoothing process is applied to the boundary value. During the calculation process, the confidence propagation of each node is monitored in real-time, and when a key condition confidence is detected to be lower than the threshold, an early termination mechanism is triggered.

[0095] After obtaining the comprehensive confidence of the root node, it is compared with the dynamic threshold. The basic decision threshold is set to 0.7, and is dynamically adjusted according to the actual scene: appropriately lower the requirement for poor image quality conditions; increase the threshold for high-risk scenarios. At the same time, the comprehensive confidence entropy is calculated, which quantifies the certainty degree of system decision through the information theory formula, considering the uniformity of the distribution of each sub-condition confidence. The final output contains three core contents: binary decision result is or not, continuous comprehensive confidence, and key evidence traceability.

[0096] The decision result is verified in multiple dimensions. First, check if there is a logical contradiction, second, evaluate the coverage of the evidence, and finally analyze the rationality of the confidence entropy. Based on the verification result, natural language explanation is generated to clearly explain the decision basis and the source of uncertainty.

[0097] The technical features of the above embodiments can be combined in any way. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combination of the technical features does not exist contradictions, it should be considered as the scope of the present disclosure.

[0098] The above description of disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for rapidly constructing complex scene visual recognition based on natural language description, characterized by: The method comprises: Obtaining natural language description information input by a user, wherein the natural language description information includes object entities, attribute descriptions, dynamic behaviors, and relationship topology; Performing language parsing and element extraction on the natural language description information to obtain a scene semantic table, wherein the scene semantic table includes an entity mapping set, a behavior chain list, a relationship constraint matrix, and a global logic expression; Decoupling the scene semantic table into M atomic conditions, encoding the atomic conditions into an instruction format, and obtaining an atomic instruction set; Use the visual encoder to extract visual features of the image to be identified, convert the image information into a feature vector, and generate an image feature vector; Convert the atomic instruction set into an atomic instruction vector set, calculate the semantic similarity entropy value between the image feature vector and the atomic instruction vector set through a cross-modal alignment engine, determine the validity status of M atomic instructions based on a dynamic threshold, and output an atomic instruction verification set; Performing spatial behavior joint reasoning based on the atomic instruction verification set and the relationship constraint matrix to generate a relationship verification matrix; A global logic expression verification operation is performed based on the atomic instruction verification set and the relational verification matrix, and a comprehensive confidence level is output.

2. According to claim 1, a method for rapidly constructing complex scene visual recognition based on natural language description is characterized in that: Get the scene semantic table, including: Preprocessing the natural language description information, wherein the preprocessing includes basic cleaning, expression standardization, and compound sentence segmentation; Based on the pre-processed natural language description information, the grammatical structure is parsed, the dominance relationship mapping between grammatical components is established, and the grammatical dependency tree is constructed; Performing a four-dimensional element extraction operation based on the grammatical dependency tree, wherein the four-dimensional element extraction operation includes entity extraction, attribute binding, behavior parsing, and relationship modeling; A scene semantic table is constructed according to the four-dimensional element extraction result. The scene semantic table includes an entity mapping set, a behavior chain list, a relationship constraint matrix, and a global logic expression.

3. The method for rapidly constructing complex scene visual recognition based on natural language description according to claim 1, characterized in that: Get the atomic instruction set, including: Decomposing the scene semantic table into indivisible basic verification conditions to obtain M atomic conditions, each of which verifies only a single fact and does not contain logical connectives; Establish conditional dependencies between M atomic conditions, mark and display dependencies, and generate a dependency graph; Generate an atomic instruction based on the atomic condition and the dependency graph, wherein the atomic instruction includes a verification type, a target object, verification parameters, and a dependency list; Topologically sort M atomic instructions and output a structured atomic instruction set.

4. The method for rapidly constructing complex scene visual recognition based on natural language description according to claim 1, characterized in that: Generate image feature vector, including: Preprocessing the image to be recognized, wherein the preprocessing includes size standardization and pixel normalization; Through a multi-layer convolutional neural network, the low-level texture features, mid-level component features, and high-level semantic features of the image to be identified are extracted step by step to obtain the final feature tensor; A dual strategy that takes into account both global information and space preservation is adopted to transform the final feature tensor into the original feature vector. Dimensionality reduction and normalization operations are performed on the original feature vector to optimize the representation density and matching efficiency, and output the image feature vector.

5. The method for rapidly constructing complex scene visual recognition based on natural language description according to claim 1, characterized in that: Output atomic instruction verification set, including: Converting the atomic instruction set into an atomic instruction vector set by a text encoder; Calculate the cosine similarity between the image feature vector and the atomic instruction vector set. The specific formula is: Among them, similarity is the cosine similarity, the range is [-1,1], V image is the image feature vector, V text is the atomic instruction vector, the molecule V image ·V text表示 Dot product operation between vectors, denominator || V image || and ||V text || respectively represent the vector modulus lengths of the image feature vector and the atomic instruction vector; The semantic similarity entropy value is calculated based on the cosine similarity. The specific formula is: Among them, S is the semantic similarity entropy value, the value is normalized to the interval [0,1], and similarity is the cosine similarity; The confidence level of the calculated semantic similarity entropy value is calibrated to ensure that the semantic similarity entropy value is accurate and objective. The specific formula is: S'=αS+β; Among them, S' is the calibrated semantic similarity entropy value, S is the semantic similarity entropy value, α represents the slope, and β represents the intercept; The atomic instruction status is determined based on the set initial threshold Q and the semantic similarity entropy value S'. If S' ≥ Q, the atomic instruction is determined to be valid. If S' < Q-0.2, the atomic instruction is determined to be invalid. If S' is between [Q-0.2, Q], it is determined that it cannot be confirmed whether the atomic instruction is valid. According to the judgment results of the M atomic instructions, an atomic instruction verification set is generated.

6. The method for rapidly constructing complex scene visual recognition based on natural language description according to claim 1, characterized in that: Generate a relationship verification matrix, including: Obtaining a relationship constraint matrix between the atomic instruction verification set and the scene semantic table, screening the relationship constraint matrix according to the atomic instruction verification set, and eliminating non-existent relationships; For each pair of entities with spatial constraints, geometric features are extracted and dynamic constraints are solved to complete spatial relationship verification. The geometric features include basic geometric parameters and advanced derived features. The dynamic constraints include orientation relationships, distance relationships, and occlusion relationships. Perform action chain analysis and physical common sense rationality check on each pair of entities with behavioral constraints to complete behavioral relationship verification. The action chain analysis includes timing verification and motion consistency verification. Based on the verification results of spatial and behavioral relationships, a hierarchical filling mechanism is used to construct a relationship verification matrix, and the matrix is ​​sparsely compressed, symmetricized and confidence fused.

7. The method for rapidly constructing complex scene visual recognition based on natural language description according to claim 1, characterized in that: Output comprehensive confidence, including: Converting the global logical expression in the scene semantic table into a tree structure to generate a logical expression tree, wherein the leaf nodes of the logical expression tree are atomic conditions and the intermediate nodes are logical operators; Traversing the logical expression tree and calculating the confidence of leaf nodes and intermediate nodes; The comprehensive confidence of the root node is obtained based on the confidence of the leaf nodes and intermediate nodes, and the comprehensive confidence is compared with the dynamic threshold to determine the existence of the scene.

Citation Information

Cited By

  • Knowledge point labeling method and system of natural language processing technology, and electronic equipment

    CN121388193A