A natural language and picture collaborative defined behavior specification supervision large model system
By constructing a set of association mapping rules between natural language and images, and combining it with a pre-trained model to generate draft behavioral norms, the ambiguity of natural language descriptions is solved, and precise supervision of behavioral norms is achieved.
Patent Information
- Application Number
- CN202511204819.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing methods of regulating behavior mainly rely on natural language descriptions, which are vague and ambiguous, making it difficult to accurately express behavioral requirements in complex scenarios. Furthermore, they lack a deep collaborative mechanism between natural language and images, resulting in inaccurate regulation.
A set of association mapping rules between natural language descriptions and normative schematic images is constructed. A text-image collaborative feature set is generated through bidirectional transformation processing. A pre-trained norm generation model is called to extract behavioral norm elements, which are then compared with the regulatory standard library. The final regulatory instructions are generated through iterative adjustments.
This improved the accuracy and efficiency of developing behavioral norms, avoided oversights and deviations inherent in manual formulation, and ensured the precision and reliability of supervision.
Smart Images

Figure CN120705641B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of behavioral regulation technology, and more specifically, to a large-scale behavioral regulation model system that uses natural language and images to collaboratively define behavioral norms. Background Technology
[0002] In today's society, effective regulation of behavioral norms in various scenarios is a crucial link in ensuring social order, public safety, and the healthy development of industries. Traditional methods of regulating behavioral norms mainly rely on manual formulation and enforcement, typically defining them in the form of written descriptions, such as laws and regulations, industry standards, and corporate rules and regulations. However, this single method of natural language description has certain limitations. On the one hand, natural language is ambiguous and polysemous; different people may have different understandings of the same text description, which can easily lead to disputes and deviations in the implementation and judgment of behavioral norms. On the other hand, for some complex behavioral scenarios, it is difficult to comprehensively and intuitively present the specific requirements and boundary conditions of behavior through textual descriptions alone, making it difficult for regulators and those being regulated to accurately grasp the content of the behavioral norms.
[0003] While some studies have attempted to incorporate visual elements such as images to aid in the expression of behavioral norms, there is currently a lack of an effective method to deeply integrate natural language and images for precise definition and efficient oversight of behavioral norms. Existing methods often simply use images as supplementary explanations to text, failing to establish a close connection and collaborative mechanism between natural language and images, thus failing to fully leverage the advantages of both in behavioral norm oversight. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a large-scale behavioral norm supervision model system collaboratively defined by natural language and images, the method comprising:
[0005] A set of association mapping rules is constructed between natural language descriptions and standardized schematic images. The set of association mapping rules includes the correspondence between text semantic elements and image visual elements and the collaborative constraints.
[0006] Receive behavioral description text and scene illustration of the scene to be monitored, and perform bidirectional conversion processing on the behavioral description text and the scene illustration according to the association mapping rule set to obtain a text-image collaborative feature set;
[0007] Based on the text and image collaborative feature set, a pre-trained specification generation model is invoked to perform behavior specification element extraction operations, generating a draft behavior specification that includes behavior boundary conditions and compliance judgment benchmarks;
[0008] The draft code of conduct is compared with a pre-set regulatory standard library to obtain a comparison result, which includes conformity identifiers and correction identifiers.
[0009] The draft code of conduct is iteratively adjusted based on the correction item identifiers in the comparison results, and the final code of conduct regulatory instructions are generated and sent to the regulatory execution terminal.
[0010] In another aspect, embodiments of the present invention also provide a large-scale model system for the supervision of behavioral norms collaboratively defined by natural language and images, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the machine-readable storage medium to implement the above-mentioned method.
[0011] Based on the above, this embodiment of the invention constructs a set of association mapping rules between natural language descriptions and standardized schematic diagrams, clarifying the correspondence and collaborative constraints between text semantic elements and image visual elements. According to this association mapping rule set, the received behavioral description text and scene schematic diagrams of the scene to be regulated are bidirectionally converted to obtain a text-image collaborative feature set. This achieves the fusion and complementarity of natural language and image information, enabling a more comprehensive and accurate expression of the content of behavioral norms. Based on the text-image collaborative feature set, a pre-trained norm generation model is invoked to extract behavioral norm elements, generating a draft behavioral norm containing behavioral boundary conditions and compliance judgment benchmarks. This improves the accuracy and efficiency of behavioral norm formulation and avoids potential omissions and deviations in manual formulation. A consistency comparison between the draft behavioral norm and a pre-set regulatory standard library allows for timely identification of problems in the draft. The draft behavioral norm is then iteratively adjusted using correction item identifiers, ultimately generating a regulatory instruction that meets regulatory requirements. This ensures the scientific and rational nature of the behavioral norm and greatly improves the accuracy and reliability of behavioral norm regulation. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the execution flow of a large-scale behavioral norm supervision model system that uses natural language and images to collaboratively define behavior norms, provided in an embodiment of the present invention.
[0013] Figure 2 This is a schematic diagram of exemplary hardware and software components of the natural language and image collaborative definition behavior regulation supervision system provided in this embodiment of the invention. Detailed Implementation
[0014] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1This is a flowchart illustrating a large-scale behavioral norm supervision model system that uses natural language and images to collaboratively define behavioral norms, according to one embodiment of the present invention. The following is a detailed description of this large-scale behavioral norm supervision model system that uses natural language and images to collaboratively define behavioral norms.
[0015] Step S110: Construct a set of association mapping rules between natural language descriptions and standard schematic images. The set of association mapping rules includes the correspondence between text semantic elements and image visual elements and collaborative constraints.
[0016] Step S111: Collect natural language description samples and corresponding specification illustration samples from historical behavior specification documents to form a sample set. The natural language description samples include the behavior subject, behavior action and scene limiting words. The specification illustration samples include the behavior subject image, action posture image and scene background image.
[0017] In this embodiment, when collecting samples, representative content needs to be selected from a large number of historical behavioral norms documents. The natural language description samples should cover behavioral requirements in different situations, such as "Within a specific area, relevant personnel must operate the equipment according to the specified procedures, check the equipment status before operation, not leave without authorization during operation, and record the operation results after operation." Here, "relevant personnel" is the subject of the behavior, "operate the equipment according to the specified procedures," "check the equipment status," "leave without authorization," and "record the operation results" are behavioral actions, and "within a specific area" is a scenario limiting term.
[0018] The corresponding sample illustrations must precisely match these descriptions, including images of personnel inspecting equipment in a specific area, operating equipment according to procedures, remaining at their posts during operation, and recording results after operation. These images must clearly show the subject, actions, and background. For potentially privacy-sensitive data, such as facial information in images, pixel blurring technology can be used, and personally identifiable information in the text will be anonymized to ensure information security.
[0019] Step S112: Perform semantic segmentation on the natural language description sample, extract the part-of-speech tag and semantic role label of each segmentation unit to obtain text semantic elements, which include subject referential identifiers, action verb identifiers and scene modification identifiers.
[0020] Taking the example of "In a specific area, relevant personnel need to operate the equipment according to the specified process. Before the operation, the equipment status needs to be checked. During the operation, they are not allowed to leave without permission. After the operation, the operation results need to be recorded", the semantic word segmentation is processed, and the word segmentation units are "in", "in a specific area", ",", "relevant personnel", "need to", "according to the specified process", "operate", "equipment", ",", "before the operation", "need to", "check", "equipment status", ",", "during the operation", "are not allowed to", "leave without permission", ",", "after the operation", "need to", "record", "operation results".
[0021] Add词性tags to each word segmentation unit. "In" is a preposition, "in a specific area" is a locative noun phrase, "relevant personnel" is a noun phrase, "need to" is an auxiliary verb, "according to the specified process" is a prepositional phrase, "operate" is a verb, "equipment" is a noun, "before the operation" is a time noun phrase, "check" is a verb, "equipment status" is a noun phrase, "during the operation" is a time noun phrase, "are not allowed to" is an auxiliary verb, "leave without permission" is a verb phrase, "after the operation" is a time noun phrase, "record" is a verb, "operation results" is a noun phrase.
[0022] When performing semantic role annotation, "relevant personnel" is the agent of all actions, that is, the subject; "operate the equipment", "check the equipment status", "leave without permission", "record the operation results" are specific action verbs; "in a specific area" is the scene where the action occurs, and "before the operation", "during the operation", "after the operation" are the time scene limitations for the occurrence of the action. Thus, the semantic elements of the text are extracted. The subject reference identifier is "relevant personnel", the action verb identifiers are "operate", "check", "leave without permission", "record", and the scene modifier identifiers are "in a specific area", "before the operation", "during the operation", "after the operation".
[0023] Step S113: Perform visual feature extraction processing on the standardized schematic picture sample, identify the subject contour area, action posture key points, and scene background area in the picture, and obtain the picture visual elements. The picture visual elements include subject contour features, posture key point coordinates, and background area texture features.
[0024] Perform visual feature extraction on the standardized schematic picture sample corresponding to the above natural language description sample. Taking the picture showing relevant personnel checking the equipment in a specific area as an example, the subject contour area is identified through an edge detection algorithm. The subject contour area presents the overall external contour of the relevant personnel, including the general shape of the body and the distribution range of the limbs.
[0025] Key points of the action posture include the position of the hand touching the device, the angle of the head facing the device, and the posture of the body standing. The coordinates of these key points are represented by the pixel coordinate system of the image. Each key point has its corresponding horizontal and vertical coordinate values to accurately reflect its position in the image.
[0026] The scene background area is the environment within a specific area, including the appearance of the equipment, the layout of the surrounding facilities, etc. Texture features of the background area are extracted through texture analysis algorithms. These features include the density, direction, and color distribution of the texture, in order to distinguish different scene backgrounds.
[0027] Step S114: Call the bidirectional attention mechanism model to calculate the correlation between the text semantic elements and the image visual elements, and generate a semantic visual correlation matrix. The elements in the semantic visual correlation matrix represent the matching strength between the text semantic elements and the image visual elements.
[0028] The extracted text semantic elements and image visual elements are input into a bidirectional attention mechanism model. The model first encodes the subject reference markers, action verb markers, and scene modification markers in the text semantic elements, and simultaneously encodes the subject outline features, pose key point coordinates, and background region texture features in the image visual elements.
[0029] After encoding, the model determines the correlation between text semantic elements and image visual elements by calculating attention weights. For each semantic element in the text, the model pays attention to the related visual elements in the image; conversely, for each visual element in the image, the model also pays attention to the related semantic elements in the text.
[0030] By employing the aforementioned two-way attention method, the matching strength between each text semantic element and each image visual element is calculated. These matching strength values constitute a semantic-visual association matrix. For example, the matching strength between the action verb tag "check device status" in the text and the coordinates of the key points of the hand touching the device in the image will correspond to a high value in the matrix.
[0031] Step S115: Based on the semantic visual association matrix, filter text visual element pairs that meet the preset conditions for matching strength, and construct the correspondence between text semantic elements and image visual elements by combining the behavioral norm logical relationship.
[0032] Set a preset condition for matching strength, such as matching strength greater than a certain threshold. Select text semantic elements and image visual elements that meet the preset condition from the semantic visual association matrix to form text visual element pairs.
[0033] By combining the logical relationships of behavioral norms, analyze the inherent connections between these elements. For example, the subject identifier "relevant personnel" should correspond to the outline features of the subject in the image; the action verb identifier "operating equipment" should correspond to the key coordinates of the posture of the relevant personnel operating the equipment in the image; and the scene embellishment identifier "within a specific area" should correspond to the texture features of the background area in the image.
[0034] Based on these logical relationships, a clear correspondence is established between text semantic elements and image visual elements, ensuring that each text semantic element can find a matching image visual element, and vice versa.
[0035] Step S116: Based on the correspondence, set collaborative constraints for text description and image display. The collaborative constraints include semantic consistency constraints, spatiotemporal correlation constraints, and expression integrity constraints. Integrate the correspondence and the collaborative constraints to form an association mapping rule set.
[0036] Based on the established correspondence, set collaborative constraints. Semantic consistency constraints require that the semantics of the text description and the content displayed in the image be consistent in meaning, and there should be no contradictions. For example, if the text description is "No one may leave without authorization during the operation," the image cannot show a scene where relevant personnel leave their operating positions.
[0037] Spatiotemporal correlation constraints require that the temporal and spatial information involved in the text description match the temporal and spatial scenes presented in the images. For example, if the text specifies the time scene as "before operation," the corresponding image should be of the relevant personnel in a state before they begin operating the equipment.
[0038] The completeness constraint requires that both text descriptions and images fully express the content of the behavioral norm, without any missing information. For example, for the action of "recording the operation result," the text should explain the recording requirements, and the image should show the recording process or result; only by combining the two can the norm be fully expressed. The established correspondences and set collaborative constraints are integrated to form a set of association mapping rules.
[0039] Step S120: Receive the behavioral description text and scene illustration of the scene to be monitored, and perform bidirectional conversion processing on the behavioral description text and the scene illustration according to the association mapping rule set to obtain a text-image collaborative feature set.
[0040] In practical applications, the system receives behavioral description text and scene illustrations from the scenario to be monitored. This information may be input by relevant personnel or collected by devices. The system then performs bidirectional conversion on these two types of information according to a set of association mapping rules, converting the text information into corresponding visual features and the visual information into corresponding semantic features, ultimately fusing them to obtain a collaborative feature set.
[0041] Step S121: Receive the behavioral description text of the scenario to be monitored, perform syntactic structure analysis on the behavioral description text, identify the subject-verb-object structure and modifying components, and extract the semantic elements of the text to be processed. The semantic elements of the text to be processed include the identifier of the subject to be monitored, the description of the action to be monitored, and the description of the scenario to be monitored.
[0042] Receive behavioral description text for the scenario to be regulated, such as "A certain group should move in a predetermined manner in a specific place, avoid specific objects when moving, and be in a designated position after moving."
[0043] Syntactic structure analysis was performed on the text, and the subject-verb-object structure was identified using a syntactic analysis algorithm. The subject is “a certain group”, the verb is “move”, and the modifiers are “in a specific place”, “in a predetermined manner”, “avoiding specific objects”, and “in a designated position”.
[0044] Based on the analysis results, semantic elements of the text to be processed are extracted. The subject to be regulated is identified as "a certain group", the action to be regulated is described as "moving in a predetermined manner", "avoiding a specific object", "being in a designated location", and the scenario to be regulated is described as "in a specific place", "while moving", and "after moving".
[0045] Step S122: Receive a scene illustration of the scene to be monitored, perform image segmentation processing on the scene illustration to separate the main body area, action area and background area, and extract the visual elements of the image to be processed. The visual elements of the image to be processed include the outline of the subject to be monitored, the action posture to be monitored and the background features of the scene to be monitored.
[0046] Receive a scenario illustration of a scene to be monitored, which shows the state of a certain group in a specific place.
[0047] Image segmentation is performed on the image, using a region growing algorithm to divide the image into different regions. The subject region is the area where a group is located, the action region is the area where the group performs actions such as movement, and the background region is the environmental area of a specific location, such as the ground and surrounding objects.
[0048] Visual elements of the image to be processed are extracted from the segmented regions. The outline of the subject to be monitored is the overall shape of a group; the action posture to be monitored is the limb posture of a group when moving, such as the direction of limb extension and position changes; the background features of the scene to be monitored are the background environmental features of a specific place, such as the texture of the ground, the shape and distribution of surrounding objects, etc.
[0049] Step S123: Based on the correspondence in the association mapping rule set, the semantic elements of the text to be processed are converted into corresponding visual feature descriptions to generate a text-to-visual feature set, which includes subject visualization parameters, action visualization parameters, and scene visualization parameters.
[0050] Based on the correspondence in the association mapping rule set, the semantic elements of the text to be processed are transformed. The identifier of the subject to be regulated, "a certain group," is converted into subject visualization parameters, which include the approximate size of the group and the visual representation of its physical characteristics.
[0051] The descriptions of the actions to be monitored, such as "moving in a predetermined manner," "avoiding specific objects," and "being in a designated position," are converted into visual parameters, such as the range of movement speed, changes in direction, the range of distances when avoiding objects, and the coordinate range of the designated position.
[0052] The descriptions of the scenarios to be regulated, such as "in a specific location," "when moving," and "after moving," are converted into scene visualization parameters, such as the spatial range of the location and the changes in the visual characteristics of the scene at different times.
[0053] These parameters together constitute a set of text-to-visual features.
[0054] Step S124: Based on the correspondence in the association mapping rule set, the visual elements of the image to be processed are converted into corresponding semantic feature descriptions to generate a visual-to-text feature set, which includes subject semantic tags, action semantic tags, and scene semantic tags.
[0055] Based on the correspondence in the association mapping rule set, the visual elements of the image to be processed are converted into semantic feature descriptions. The semantic label corresponding to the outline of the subject to be regulated is semantic information such as the type and composition of "a certain group".
[0056] The semantic tags corresponding to the actions and postures to be monitored are semantic descriptions such as "movement method", "behavior of avoiding objects", and "state of being in a specified position".
[0057] The semantic tags corresponding to the background features of the scenarios to be regulated are semantic information such as "type of a specific place", "environmental characteristics within the place", and "scenarios at different points in time".
[0058] These semantic tags constitute a set of visual-to-text features.
[0059] Step S125: Call the feature fusion module to perform collaborative verification processing on the text-to-visual feature set and the visual-to-text feature set, remove conflicting features that violate collaborative constraints, and obtain a preliminary collaborative feature set.
[0060] Invoke the feature fusion module by inputting the text-to-visual feature set and the visual-to-text feature set. The module will then validate the features in both sets based on the collaborative constraints in the association mapping rule set.
[0061] Check semantic consistency to see if the semantics expressed by the text-to-visual features are consistent with those expressed by the visual-to-text features. For example, does the statement "a group moves relatively fast" in the text-to-visual features conflict with the statement "a group moves slowly" in the visual-to-text features?
[0062] Examine the spatiotemporal correlation to confirm whether the temporal and spatial parameters in the text-to-visual features match the temporal and spatial semantics in the visual-to-text features. For example, is there a contradiction between "the movement occurred during a certain period" in the text-to-visual features and "the movement occurred during another period" in the visual-to-text features?
[0063] Check the completeness of the description to ensure that the two feature sets combined can fully express the information of the scenario to be regulated, without omitting any key content.
[0064] For conflicting features that violate the collaborative constraints, such as features with semantic contradictions, spatiotemporal mismatches, or missing information, these features are removed, and features that meet the conditions are retained to form a preliminary collaborative feature set.
[0065] Step S126: Perform feature dimension adaptation processing on the preliminary collaborative feature set to make the representation dimensions of text-derived features and visual-derived features consistent, and generate a text-image collaborative feature set containing association weight parameters.
[0066] The text-derived features and visual-derived features in the initial collaborative feature set may differ in their representational dimensions. For example, text-derived features may be represented by the number of semantic labels, while visual-derived features may be represented by the numerical value of parameters.
[0067] These features undergo dimensionality adaptation processing, using feature transformation algorithms to convert text-derived features and visual-derived features to the same dimensional space. For example, the semantic labels of text-derived features are converted into corresponding vector representations, ensuring they maintain dimensionality consistency with the parameter vectors of visual-derived features.
[0068] During the adaptation process, a correlation weight parameter is assigned to each feature based on the degree of correlation between features; the higher the correlation, the larger the weight parameter. The adapted features and their corresponding correlation weight parameters are then integrated to generate a text-image collaborative feature set.
[0069] Step S130: Based on the text image collaborative feature set, call the pre-trained specification generation model to perform behavior specification element extraction operation, and generate a draft behavior specification containing behavior boundary conditions and compliance judgment benchmarks.
[0070] By utilizing a text-image collaborative feature set, a pre-trained norm generation model is invoked. This norm generation model, trained on a large number of samples, can extract key elements of behavioral norms from the collaborative features, such as the scope of the behavioral subject, the boundary conditions of the behavior, and the benchmarks for judging whether the behavior is compliant, thereby generating a draft of behavioral norms.
[0071] Step S131: Input the text image collaborative feature set into the feature encoding layer of the pre-trained canonical generation model, and model the correlation between features through a multi-head self-attention mechanism to generate encoded feature vectors.
[0072] The text-image collaborative feature set is input into the feature encoding layer of the standardized generative model. The feature encoding layer first preprocesses the input features, converting different types of features into a format that the model can process.
[0073] Then, the relationships between features are modeled using a multi-head self-attention mechanism. This mechanism simultaneously calculates attention weights between features from multiple different perspectives, with each perspective focusing on different relationships between features.
[0074] The above method can capture complex correlation information in the feature set. After integrating this information, an encoded feature vector is generated, which contains key information and correlations in the collaborative feature set.
[0075] Step S1311: Perform feature splitting on the text image collaborative feature set to obtain a text-derived feature subset and a visual-derived feature subset. The text-derived feature subset includes semantic label features and association weight features, and the visual-derived feature subset includes visual parameter features and matching strength features.
[0076] The text-image collaborative feature set is split into text-derived feature subsets and visual-derived feature subsets according to the source and type of the features.
[0077] The semantic label features in the text-derived feature subset are various semantic labels extracted from text information, such as subject semantic labels and action semantic labels; the association weight features are the weight parameters assigned to each text feature during the feature fusion process.
[0078] The visual parameter features in the visual derived feature subset are various visual parameters extracted from image information, such as subject visualization parameters and action visualization parameters; the matching strength features are parameters representing the degree of matching between text features and visual features.
[0079] Step S1312: Input the text-derived feature subset and the visual-derived feature subset into the embedding module of the feature encoding layer, perform vector transformation processing respectively, and generate text feature vector sequence and visual feature vector sequence, wherein the dimensions of the text feature vector sequence and the visual feature vector sequence are consistent.
[0080] The text-derived feature subset and the visual-derived feature subset are input into the embedding module of the feature encoding layer. The embedding module performs vector transformation on each feature, converting text-derived features such as semantic label features and association weight features into text feature vectors, and visual-derived features such as visual parameter features and matching strength features into visual feature vectors.
[0081] During the transformation process, it is ensured that the generated text feature vector sequence and visual feature vector sequence have the same dimension to facilitate subsequent feature interaction processing. For example, both the text feature vector sequence and the visual feature vector sequence use vectors of the same length to represent each feature.
[0082] Step S1313: Call the multi-head self-attention mechanism to perform cross-attention calculation on the text feature vector sequence and the visual feature vector sequence to generate a text-visual interaction attention weight matrix. The text-visual interaction attention weight matrix is used to represent the correlation strength between text features and visual features.
[0083] A multi-head self-attention mechanism is invoked, which consists of multiple parallel attention heads. Each attention head processes the text feature vector sequence and the visual feature vector sequence separately, calculating the cross-attention weights between the text features and the visual features.
[0084] Specifically, each attention head performs a linear transformation on the text feature vector sequence and the visual feature vector sequence to obtain a query vector, a key vector, and a value vector. For the text feature vector sequence, the linear transformation yields a text query vector, a text key vector, and a text value vector; for the visual feature vector sequence, the linear transformation yields a visual query vector, a visual key vector, and a visual value vector.
[0085] Next, each attention head calculates the similarity between the text query vector and the visual key vector. These similarities are then normalized using the softmax function to obtain the cross-attention weights, which reflect the strength of the association between text features and visual features.
[0086] After multiple attention heads have been computed, the cross-attention weights obtained from each attention head are concatenated to generate a text-visual interaction attention weight matrix. Each element in this text-visual interaction attention weight matrix represents the correlation strength between a specific text feature and a specific visual feature.
[0087] Step S1314: Based on the text visual interaction attention weight matrix, perform weighted fusion processing on the text feature vector sequence and the visual feature vector sequence to generate an interactive fusion feature vector sequence.
[0088] We utilize a text-visual interaction attention weight matrix to perform weighted fusion of text feature vector sequences and visual feature vector sequences. For each text feature vector, the corresponding visual feature vectors are weighted and summed based on their association strength with each visual feature vector to obtain the fused vector corresponding to that text feature vector; similarly, for each visual feature vector, the corresponding text feature vectors are weighted and summed based on their association strength with each text feature vector to obtain the fused vector corresponding to that visual feature vector.
[0089] These fused vectors are arranged in their original sequence order to generate an interactive fused feature vector sequence. This interactive fused feature vector sequence contains information from both textual and visual features, and reflects the interaction between the two.
[0090] Step S1315: The interactive fused feature vector sequence is subjected to nonlinear transformation processing through the feedforward neural network of the feature encoding layer to enhance the expressive power of the features and generate an intermediate feature vector sequence.
[0091] The interactively fused feature vector sequence is input into a feedforward neural network of the feature encoding layer. This feedforward neural network contains multiple fully connected layers, each of which uses a non-linear activation function to process the input feature vector.
[0092] First, the first fully connected layer performs a linear transformation on the interactive fusion feature vector sequence, mapping the features to a higher-dimensional space. Then, a non-linear activation function is used to perform a non-linear transformation on the transformed features, introducing non-linear factors to enhance the expressive power of the features. Next, the second fully connected layer maps the high-dimensional features back to the original dimensional space.
[0093] After processing by the feedforward neural network, an intermediate feature vector sequence is generated. This intermediate feature vector sequence has a stronger feature expression ability while retaining the original key information.
[0094] Step S1316: Perform layer equalization on the intermediate feature vector sequence to obtain encoded feature vectors with a unified representation.
[0095] The individual feature vectors in the intermediate feature vector sequence may differ in numerical range, distribution, etc. To give these feature vectors a unified representation, layer equalization is required.
[0096] Layer equalization employs a standardization operation, transforming each eigenvector in the intermediate eigenvector sequence to have a mean of zero and a variance of one. Specifically, the mean and variance of the intermediate eigenvector sequence are calculated, and then the mean is subtracted from each eigenvector, and the result is divided by the square root of the variance to obtain the standardized eigenvector.
[0097] After layer equalization, a coded feature vector with a unified representation is obtained, which can be better processed by subsequent feature extraction layers.
[0098] Step S132: Receive the encoded feature vector through the feature extraction layer of the specification generation model, perform the behavior subject definition operation, determine the subject scope and subject attribute features to which the behavior specification applies, and generate the subject definition result.
[0099] After receiving the encoded feature vector, the feature extraction layer of the normative generation model first focuses on defining the subject of the behavior. By analyzing and extracting the subject-related information from the encoded feature vector, it determines the scope of subjects to which the behavioral norm applies, as well as the attribute characteristics of these subjects, ultimately forming the subject definition result.
[0100] Step S1321: Receive the encoded feature vector through the feature extraction layer of the standard generation model, perform matching processing through the preset subject feature template, identify the feature components related to the behavioral subject in the encoded feature vector, and obtain the subject feature vector.
[0101] After receiving the encoded feature vector, the feature extraction layer calls the preset subject feature template. The subject feature template is pre-defined based on a large amount of historical data and behavioral norms, and contains typical feature patterns related to the behavioral subject.
[0102] The encoded feature vector is matched with the subject feature template. By calculating the similarity between the two, the feature components in the encoded feature vector that are related to the subject are identified. These feature components together constitute the subject feature vector, which reflects the information related to the subject.
[0103] Step S1322: Analyze the type attributes of the behavioral subject based on the subject feature vector, determine the category identifier and category feature description of the subject, and generate the subject type definition result.
[0104] A thorough analysis of the subject's feature vector is conducted to extract feature information that reflects the subject's type attributes. This information may include the subject's composition, functional roles, etc.
[0105] Based on these characteristics and referring to pre-defined subject type classification standards, the category identifier to which the subject belongs is determined. For example, the category identifier may include "individual," "group," or "organization." Simultaneously, the characteristics of this category are described, such as the size of the group or the structural characteristics of the organization, generating the subject type definition result.
[0106] Step S1323: Based on the subject type definition result, extract the constraint information related to the subject range from the encoded feature vector, determine the number range, identity limitation and qualification conditions of subjects, and generate the subject range definition result.
[0107] Based on the subject type definition, further constraint information related to the subject scope is extracted from the encoded feature vector. This information may involve the number of subjects, identity requirements, and required qualifications.
[0108] Analyze this constraint information to determine the range of subjects, such as a single subject, multiple subjects, or a group of a specific size; clarify the identity limitations of the subjects, such as specific occupations or identity markers; determine the qualification requirements of the subjects, such as whether they need to possess certain skills or pass certain certifications. Integrate this information to generate the subject scope definition result.
[0109] Step S1324: Analyze the attribute feature information in the subject feature vector, identify the subject's inherent attribute features and dynamic attribute features, wherein the inherent attribute features include the subject's static feature parameters, and the dynamic attribute features include the subject's state feature parameters.
[0110] The attribute feature information in the subject's feature vector is analyzed to distinguish between inherent attribute features and dynamic attribute features. Inherent attribute features are relatively stable feature parameters that are inherent to the subject itself, such as static feature parameters like the subject's physical features and structural features.
[0111] Dynamic attributes are the characteristic parameters exhibited by an entity in different states, such as the entity's activity state and its adaptation to its environment. By distinguishing between these, we can gain a more comprehensive understanding of the entity's attribute characteristics.
[0112] Step S1325: Integrate the subject type definition result, the subject scope definition result, and the attribute feature information, and generate a subject definition result containing subject identifier, type description, scope limitation, and attribute parameters according to the subject definition rules.
[0113] The results of subject type definition, subject scope definition, and attribute feature information are integrated. The subject definition rules are pre-defined and used to standardize the generation format and content requirements of the subject definition results.
[0114] According to the subject definition rules, the integrated information is organized into a subject definition result that includes subject identifier (used to uniquely identify the subject), type description (detailed description of the subject type), scope limitation (specific definition of the subject's scope), and attribute parameters (specific parameters of the subject's inherent and dynamic attributes).
[0115] Step S133: Based on the subject definition result, analyze the spatiotemporal range of the behavior through the element extraction layer, determine the time interval limit and spatial area limit of the behavior, and generate the behavior boundary conditions.
[0116] Based on the subject definition results, the element extraction layer further analyzes the spatiotemporal limitations of the behavior. By extracting and analyzing time- and space-related information from the encoded feature vector, the time interval and spatial region where the behavior occurs are determined, thereby generating the behavior boundary conditions.
[0117] Specifically, by analyzing the time-related feature components in the encoded feature vector, information such as the start time, end time, and duration of the behavior can be identified, and time intervals can be defined, such as "a specific period on a weekday" or "a continuous number of hours".
[0118] At the same time, by analyzing the spatially related feature components in the encoded feature vector, information such as the specific location and area where the behavior occurs can be identified, and spatial area limits can be determined, such as "a certain area within a specific building" or "a certain geographical area".
[0119] By integrating time interval constraints and spatial region constraints, behavioral boundary conditions are generated, clarifying the spatiotemporal scope in which the behavior occurs.
[0120] Step S134: Combining compliance clues in text and image collaborative features, identify permitted and prohibited forms of behavior and construct compliance judgment benchmarks for behavior and actions. The compliance judgment benchmarks include positive judgment indicators and negative judgment indicators.
[0121] Extract compliance-related clues from the text-image collaborative feature set. These clues may include descriptions of permitted or prohibited behaviors in the text, examples of compliant or non-compliant behaviors shown in the image, etc.
[0122] Based on these compliance clues, identify the permissible forms of behavior, i.e., which behaviors comply with the requirements of the regulations; at the same time, identify the prohibited forms of behavior, i.e., which behaviors do not comply with the requirements of the regulations.
[0123] Compliance judgment benchmarks are constructed based on permitted and prohibited forms. Positive judgment indicators are used to determine whether an action conforms to a permitted form, including the specific manifestation and implementation of the action; negative judgment indicators are used to determine whether an action falls under a prohibited form, similarly including the specific manifestation and implementation of the action. Through these positive and negative judgment indicators, the compliance of an action can be clearly determined.
[0124] Step S135: Integrate the subject definition results, the behavioral boundary conditions, and the compliance judgment criteria, and arrange them according to the preset standard text structure to generate a draft of behavioral norms. The draft of behavioral norms includes clause numbers, clause content, and corresponding illustration indexes.
[0125] The system integrates the definition of the subject, the boundary conditions of behavior, and the compliance judgment criteria. The pre-defined structure of the normative text is based on common behavioral norm text formats, including the order of clauses and the organization of content.
[0126] The integrated information is arranged according to a pre-defined standardized text structure. Each clause is assigned a unique clause number and its content is clearly defined, including information such as the subject, boundaries of action, and compliance determination. Simultaneously, a corresponding illustrative image index is added to each clause, allowing users to link to relevant illustrative images.
[0127] After editing and processing, a draft code of conduct is generated, which fully presents all the elements of the code of conduct.
[0128] Step S140: Perform a consistency comparison process between the draft code of conduct and the preset regulatory standard library to obtain the comparison result, which includes compliance item identifiers and correction item identifiers.
[0129] To ensure the accuracy and compliance of the draft code of conduct, it needs to be compared with a pre-established regulatory standard library. By comparing the content of the draft with that in the standard library, the similarities and differences between the two are identified and marked as conforming items and correction items, respectively, forming the code comparison results.
[0130] Step S141: Obtain a preset regulatory standard library, which contains historically valid behavioral norms standard texts and corresponding standard illustration sets. Each standard text contains standard clauses, scope of application, and judgment basis.
[0131] Access a pre-defined regulatory standards library, which has been accumulated and organized over a long period and contains a large number of historically valid code-of-conduct texts. Each code-of-conduct text details the standard clauses, the scope of application of the clauses, and the basis for determining whether a behavior complies with the clauses.
[0132] In addition, the standard library also contains a set of standard illustration images corresponding to the standard text. These images visually demonstrate the behavioral norms described in the standard clauses, helping to better understand the standard content.
[0133] Step S142: The draft code of conduct is split into clauses to obtain multiple independent draft clause units. Each draft clause unit contains clause content, corresponding illustrations, and applicable conditions.
[0134] The draft code of conduct was broken down into clauses, and the draft was divided into multiple independent draft clause units according to the logical relationship and content independence of the clauses.
[0135] Each draft clause unit contains complete information: the clause content is the specific provision for the behavior stipulated in the clause; the corresponding illustration is an image used to explain the content of the clause; and the applicable conditions are the specific circumstances and prerequisites under which the clause applies.
[0136] Step S143: The standard text in the regulatory standard library is split into clauses to obtain multiple independent standard clause units. Each standard clause unit contains the standard clause content, the corresponding standard image, and the standard application conditions.
[0137] Using the same method as splitting the draft code of conduct, the standard texts in the regulatory standard library were split into clauses to obtain multiple independent standard clause units.
[0138] Each standard clause unit contains the standard clause content, that is, the specific provisions of the standard for behavior; the corresponding standard image, used to visually demonstrate the standard clause; and the standard application conditions, which clarify the circumstances under which the standard clause applies.
[0139] Step S144: Perform semantic matching processing on the draft clause content of the draft clause unit and the standard clause content of the standard clause unit to generate a text similarity score.
[0140] Semantic matching is performed between the draft clauses in the draft clause unit and the standard clauses in the standard clause unit. Natural language processing technology is used to analyze the semantic similarity between the two, generating a text similarity score that reflects the closeness in meaning between the two texts.
[0141] Step S1441: Perform word segmentation and词性 tagging on the draft clause content of the draft clause unit. After removing stop words, an effective word sequence is obtained. Perform word vector conversion on each effective word to generate a draft word vector sequence.
[0142] Perform word segmentation on the draft clause content to split the text into individual words. Then, perform词性 tagging on each word to clarify its词性, such as noun, verb, adjective, etc.
[0143] Remove stop words from the text. These words usually have no actual semantic meaning or have little impact on the semantics, such as "的", "在", "和", etc. The remaining words form an effective word sequence.
[0144] Perform word vector conversion on each effective word in the effective word sequence to convert the words into corresponding vector representations. These vectors can reflect the semantic information of the words. Multiple word vectors are arranged in the original word order to generate a draft word vector sequence.
[0145] Step S1442: Perform word segmentation and词性 tagging on the standard clause content of the standard clause unit. After removing stop words, a standard word sequence is obtained. Perform word vector conversion on each standard word to generate a standard word vector sequence.
[0146] Use the same method as for processing the draft clause content to perform word segmentation and词性 tagging on the standard clause content. After removing stop words, a standard word sequence is obtained.
[0147] Perform word vector conversion on each standard word in the standard word sequence to generate a standard word vector sequence. This standard word vector sequence has the same vector dimension as the draft word vector sequence.
[0148] Step S1443: Calculate the cosine similarity of the corresponding position word vectors in the draft word vector sequence and the standard word vector sequence to obtain a word-level similarity matrix.
[0149] Calculate the cosine similarity of the corresponding position word vectors in the draft word vector sequence and the standard word vector sequence. Cosine similarity is an index to measure the similarity of the directions of two vectors. The larger its value, the closer the directions of the two vectors are, and the more similar the corresponding word semantics are.
[0150] Arrange the cosine similarity values of all corresponding position word vectors in matrix form to obtain a word-level similarity matrix. Each element in this word-level similarity matrix represents the semantic similarity of the corresponding position words in the two word sequences.
[0151] Step S1444: Perform an optimal path search on the word-level similarity matrix to determine the best matching path between the two word sequences and generate a path matching score.
[0152] A dynamic programming algorithm is used to search for the optimal path in the word-level similarity matrix. The optimal path is the path from the top left to the bottom right in the word-level similarity matrix, where the sum of the cosine similarities is maximized, representing the most matching word correspondence between two word sequences.
[0153] By calculating the sum of cosine similarities along the optimal path and normalizing them, a path matching score is generated, which reflects the overall matching degree of the two word sequences.
[0154] Step S1445: Extract the syntactic structure features of the draft clause content and the standard clause content, calculate the edit distance of the syntactic tree, and obtain the structural similarity score.
[0155] Syntactic analysis is performed on the draft and standard clauses to generate their respective syntax trees. The syntax trees reflect the syntactic relationships between words in a sentence, such as subject-predicate and verb-object relationships.
[0156] Extract structural features of the syntax tree, such as tree depth, number and type of nodes, and connections between nodes. Calculate the edit distance between two syntax trees, which is the minimum number of editing operations (such as insertion, deletion, and replacement of nodes) required to transform one syntax tree into another.
[0157] Based on the edit distance, a structural similarity score is obtained. The smaller the edit distance, the higher the structural similarity score, indicating that the syntactic structures of the two texts are more similar.
[0158] Step S1446: The path matching score and the structural similarity score are weighted and summed according to preset weights to obtain the preliminary text similarity.
[0159] The weights of the preset path matching score and structural similarity score are pre-set based on their importance in semantic matching.
[0160] The path matching score is multiplied by its corresponding weight, and the structural similarity score is multiplied by its corresponding weight. The two products are then added together to obtain the preliminary text similarity score, which comprehensively reflects the degree of similarity between the two texts in terms of word semantics and syntactic structure.
[0161] Step S1447: Perform interval mapping processing on the preliminary text similarity, mapping it to a preset numerical interval to generate a text similarity score.
[0162] In this embodiment, a preset numerical range, such as between 0 and 1, is used to represent the level of text similarity. The initial text similarity is processed by range mapping, transforming its value into the preset numerical range through linear transformation or other methods. The transformed value is the text similarity score; the closer the score is to 1, the more semantically similar the draft clauses are to the standard clauses; the closer the score is to 0, the greater the semantic difference between the two.
[0163] Step S145: Perform visual matching processing on the schematic diagrams corresponding to the draft clause units and the standard images corresponding to the standard clause units to generate image similarity scores.
[0164] Visual matching was performed on the illustrative images corresponding to the draft clause units and the standard images corresponding to the standard clause units. Image processing techniques were used to analyze the similarity of the two images in terms of visual features, generating an image similarity score. This score reflects the degree of similarity between the two images in terms of content and visual presentation.
[0165] Specifically, visual features can be extracted from two images, such as color histograms, texture features, and shape features. Then, the similarity between these visual features is calculated, for example, by calculating the intersection distance of the color histograms and the Euclidean distance of the texture features. These similarity values are then combined to obtain an image similarity score.
[0166] Step S146: Calculate the comprehensive matching degree based on the text similarity score and the image similarity score. When the comprehensive matching degree reaches a preset threshold, mark the draft clause unit as a compliant item and add a compliant item identifier.
[0167] Assign appropriate weights to text similarity scores and image similarity scores, with the weights determined based on the importance of the text and image in the matching process.
[0168] Multiply the text similarity score by its weight, multiply the image similarity score by its weight, and then add the two results together to obtain the overall matching degree.
[0169] A comprehensive matching threshold is preset. When the calculated comprehensive matching degree reaches the threshold, it means that the draft clause unit and the standard clause unit are relatively consistent in content and visual appearance. The draft clause unit is marked as a conforming item and a conforming item identifier is added. The identifier may contain the information of the matching standard clause unit.
[0170] Step S147: When the overall matching degree does not reach the preset threshold, analyze the location and type of the difference points, mark the draft clause unit as a correction item and add a correction item identifier. The correction item identifier includes a description of the difference type and a reference standard index.
[0171] In this embodiment, when the overall matching degree does not reach the preset threshold, it indicates that there are significant differences between the draft clause unit and the standard clause unit, and the differences need to be analyzed. First, determine the specific location of the differences: whether it is a difference in text content, a difference in illustrations, or both.
[0172] Regarding the differences in text content, we can further analyze the types of differences. They may be deviations in semantic expression, such as inaccurate descriptions of actions; they may also be differences in applicable conditions, such as discrepancies between the scope of application of draft clauses and standard clauses; or they may be differences in the basis for compliance judgment, such as inconsistencies in the indicators for judging compliance.
[0173] Regarding the differences in the illustrations, we can also analyze the types of differences. It could be that the behavior shown in the illustration does not match the standard illustration, such as a deviation in the posture; or that the background of the illustration differs from the standard illustration, such as differences in the details of the background environment.
[0174] After identifying the location and type of the differences, mark the draft clause unit as a revision item and add a revision item identifier. The difference type description in the revision item identifier details the specific type and manifestation of the difference, while the reference standard index points to the corresponding standard clause unit in the regulatory standard library.
[0175] Step S148: Integrate the conformity identifiers and amendment identifiers of all draft clause units to generate specification comparison results.
[0176] After comparing all draft clause units, the compliance or amendment identifiers for each unit are summarized. These identifiers are then integrated according to the order of the draft clause units to form a complete regulatory comparison result. This comparison result clearly shows which clauses in the draft code of conduct meet regulatory standards and which clauses require amendment.
[0177] Step S150: Iteratively adjust the draft code of conduct according to the correction item identifier in the comparison result of the code of conduct, generate the final code of conduct supervision instruction and send it to the supervision execution terminal.
[0178] Based on the correction markers in the comparison results, the draft code of conduct was modified and improved in a targeted manner. After multiple iterations and adjustments to ensure that the draft met regulatory standards, it was converted into an instruction format recognizable by the regulatory execution terminal and sent to the terminal to achieve effective supervision of relevant behaviors.
[0179] Step S151: Parse the correction item identifiers in the specification comparison results, and extract the draft clause unit, difference type description, and reference standard index corresponding to each correction item.
[0180] The correction item identifiers in the standard comparison results are analyzed, and the draft clause unit corresponding to each correction item is extracted one by one to clarify the specific clauses that need to be modified. At the same time, the difference type description is extracted to understand the specific difference types between the clause and the standard clause, and the standard index is referenced to determine the standard clause unit on which the modification is based.
[0181] For example, if a certain amendment identifier indicates that the description of behavior and action in the draft clause unit differs semantically from that in the standard clause, and the reference standard index points to a certain standard clause unit in the regulatory standard library, then after parsing, it can be clearly determined that the part of the draft clause unit that needs to be modified is the description of behavior and action, and the corresponding standard clause unit is used as a reference.
[0182] Step S152: Retrieve the corresponding standard clause unit and standard illustration from the regulatory standard library according to the reference standard index, as a reference for correction.
[0183] Based on the reference standard index obtained from the analysis, the corresponding standard clause units and standard illustrations are accurately retrieved from the regulatory standard library. The content of the standard clauses, the applicable conditions, and the basis for judgment in the standard clause units, as well as the visual information displayed in the standard illustrations, together constitute the reference basis for the revision work, ensuring that the revised draft clauses are consistent with the regulatory standards.
[0184] Step S153: For each amendment, compare the draft clause unit with the corresponding standard clause unit, identify semantic differences and expression differences, and generate a difference analysis report.
[0185] For each amendment, the draft clause unit was meticulously compared with the retrieved standard clause unit. Regarding the text content, the clauses were compared word by word to identify semantic differences, such as different meanings in the description of the same action; at the same time, differences in expression were identified, such as differences in word choice and sentence structure.
[0186] Regarding the illustrations, the illustrations corresponding to the draft clauses are compared with the standard images to identify visual differences, such as the angle of display of actions and the details of the scene background. These differences are then organized and categorized to generate a detailed difference analysis report, which clearly specifies the content, location, and form of the differences.
[0187] Step S154: Based on the difference analysis report, and in conjunction with the association mapping rule set, perform semantic correction processing on the draft clause content of the draft clause unit to ensure that the corrected draft clause content is consistent with the core semantics of the standard clause.
[0188] Based on the difference analysis report and combined with the previously constructed set of association mapping rules, the semantic content of the draft clauses in the draft clause units was semantically revised. For the identified semantic differences, the wording of the draft clauses was adjusted with reference to the core semantics of the standard clauses.
[0189] During the revision process, ensure that the revised content conforms to the semantic consistency constraints of the association mapping rule set, that is, the semantics of the text description matches the corresponding visual elements of the image. For example, if the core semantics of the standard clause "safety devices must be checked before operating the equipment" emphasizes the necessity of the check, while the draft clause states "safety devices may be checked when operating the equipment," then "may" needs to be changed to "must" to maintain consistency with the core semantics of the standard clause.
[0190] For example, step S1541: extract semantic difference points from the difference analysis report, identify the parts of the draft clauses that conflict with the core semantics of the standard clauses, and mark them as semantic fragments to be corrected.
[0191] Carefully study the difference analysis report and extract all semantic differences. Based on these differences, locate the specific expressions in the draft clauses that conflict with the core semantics of the standard clauses, and mark these parts as semantic fragments to be revised. Semantic fragments to be revised may be a word, a phrase, or a complete sentence; their common characteristic is inconsistency with the core semantics of the standard clauses.
[0192] Step S1542: Call the semantic parsing model to perform deep semantic analysis on the semantic fragment to be corrected, identify the type and cause of semantic conflict, and generate semantic conflict analysis results.
[0193] The semantic parsing model is invoked, and the semantic fragment to be corrected is input into the model for deep semantic analysis. The model identifies the type of semantic conflict, such as conceptual confusion, incorrect scope definition, and inverted logical relationships, by analyzing the syntactic structure, semantic collocation, and contextual relationships of the fragment.
[0194] Meanwhile, the causes of semantic conflicts may be due to inaccurate understanding of terminology or insufficient rigor in expression.
[0195] Step S1543: Based on the semantic conflict analysis results and referring to the corresponding standard clause content, generate multiple candidate corrected semantic fragments.
[0196] Based on the semantic conflict analysis results and the corresponding standard clause content, several possible revision schemes were conceived, generating multiple candidate revised semantic fragments. Each candidate fragment was adjusted to address the semantic conflict points, attempting to ensure that the revised semantics remained consistent with the core semantics of the standard clause.
[0197] For example, if a semantic fragment to be corrected causes semantic conflict due to conceptual confusion, multiple different expressions can be generated as candidate semantic fragments for correction, referring to the accurate expression of the concept in the standard clause.
[0198] Step S1544: Convert each candidate semantic segment for correction into a corresponding visual feature description, check whether it matches the visual elements of the illustration corresponding to the clause, and filter out the candidate semantic segments for correction that conform to the association mapping rule set.
[0199] For each candidate semantic fragment to be corrected, it is converted into a corresponding visual feature description according to the correspondence in the association mapping rule set. Then, it is checked whether these visual feature descriptions match the visual elements of the illustration corresponding to the clause, such as the main body contour features, action pose key points, and background area texture features, to see if they are consistent with the visual feature descriptions.
[0200] Candidate semantic fragments that match the visual elements of the illustration and conform to the association mapping rule set are selected to ensure that the corrected text content is consistent with the image information.
[0201] Step S1545: Perform semantic fluency evaluation on the selected candidate modified semantic segments, and calculate the grammatical correctness score and semantic coherence score for each segment.
[0202] The selected candidate semantic segments are evaluated for semantic fluency. The evaluation includes two aspects: grammatical correctness and semantic coherence. The grammatical correctness score mainly examines whether the sentence structure of the segment is complete, whether the word choice is appropriate, and whether the punctuation is correct; the semantic coherence score mainly examines whether the logical relationship within the segment and with the context is smooth and whether the meaning is coherent.
[0203] Using preset evaluation metrics and algorithms, the grammatical correctness score and semantic coherence score of each candidate segment are calculated.
[0204] Step S1546: Select the candidate semantic segment with the highest grammatical correctness score and semantic coherence score, replace the semantic segment to be corrected in the draft clause content, and complete the single semantic correction.
[0205] Compare the grammatical correctness and semantic coherence scores of each candidate semantic segment to be corrected, and select the candidate segment with the highest score. Replace the semantic segment to be corrected in the draft clause with this candidate segment to complete a single correction of the semantic difference.
[0206] Step S155: Perform visual adjustments on the illustrations corresponding to the revised draft clauses to ensure that the visual elements of the images match the semantic elements of the revised text and comply with the collaborative constraints.
[0207] After revising the text, visual adjustments were made to the illustrations corresponding to the draft clauses. Based on the semantic elements of the revised text, the parts of the images that needed adjustment were analyzed, such as the subject's posture and the details of the background.
[0208] Image processing tools are used to edit images, such as adjusting the angle of the subject's movement and modifying the distribution of objects in the background, so that the visual elements of the image match the semantic elements of the corrected text, and meet the collaborative constraints in the association mapping rule set, such as semantic consistency constraints and spatiotemporal correlation constraints.
[0209] Step S156: Re-compare the revised clause unit with the regulatory standard library for consistency. If there are still revisions, repeat the above revision process until all revisions are converted into compliant items.
[0210] The revised clause units, after text and image corrections, are again compared with the corresponding standard clause units in the regulatory standard library to calculate the overall matching degree. If the overall matching degree reaches the preset threshold, it means that the clause unit has been converted into a compliant item; if it does not reach the threshold, it means that there are still differences, and the correction process of steps S151 to S155 needs to be repeated.
[0211] This process is repeated iteratively until all revisions are converted into compliance items, ensuring that all provisions in the draft code of conduct are consistent with the standards in the regulatory standards library.
[0212] Step S157: Integrate all the clause units that have passed the comparison, rearrange them according to the logical order of the code of conduct, and generate the final code of conduct text and the corresponding set of illustrations.
[0213] After all clauses have been compared, they are integrated. Following the inherent logical order of the code of conduct, such as the order of subject definition, behavioral boundaries, and compliance determination, the clauses are rearranged to form a final code of conduct text that is clearly structured and logically rigorous.
[0214] At the same time, the illustrations corresponding to each clause unit are compiled into a set of illustrations that match the final code of conduct text. The illustration set corresponds one-to-one with the clauses in the text, which facilitates understanding and implementation.
[0215] Step S158: Convert the final behavioral norm text into an instruction format recognizable by the regulatory execution terminal, add an execution time identifier and an execution scope identifier, and generate the final behavioral norm regulatory instruction.
[0216] Based on the interface specifications and data format requirements of the regulatory enforcement terminal, the final behavioral specification text is converted into an instruction format that the terminal can recognize, such as a specific XML format or JSON format.
[0217] The converted instructions include an execution time identifier to clearly define the effective and expiration dates of the regulatory instruction; and an execution scope identifier to clearly define the geographical and subject-specific scope to which the instruction applies. These identifiers ensure that regulatory instructions are executed at the correct time and within the correct scope.
[0218] Step S159: Send the final behavioral norm supervision instruction to the supervision execution terminal to trigger the terminal's norm storage and execution preparation operations.
[0219] The final behavioral regulation instructions are sent to the regulatory enforcement terminal via network communication protocols. Upon receiving the instructions, the terminal first verifies them, checking their completeness and legality.
[0220] After successful verification, the terminal's standard storage operation is triggered, storing the regulatory instructions in the terminal's local database for subsequent querying and retrieval. Simultaneously, execution preparation operations are performed, such as loading relevant executable programs and initializing execution parameters, to prepare for subsequent behavior monitoring.
[0221] Pre-training the normative generative model is a crucial step in ensuring its accurate extraction of behavioral normative elements, and specifically includes the following steps:
[0222] Step S211: Collect a large amount of behavioral norm sample data, including historical behavioral norm documents, corresponding illustrations, and relevant compliance judgment cases, and build a model training dataset.
[0223] In this embodiment, the collected data needs to cover different behavior types, subject types, and scene types to ensure the model's generalization ability. The collected data is cleaned and preprocessed, such as removing duplicate data, correcting errors, segmenting and standardizing text, and resizing and converting images.
[0224] Step S212: Divide the training dataset into a training set, a validation set, and a test set. The training set is used to learn the model parameters, the validation set is used to adjust the parameters during the model training process, and the test set is used to evaluate the final performance of the model.
[0225] The division ratio can be set according to the size of the data, for example, a ratio of 7:2:1 can be used.
[0226] Step S213: Construct the network structure of the canonical generative model, which includes a feature encoding layer and a feature extraction layer. The feature encoding layer adopts a Transformer structure with a multi-head self-attention mechanism, and the feature extraction layer adopts a structure combining a multilayer perceptron and a conditional random field.
[0227] Set the hyperparameters of the model, such as the dimension of the feature vector, the number of heads in the multi-head self-attention mechanism, the number of neurons in the hidden layer, the learning rate, and the number of iterations.
[0228] Step S214: Input the text-image collaborative features from the training set into the model for training. During training, the model parameters are continuously adjusted using the backpropagation algorithm to ensure that the model's output, such as the subject definition results, behavioral boundary conditions, and compliance judgment benchmarks, is as consistent as possible with the labeled results in the training data.
[0229] The training effect of the model is monitored using a validation set. Training is stopped when the loss function value on the validation set no longer decreases, in order to avoid overfitting of the model.
[0230] Step S215: Use the test set to evaluate the performance of the trained model. Evaluation metrics include subject definition accuracy, behavior boundary condition extraction accuracy, and compliance judgment benchmark generation accuracy.
[0231] Based on the evaluation results, if the model performance does not meet the preset requirements, the network structure or hyperparameters of the model will be adjusted and retrained; if the requirements are met, the model parameters will be saved and the pre-training of the standardized generative model will be completed.
[0232] The training of the bidirectional attention mechanism model aims to improve its accuracy in calculating the correlation between text semantic elements and image visual elements. The specific steps are as follows:
[0233] Step S311: Collect paired sample data of text semantic elements and image visual elements. These sample data should contain known correlation labels to build a model training dataset.
[0234] Data preprocessing includes converting text semantic elements into vector form and performing feature extraction and vector transformation on image visual elements.
[0235] Step S312: Divide the dataset into training set, validation set, and test set. Refer to step S212 for the division method.
[0236] Step S313: Construct the network structure of the bidirectional attention mechanism model, which includes a text encoding subnetwork, an image encoding subnetwork, and a bidirectional attention computation subnetwork.
[0237] Set the model's hyperparameters, such as the dimension of the encoding vector, the number of attention heads, and the learning rate.
[0238] Step S314: Input the text semantic elements and image visual elements from the training set into the model for training. Adjust the model parameters through the backpropagation algorithm to minimize the error between the correlation degree of the model output and the correlation degree of the sample annotation.
[0239] Use a validation set to monitor the model training process and prevent overfitting.
[0240] Step S315: Evaluate the model's performance using the test set, with the mean squared error of the correlation coefficient as the evaluation metric. Adjust the model based on the evaluation results until its performance meets the requirements, and then save the model parameters.
[0241] Figure 2 The illustration shows exemplary hardware and software components of a natural language and image collaborative definition behavior regulation supervision system 100 that can implement the ideas of this application, according to some embodiments of this application. For example, a processor 120 can be used in the natural language and image collaborative definition behavior regulation supervision system 100 and to perform the functions in this application.
[0242] The Natural Language and Image Collaborative Definition Behavioral Norms Supervision System 100 can be a general-purpose server or a special-purpose server; both can be used to implement the large-scale model system for natural language and image collaborative definition behavioral norms supervision as described in this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the load.
[0243] For example, the natural language and image collaborative behavior regulation supervision system 100 may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and various forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the natural language and image collaborative behavior regulation supervision system 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The methods of this application can be implemented according to these program instructions. The natural language and image collaborative behavior regulation supervision system 100 also includes an I / O interface 150 between the computer and other input / output devices.
[0244] For ease of explanation, only one processor is described in the Natural Language and Image Collaborative Definition of Behavioral Norms Supervision System 100. However, it should be noted that the Natural Language and Image Collaborative Definition of Behavioral Norms Supervision System 100 may also include multiple processors. Therefore, the steps performed by one processor as described in this application may also be performed jointly or individually by multiple processors. For example, if the processor of the Natural Language and Image Collaborative Definition of Behavioral Norms Supervision System 100 performs steps A and B, it should be understood that steps A and B may also be performed jointly by two different processors or individually by one processor. For example, the first processor performs step A, the second processor performs step B, or the first processor and the second processor jointly perform steps A and B.
[0245] Furthermore, embodiments of the present invention also provide a readable storage medium, wherein computer-executable instructions are preset in the readable storage medium, and when the processor executes the computer-executable instructions, the above-mentioned behavior norm supervision model system defined collaboratively by natural language and images is implemented.
[0246] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A method for a large-scale behavioral norm supervision model system for collaborative definition of natural language and images, characterized in that, The method includes: A set of association mapping rules is constructed between natural language descriptions and standardized schematic images. The set of association mapping rules includes the correspondence between text semantic elements and image visual elements and the collaborative constraints. Receive behavioral description text and scene illustration of the scene to be monitored, and perform bidirectional conversion processing on the behavioral description text and the scene illustration according to the association mapping rule set to obtain a text-image collaborative feature set; Based on the text and image collaborative feature set, a pre-trained specification generation model is invoked to perform behavior specification element extraction operations, generating a draft behavior specification that includes behavior boundary conditions and compliance judgment benchmarks; The draft code of conduct is compared with a pre-set regulatory standard library to obtain a comparison result, which includes conformity identifiers and correction identifiers. The draft code of conduct is iteratively adjusted based on the correction item identifiers in the comparison results, and the final code of conduct regulatory instructions are generated and sent to the regulatory execution terminal.
2. The method for a large-scale behavioral norm supervision model system for collaborative definition of natural language and images according to claim 1, characterized in that, The rule set for constructing the association mapping between natural language descriptions and specification diagram images includes: Collect natural language description samples and corresponding specification illustration samples from historical behavior specification documents to form a sample set. The natural language description samples include behavior subjects, behavior actions and scene limiting words. The specification illustration samples include images of behavior subjects, action posture images and scene background images. The natural language description sample is subjected to semantic word segmentation processing, and the part-of-speech tags and semantic role annotations of each segmentation unit are extracted to obtain text semantic elements. The text semantic elements include subject referential identifiers, action verb identifiers and scene modification identifiers. Visual feature extraction processing is performed on the standard schematic image sample to identify the main body outline region, action posture key points and scene background region in the image to obtain the image visual elements. The image visual elements include the main body outline features, posture key point coordinates and background region texture features. A bidirectional attention mechanism model is invoked to calculate the correlation between the text semantic elements and the image visual elements, generating a semantic visual correlation matrix. The elements in the semantic visual correlation matrix represent the matching strength between the text semantic elements and the image visual elements. Based on the semantic visual association matrix, text visual element pairs that meet the preset conditions for matching strength are selected, and the correspondence between text semantic elements and image visual elements is constructed by combining the behavioral norm logical relationship. Based on the correspondence, collaborative constraints are set for text description and image display. These collaborative constraints include semantic consistency constraints, spatiotemporal correlation constraints, and expression integrity constraints. The correspondence and collaborative constraints are integrated to form an association mapping rule set.
3. The method for a large-scale behavioral norm supervision model system for collaborative definition of natural language and images according to claim 1, characterized in that, The process involves receiving behavioral description text and scene illustrations of the scenario to be monitored, and then performing bidirectional conversion processing on the behavioral description text and scene illustrations according to the association mapping rule set to obtain a text-image collaborative feature set, including: Receive behavioral description text of the scenario to be monitored, perform syntactic structure analysis on the behavioral description text, identify subject-verb-object structure and modifiers, and extract semantic elements of the text to be processed. The semantic elements of the text to be processed include the identifier of the subject to be monitored, the description of the action to be monitored, and the description of the scenario to be monitored. Receive a scene illustration of the scene to be monitored, perform image segmentation processing on the scene illustration to separate the main body area, action area and background area, and extract the visual elements of the image to be processed. The visual elements of the image to be processed include the outline of the subject to be monitored, the action posture to be monitored and the background features of the scene to be monitored. Based on the correspondence in the association mapping rule set, the semantic elements of the text to be processed are converted into corresponding visual feature descriptions to generate a text-to-visual feature set, which includes subject visualization parameters, action visualization parameters, and scene visualization parameters. Based on the correspondence in the association mapping rule set, the visual elements of the image to be processed are converted into corresponding semantic feature descriptions to generate a visual-to-text feature set, which includes subject semantic tags, action semantic tags and scene semantic tags. The feature fusion module is invoked to perform collaborative verification processing on the text-to-visual feature set and the visual-to-text feature set, eliminating conflicting features that violate collaborative constraints, and obtaining a preliminary collaborative feature set. The preliminary collaborative feature set is subjected to feature dimension adaptation processing to make the representation dimensions of text-derived features and visual-derived features consistent, thereby generating a text-image collaborative feature set containing association weight parameters.
4. The method for a large-scale behavioral norm supervision model system for collaborative definition of natural language and images according to claim 1, characterized in that, The process involves using the pre-trained specification generation model based on the text-image collaborative feature set to perform behavior specification element extraction, generating a draft behavior specification that includes behavior boundary conditions and compliance judgment benchmarks, including: The text-image collaborative feature set is input into the feature encoding layer of a pre-trained canonical generation model. The correlation between features is modeled and processed through a multi-head self-attention mechanism to generate encoded feature vectors. The element extraction layer of the above-mentioned standard generation model receives the encoded feature vector, performs the behavior subject definition operation, determines the subject scope and subject attribute features to which the behavior standard applies, and generates subject definition results. Based on the subject definition results, the spatiotemporal range of the behavior is analyzed through the element extraction layer to determine the time interval and spatial region limits of the behavior, and to generate the behavior boundary conditions. By combining compliance clues in text and image collaborative features, permitted and prohibited forms of behavior are identified, and compliance judgment benchmarks for behavior are constructed. The compliance judgment benchmarks include positive judgment indicators and negative judgment indicators. The subject definition results, the behavioral boundary conditions, and the compliance judgment criteria are integrated and processed according to the preset standard text structure to generate a draft code of conduct. The draft code of conduct includes clause numbers, clause content, and corresponding illustration indexes.
5. The method for a large-scale behavioral norm supervision model system for collaborative definition of natural language and images according to claim 4, characterized in that, The step of inputting the text-image collaborative feature set into the feature encoding layer of a pre-trained canonical generative model, and modeling the correlation between features through a multi-head self-attention mechanism to generate encoded feature vectors includes: The text-image collaborative feature set is subjected to feature decomposition processing to obtain a text-derived feature subset and a visual-derived feature subset. The text-derived feature subset includes semantic label features and association weight features, and the visual-derived feature subset includes visual parameter features and matching strength features. The text-derived feature subset and the visual-derived feature subset are input into the embedding module of the feature encoding layer, and vector transformation processing is performed respectively to generate a text feature vector sequence and a visual feature vector sequence, wherein the dimensions of the text feature vector sequence and the visual feature vector sequence are consistent. A multi-head self-attention mechanism is invoked to perform cross-attention calculation on the text feature vector sequence and the visual feature vector sequence to generate a text-visual interaction attention weight matrix, which is used to represent the correlation strength between text features and visual features. Based on the text visual interaction attention weight matrix, the text feature vector sequence and the visual feature vector sequence are weighted and fused to generate an interactive fused feature vector sequence. The interactive fused feature vector sequence is subjected to nonlinear transformation processing through a feedforward neural network of the feature encoding layer to enhance the expressive power of the features and generate an intermediate feature vector sequence. The intermediate feature vector sequence is subjected to layer equalization to obtain encoded feature vectors with a unified representation.
6. The method for a large-scale behavioral norm supervision model system for collaborative definition of natural language and images according to claim 4, characterized in that, The feature extraction layer of the standard generation model receives the encoded feature vector, performs a behavior subject definition operation, determines the subject scope and subject attribute features to which the behavior standard applies, and generates subject definition results, including: The element extraction layer of the above-mentioned standard generation model receives the encoded feature vector, performs matching processing through a preset subject feature template, identifies the feature components in the encoded feature vector that are related to the behavioral subject, and obtains the subject feature vector. Based on the subject feature vector analysis, the type attribute of the subject is determined, the category identifier and category feature description of the subject are determined, and the subject type definition result is generated; Based on the subject type definition result, extract the constraint information related to the subject scope from the encoded feature vector, determine the number range of subjects, identity restrictions and qualification conditions, and generate the subject scope definition result; Analyze the attribute feature information in the subject feature vector to identify the subject's inherent attribute features and dynamic attribute features. The inherent attribute features include the subject's static feature parameters, and the dynamic attribute features include the subject's state feature parameters. By integrating the subject type definition results, the subject scope definition results, and the attribute feature information, a subject definition result containing subject identifier, type description, scope limitation, and attribute parameters is generated according to the subject definition rules.
7. The method for a large-scale behavioral norm supervision model system for collaborative definition of natural language and images according to claim 1, characterized in that, The process of comparing the draft code of conduct with a pre-set regulatory standard library to obtain the comparison results includes: Obtain a pre-set regulatory standard library, which contains historically valid behavioral norms standard texts and corresponding standard illustration image sets. Each standard text contains standard clauses, scope of application, and judgment basis. The draft code of conduct is broken down into multiple independent draft clause units, each of which includes clause content, corresponding illustrations, and applicable conditions. The standard texts in the regulatory standard library are split into clauses to obtain multiple independent standard clause units. Each standard clause unit contains the standard clause content, the corresponding standard image, and the standard application conditions. Semantic matching is performed between the draft clause content of the draft clause unit and the standard clause content of the standard clause unit to generate a text similarity score; Visual matching processing is performed on the schematic diagrams corresponding to the draft clause units and the standard images corresponding to the standard clause units to generate image similarity scores; The overall matching degree is calculated based on the text similarity score and the image similarity score. When the overall matching degree reaches a preset threshold, the draft clause unit is marked as a compliant item and a compliant item identifier is added. When the overall matching degree does not reach the preset threshold, the location and type of the differences are analyzed, the draft clause unit is marked as a correction item and a correction item identifier is added. The correction item identifier includes a description of the difference type and a reference standard index. Integrate the conformity and amendment identifiers of all draft clause units to generate specification comparison results.
8. The method for a large-scale behavioral norm supervision model system for collaborative definition of natural language and images according to claim 7, characterized in that, The semantic matching process between the draft clause content of the draft clause unit and the standard clause content of the standard clause unit to generate a text similarity score includes: The draft clause content of the draft clause unit is processed by word segmentation and part-of-speech tagging. After removing stop words, a valid word sequence is obtained. Each valid word is processed by word vector transformation to generate a draft word vector sequence. The standard clause content of the standard clause unit is processed by word segmentation and part-of-speech tagging. After removing stop words, a standard word sequence is obtained. Each standard word is processed by word vector transformation to generate a standard word vector sequence. Calculate the cosine similarity between the draft word vector sequence and the corresponding word vectors in the standard word vector sequence to obtain a word-level similarity matrix; Perform optimal path search processing on the word-level similarity matrix to determine the best matching path between two word sequences and generate a path matching score. Extract the syntactic structure features of the draft clauses and the standard clauses, calculate the edit distance of the syntactic tree, and obtain the structural similarity score; The path matching score and the structural similarity score are weighted and summed according to preset weights to obtain a preliminary text similarity. The initial text similarity is processed by interval mapping, which maps it to a preset numerical range to generate a text similarity score.
9. The method for a large-scale behavioral norm supervision model system for collaborative definition of natural language and images according to claim 1, characterized in that, The step of iteratively adjusting the draft code of conduct based on the correction item identifiers in the comparison results, generating a final code of conduct regulatory instruction, and sending it to the regulatory execution terminal includes: Parse the correction item identifiers in the standard comparison results, and extract the draft clause unit, difference type description, and reference standard index corresponding to each correction item; According to the reference standard index, the corresponding standard clause unit and standard illustration are retrieved from the regulatory standard library as a reference for revision; For each amendment, the draft clause unit is compared with the corresponding standard clause unit to identify semantic and expression differences and generate a difference analysis report. Based on the aforementioned difference analysis report, and in conjunction with the association mapping rule set, the content of the draft clauses of the draft clause units is semantically corrected to ensure that the corrected draft clause content is consistent with the core semantics of the standard clauses. Visual adjustments were made to the illustrations corresponding to the revised draft clauses to ensure that the visual elements of the images match the semantic elements of the revised text and comply with the collaborative constraints. The revised clause unit is re-compared with the regulatory standard library for consistency. If there are still revisions, the above revision process is repeated until all revisions are converted into compliant items. Integrate all the clause units that have passed the comparison, rearrange them according to the logical order of the code of conduct, and generate the final code of conduct text and the corresponding set of illustrations; The final behavioral guidelines text is converted into an instruction format recognizable by the regulatory execution terminal, and execution time and execution scope identifiers are added to generate the final behavioral guidelines regulatory instructions; The final behavioral norm supervision instruction is sent to the supervision execution terminal, triggering the terminal's norm storage and execution preparation operations.
10. A large-scale behavioral norm supervision model system collaboratively defined by natural language and images, characterized in that, The system includes a processor and a memory connected to the processor. The memory is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the memory to implement the method of a behavior norm supervision big data model system for collaborative definition of natural language and images as described in any one of claims 1-9.
Citation Information
Patent Citations
Research and development document processing method and device
CN120087351A
Fabricated building component intelligent generation and real-time detection method and system based on multi-modal AI
CN120449262A