Behavior specification supervision large model system cooperatively defined by natural language and picture

By constructing a set of association mapping rules between natural language and images, combining it with a pre-trained model to generate a draft code of conduct, and comparing it with the regulatory standards library, the ambiguity of natural language descriptions is resolved and precise supervision of behavioral norms is achieved.

CN120705641AActive Publication Date: 2025-09-26NAT CERTIFICATION TECH (HANGZHOU) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511204819.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-09-26
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

The existing behavioral norms supervision method mainly relies on natural language descriptions, which are vague and ambiguous, making it difficult to fully and intuitively express the requirements of complex behavioral scenarios. It lacks a deep collaborative mechanism between natural language and images, resulting in inaccurate supervision.

Method used

Construct a set of association mapping rules between natural language descriptions and standard schematic images, generate a text-image collaborative feature set through bidirectional conversion processing, call the pre-trained standard generation model to extract behavioral standard elements, compare them with the regulatory standard library, and iteratively adjust to generate the final regulatory instructions.

Benefits of technology

It has achieved improvements in the accuracy and efficiency of behavioral norms, avoided omissions and deviations in manual formulation, ensured the scientific nature and rationality of supervision, and improved the precision and reliability of supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705641A_ABST
    Figure CN120705641A_ABST
Patent Text Reader

Abstract

The invention provides a behavior specification supervision large model system cooperatively defined by natural languages and pictures, and belongs to the technical field of behavior specification supervision. Firstly, an association mapping rule set between natural language description and specification schematic pictures is constructed, and the corresponding relation and cooperative constraint conditions of text semantic elements and picture visual elements are determined; then, receiving a behavior description text and a scene schematic picture of a scene to be supervised, performing bidirectional conversion processing according to the association mapping rule set to obtain a text and picture collaborative feature set, and then extracting behavior specification elements based on the calling of a pre-trained specification generation model; generating a behavioral specification draft containing behavioral boundary conditions and compliance criterion; according to the method, consistency comparison processing is carried out on the behavioral specification draft and the preset supervision standard library, iterative adjustment is carried out on the draft according to the correction item identifier in the comparison result, and finally, the behavioral specification supervision instruction is generated and sent to the supervision execution terminal, so that the accuracy and reliability of behavioral specification supervision are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of behavioral norm supervision, and in particular to a large-scale behavioral norm supervision model system collaboratively defined by natural language and images. Background Art

[0002] In today's society, effective regulation of behavioral norms in various scenarios is crucial for maintaining social order, public safety, and the healthy development of industries. Traditional methods of regulating behavioral norms rely primarily on manual formulation and implementation, typically defining behavioral norms through textual descriptions, such as laws and regulations, industry guidelines, and corporate rules and regulations. However, this single natural language description approach has certain limitations. On the one hand, natural language is ambiguous and polysemic, and different people may have different understandings of the same text description, which can easily lead to disputes and deviations in the implementation and judgment of behavioral norms. On the other hand, for some complex behavioral scenarios, it is difficult to fully and intuitively present the specific requirements and boundary conditions of the behavior through textual descriptions alone, making it difficult for regulators and regulated entities to accurately grasp the content of the behavioral norms.

[0003] While some research has attempted to incorporate visual elements like images to aid in the expression of behavioral norms, there is currently a lack of effective methods that deeply integrate natural language and images to achieve precise definition and efficient regulation of behavioral norms. Existing methods often simply use images as supplementary text, failing to establish a close connection and synergy between natural language and images, and thus failing to fully leverage the advantages of both in regulating behavioral norms. Summary of the Invention

[0004] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a large-scale model system for behavior regulation collaboratively defined by natural language and images, the method comprising: Constructing an association mapping rule set between natural language descriptions and standard schematic images, wherein the association mapping rule set includes correspondences between text semantic elements and image visual elements and collaborative constraints; Receiving a behavior description text and a scene diagram image of a scene to be supervised, and performing a bidirectional conversion process on the behavior description text and the scene diagram image according to the association mapping rule set to obtain a text-image collaborative feature set; Based on the text-image collaborative feature set, a pre-trained specification generation model is called to perform a behavioral specification element extraction operation to generate a draft behavioral specification including behavioral boundary conditions and compliance determination benchmarks; Performing a consistency comparison between the draft code of conduct and a preset regulatory standard library to obtain a standard comparison result, wherein the standard comparison result includes a conforming item identifier and a revised item identifier; The draft code of conduct is iteratively adjusted according to the correction item identifier in the code comparison result, and a final code of conduct supervision instruction is generated and sent to the supervision execution terminal.

[0005] On the other hand, an embodiment of the present invention also provides a large model system for behavioral norm supervision collaboratively defined by natural language and images, including a processor and a machine-readable storage medium, the machine-readable storage medium being connected to the processor, the machine-readable storage medium being used to store programs, instructions or codes, and the processor being used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0006] Based on the above aspects, the embodiment of the present invention clarifies the correspondence between text semantic elements and image visual elements and collaborative constraints by constructing an association mapping rule set between natural language descriptions and standard schematic images. According to the association mapping rule set, the received behavior description text and scene schematic images of the scene to be supervised are bidirectionally converted to obtain a text-image collaborative feature set, which realizes the fusion and complementarity of natural language and image information, and can express the content of the behavioral norms more comprehensively and accurately. Based on the text-image collaborative feature set, the pre-trained standard generation model is called to extract behavioral norm elements, and a draft behavioral norm is generated that includes behavioral boundary conditions and compliance judgment benchmarks, thereby improving the accuracy and efficiency of the formulation of behavioral norms and avoiding possible omissions and deviations in manual formulation. The draft behavioral norm is compared with the preset regulatory standard library for consistency, which can timely discover problems in the draft, and iteratively adjust the draft behavioral norm through the correction item identification, and finally generate behavioral norm supervision instructions that meet regulatory requirements, ensuring the scientificity and rationality of the behavioral norms and greatly improving the accuracy and reliability of behavioral norm supervision. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 This is a schematic diagram of the execution flow of a large model system for behavioral norm supervision that is collaboratively defined by natural language and images, provided by an embodiment of the present invention.

[0008] Figure 2 Schematic diagram of exemplary hardware and software components of a system for collaboratively defining behavioral norms using natural language and images, provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0009] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a large model system for regulating behavioral norms through collaborative definition of natural language and images provided by an embodiment of the present invention. The following is a detailed introduction to the large model system for regulating behavioral norms through collaborative definition of natural language and images.

[0010] Step S110: constructing an association mapping rule set between the natural language description and the standard schematic image, wherein the association mapping rule set includes the correspondence between the text semantic elements and the image visual elements and the collaborative constraint conditions.

[0011] Step S111: Collect natural language description samples and corresponding standard schematic image samples in the historical behavior specification file to form a sample set, wherein the natural language description samples include behavior subjects, behavior actions and scene qualifiers, and the standard schematic image samples include behavior subject images, action posture images and scene background images.

[0012] In this embodiment, when collecting samples, representative content must be screened from a large number of historical behavioral specification documents. Natural language description samples should cover behavioral requirements in different situations. For example, "Within a specific area, relevant personnel must operate equipment according to specified procedures, check equipment status before operation, must not leave without authorization during operation, and record operation results after operation." Here, "relevant personnel" is the subject of the behavior, "operate equipment according to specified procedures," "check equipment status," "leave without authorization," and "record operation results" are the behavioral actions, and "within a specific area" is a scenario qualifier.

[0013] The corresponding standard diagram image samples must accurately match these descriptions. These images include images showing relevant personnel inspecting equipment in a specific area, operating equipment according to procedures, maintaining their positions during operations, and recording the results after operations. These images must clearly depict the subject, posture, and scene background. For potentially privacy-sensitive data, such as facial information of people in images, pixel blurring technology can be used to process it. Personal identification information involved in the text will also be anonymized to ensure information security.

[0014] Step S112: performing semantic word segmentation processing on the natural language description sample, extracting the part-of-speech tag and semantic role labeling of each word segmentation unit, and obtaining text semantic elements, wherein the text semantic elements include a subject reference identifier, an action verb identifier, and a scene modification identifier.

[0015] Taking "In a specific area, relevant personnel need to operate the equipment according to the specified process, check the equipment status before operation, must not leave without permission during the operation, and must record the operation results after the operation" as an example, semantic word segmentation processing is performed, and the word segmentation units obtained are "in" "specific area", "relevant personnel" "need" "to" "operate" "equipment" "according to the specified process", "before operation", "need" "check" "equipment status", "during operation", "must not" "leave without permission", "after the operation is completed", "need" "record" "operation results".

[0016] Add a part-of-speech tag to each word segment unit, "in" is a preposition, "within a specific area" is a directional noun phrase, "relevant personnel" is a noun phrase, "need" is an auxiliary verb, "according to the specified process" is a prepositional phrase, "operate" is a verb, "equipment" is a noun, "before operation" is a time noun phrase, "check" is a verb, "equipment status" is a noun phrase, "during operation" is a time noun phrase, "shall not" is an auxiliary verb, "leave without permission" is a verb phrase, "after the operation is completed" is a time noun phrase, "record" is a verb, and "operation results" is a noun phrase.

[0017] When performing semantic role annotation, "relevant personnel" is the initiator of all actions, i.e., the subject; "operating equipment," "checking equipment status," "leaving without permission," and "recording operation results" are specific behavioral actions; "within a specific area" is the scene where the action occurs, and "before operation," "during operation," and "after operation" are the time and scene restrictions for the action. This extracts text semantic elements, with the subject reference identified as "relevant personnel," the action verbs identified as "operate," "check," "leave without permission," and "record," and the scene modifiers identified as "within a specific area," "before operation," "during operation," and "after operation."

[0018] Step S113: Perform visual feature extraction processing on the standard schematic picture sample to identify the subject contour area, action posture key points and scene background area in the picture to obtain picture visual elements, which include subject contour features, posture key point coordinates and background area texture features.

[0019] Visual feature extraction is performed on the canonical schematic image samples corresponding to the aforementioned natural language description samples. For example, in an image showing a person inspecting equipment within a specific area, an edge detection algorithm is used to identify the subject's outline. This outline represents the person's overall appearance, including the general shape of their body and the distribution of their limbs.

[0020] The key points of the action posture include the position where the hand touches the device, the angle of the head toward the device, the standing posture of the body, etc. The coordinates of these key points are represented by the pixel coordinate system of the image. Each key point has its corresponding horizontal and vertical coordinate values ​​to accurately reflect its position in the image.

[0021] The scene background area refers to the environment within a specific area, including the appearance of the equipment, the layout of surrounding facilities, etc. The texture features of the background area are extracted through the texture analysis algorithm. The above features include the density, direction, color distribution, etc. of the texture to distinguish different scene backgrounds.

[0022] Step S114: calling the bidirectional attention mechanism model to calculate the correlation between the text semantic elements and the image visual elements to generate a semantic visual correlation matrix, wherein the elements in the semantic visual correlation matrix represent the matching strength between the text semantic elements and the image visual elements.

[0023] The extracted text semantic elements and image visual elements are fed into a bidirectional attention mechanism model. The model first encodes the subject reference identifier, action verb identifier, and scene modifier identifier in the text semantic elements, and simultaneously encodes the subject contour features, posture key point coordinates, and background area texture features in the image visual elements.

[0024] After encoding, the model determines the degree of association between the semantic elements of the text and the visual elements of the image by calculating the attention weights between them. For each semantic element in the text, the model pays attention to the related visual elements in the image; conversely, for each visual element in the image, the model also pays attention to the related semantic elements in the text.

[0025] Through this bidirectional attention approach, the matching strength between each semantic element in the text and each visual element in the image is calculated. These matching strength values ​​form a semantic-visual association matrix. For example, the matching strength between the action verb "check device status" in the text and the coordinates of the key points of the hand touching the device in the image will correspond to a high value in the matrix.

[0026] Step S115: screening text-visual element pairs whose matching strengths meet preset conditions according to the semantic-visual association matrix, and constructing corresponding relationships between text semantic elements and image visual elements in combination with behavioral normative logical relationships.

[0027] A preset matching strength condition is set, such as the matching strength being greater than a certain threshold. Combinations of text semantic elements and image visual elements that meet the preset condition are selected from the semantic visual association matrix to form text visual element pairs.

[0028] Combined with the logical relationships within the behavioral norms, analyze the inherent connections between these pairs of elements. For example, the subject reference identifier "relevant personnel" should correspond to the subject's outline features in the image; the action verb identifier "operate equipment" should correspond to the coordinates of the key points of the person's posture when operating the equipment in the image; and the scene modification identifier "within a specific area" should correspond to the texture features of the background area in the image.

[0029] Based on these logical relationships, a clear correspondence between text semantic elements and image visual elements is established to ensure that each text semantic element can find a matching image visual element, and vice versa.

[0030] Step S116: setting collaborative constraints between text description and picture display based on the corresponding relationship, wherein the collaborative constraints include semantic consistency constraints, spatiotemporal correlation constraints, and expression integrity constraints, and integrating the corresponding relationship and the collaborative constraints to form an association mapping rule set.

[0031] Based on the established correspondence, collaborative constraints are set. Semantic consistency constraints require that the semantics of the text description and the content of the image display are consistent in meaning and cannot be contradictory. For example, if the text description says "No one is allowed to leave the operation without permission", the image cannot show the relevant personnel leaving the operation station.

[0032] The spatiotemporal relevance constraint requires that the temporal and spatial information in the text description match the temporal and spatial context presented in the image. For example, if the temporal context constraint "before operation" in the text is used, the corresponding image should depict the state of the relevant personnel before they begin operating the equipment.

[0033] The completeness constraint requires that both the textual description and the image display fully express the content of the behavioral specification, without missing any information. For example, for the action "record operation results," the text should explain the recording requirements, and the image should show the recording process or results. The combination of these two fully expresses the specification. The established correspondence and the set collaborative constraints are integrated to form an association mapping rule set.

[0034] Step S120: receiving a behavior description text and a scene diagram image of a scene to be supervised, and performing a bidirectional conversion process on the behavior description text and the scene diagram image according to the association mapping rule set to obtain a text-image collaborative feature set.

[0035] In practice, the system receives both textual descriptions of behaviors and images of scenarios to be monitored. This information may be input by personnel or collected through devices. The system then performs a bidirectional conversion of these two types of information based on a set of association mapping rules, converting textual information into corresponding visual features and visual information into corresponding semantic features, ultimately fusing them into a collaborative feature set.

[0036] Step S121: Receive a behavior description text of a scenario to be supervised, perform syntactic structure analysis on the behavior description text, identify the subject-verb-object structure and modifying components, and extract semantic elements of the text to be processed. The semantic elements of the text to be processed include an identifier of the subject to be supervised, a description of the action to be supervised, and a description of the scenario to be supervised.

[0037] Receive behavioral description text of the scenario to be supervised, such as "a group of people in a specific place should move in a predetermined manner, avoid specific objects when moving, and stay in a designated location after moving."

[0038] The text was analyzed for syntactic structure, and a syntactic analysis algorithm was used to identify the subject-verb-object structure. "A certain group" is the subject, "moving" is the predicate, and "in a specific place," "in a predetermined manner," "avoiding specific objects," and "in a designated location" are modifiers.

[0039] Based on the analysis results, semantic elements of the text to be processed are extracted. The subject to be supervised is identified as "a certain group," the action to be supervised is described as "moving in a predetermined manner," "avoiding specific objects," and "being in a designated location," and the scene to be supervised is described as "in a specific place," "while moving," and "after moving."

[0040] Step S122: Receive a scene schematic image of the scene to be supervised, perform image segmentation processing on the scene schematic image, separate the subject area, action area and background area, and extract visual elements of the image to be processed, wherein the visual elements of the image to be processed include the outline of the subject to be supervised, the action posture to be supervised and the background features of the scene to be supervised.

[0041] A scene schematic diagram of a scene to be supervised is received, where the scene schematic diagram shows the status of a group of people in a specific place.

[0042] The image is segmented and a region growing algorithm is used to divide the image into different regions. The subject region is the area where a group of people are located, the action region is the area where the group is moving, and the background region is the environment of a specific place, such as the ground and surrounding objects.

[0043] Extract visual elements from the segmented regions. The subject outline is the overall outline of a group. The action posture is the body posture of a group when moving, such as the direction of limb extension and position changes. The background features are the background environment characteristics of a specific location, such as the texture of the ground and the shape and distribution of surrounding objects.

[0044] Step S123: According to the corresponding relationship in the association mapping rule set, the semantic elements of the text to be processed are converted into corresponding visual feature descriptions to generate a text-to-visual feature set, which includes subject visualization parameters, action visualization parameters and scene visualization parameters.

[0045] According to the corresponding relationships in the association mapping rule set, the semantic elements of the text to be processed are converted. The identification of the subject to be supervised, "a certain group", is converted into the subject's visual parameters. These parameters include the approximate size of the group and the visual representation of its appearance characteristics.

[0046] The description of the action to be supervised, such as "moving in a predetermined manner", "avoiding specific objects", and "being at a specified location", is converted into action visualization parameters, such as the speed range of movement, direction change, distance range when avoiding objects, coordinate range of the specified location, and other visual parameter representations.

[0047] The descriptions of the scenes to be supervised, such as "in a specific place", "while moving" and "after moving", are converted into scene visualization parameters, such as the spatial scope of the place, changes in the visual characteristics of the scene at different time points and other parameters.

[0048] These parameters together constitute the text-to-visual feature set.

[0049] Step S124: According to the corresponding relationship in the association mapping rule set, the visual elements of the image to be processed are converted into corresponding semantic feature descriptions to generate a visual-to-text feature set, which includes a subject semantic label, an action semantic label and a scene semantic label.

[0050] Based on the corresponding relationships in the association mapping rule set, the visual elements of the image to be processed are converted into semantic feature descriptions. The subject semantic label corresponding to the outline of the subject to be supervised is the semantic information such as the type and composition of the "certain group".

[0051] The action semantic labels corresponding to the action posture to be supervised are semantic descriptions such as "way of moving", "behavior of avoiding objects", and "state of being in a specified position".

[0052] The scene semantic labels corresponding to the background features of the scene to be supervised are semantic information such as "the type of specific place", "the environmental characteristics within the place", and "the state of the scene at different time points".

[0053] These semantic labels constitute the vision-to-text feature set.

[0054] Step S125: calling a feature fusion module to perform collaborative verification processing on the text-to-visual feature set and the visual-to-text feature set, eliminating conflicting features that violate collaborative constraints, and obtaining a preliminary collaborative feature set.

[0055] The feature fusion module is called and the text-to-visual feature set and the visual-to-text feature set are input into it. The module verifies the features in the two sets according to the collaborative constraints in the association mapping rule set.

[0056] Check semantic consistency to see if the semantics expressed by the text-to-visual features are consistent with the semantics expressed by the visual-to-text features. For example, whether the text-to-visual feature "a group of people moves quickly" conflicts with the visual-to-text feature "a group of people moves slowly."

[0057] Check spatiotemporal correlation to confirm whether the temporal and spatial parameters in the text-to-visual features match the temporal and spatial semantics in the visual-to-text features. For example, whether the text-to-visual feature that "movement occurred in a certain time period" contradicts the visual-to-text feature that "movement occurred in another time period."

[0058] Check the completeness of the expression to ensure that the two feature sets combined can fully express the information of the scenario to be regulated without missing any key content.

[0059] For conflicting features that violate collaborative constraints, such as features with semantic contradictions, spatiotemporal mismatches, or information missing, they are eliminated and features that meet the conditions are retained to form a preliminary collaborative feature set.

[0060] Step S126: performing feature dimension adaptation processing on the preliminary collaborative feature set so as to make the expression dimensions of the text-derived features and the visual-derived features consistent, and generating a text-image collaborative feature set including associated weight parameters.

[0061] The text-derived features and visual-derived features in the preliminary collaborative feature set may differ in their representation dimensions. For example, text-derived features may be presented in the dimension of the number of semantic labels, while visual-derived features may be presented in the dimension of the numerical value of parameters.

[0062] These features are then dimensionalized, using a feature conversion algorithm to convert text-derived and visual-derived features into the same dimensional space. For example, the semantic labels of text-derived features are converted into corresponding vector representations to keep them dimensionally consistent with the parameter vectors of visual-derived features.

[0063] During the adaptation process, each feature is assigned an association weight parameter based on the degree of correlation between the features. The higher the correlation, the larger the weight parameter. The adapted features and the corresponding association weight parameters are integrated together to generate a text-image collaborative feature set.

[0064] Step S130: Based on the text-image collaborative feature set, a pre-trained specification generation model is called to perform a behavior specification element extraction operation to generate a behavior specification draft including behavior boundary conditions and compliance judgment criteria.

[0065] Leveraging the collaborative feature set of text and images, a pre-trained code generation model is invoked. This model, trained on a large number of samples, can extract key elements of behavioral norms from these collaborative features, such as the scope of the behavioral subject, the boundary conditions of the behavior, and the benchmarks for determining compliance. This model then generates a draft code of conduct.

[0066] Step S131: Input the text-image collaborative feature set into the feature encoding layer of the pre-trained canonical generation model, model the correlation relationship between the features through the multi-head self-attention mechanism, and generate a coding feature vector.

[0067] The text-image collaborative feature set is fed into the feature encoding layer of the canonical generation model. The feature encoding layer first preprocesses the input features, converting different types of features into a format that the model can process.

[0068] Then, the correlation between features is modeled through a multi-head self-attention mechanism. The multi-head self-attention mechanism simultaneously calculates the attention weights between features from multiple different perspectives, with each perspective focusing on a different correlation between features.

[0069] Through the above method, the complex correlation information in the feature set can be captured, and after integrating this information, a coded feature vector is generated. The coded feature vector includes the key information and correlation relationship in the collaborative feature set.

[0070] Step S1311: performing feature splitting processing on the text-image collaborative feature set to obtain a text-derived feature subset and a visual-derived feature subset, wherein the text-derived feature subset includes semantic label features and associated weight features, and the visual-derived feature subset includes visual parameter features and matching strength features.

[0071] The text-image collaborative feature set is split into text-derived feature subsets and visual-derived feature subsets according to the source and type of the features.

[0072] The semantic label features in the text-derived feature subset are various semantic labels extracted from text information, such as subject semantic labels, action semantic labels, etc.; the associated weight features are the weight parameters assigned to each text feature during the feature fusion process.

[0073] The visual parameter features in the visual derived feature subset are various visual parameters extracted from image information, such as subject visualization parameters, action visualization parameters, etc.; the matching strength features are parameters that represent the degree of matching between text features and visual features.

[0074] Step S1312: Input the text-derived feature subset and the visual-derived feature subset into the embedding module of the feature coding layer, perform vector conversion processing on them respectively, and generate a text feature vector sequence and a visual feature vector sequence, wherein the dimensions of the text feature vector sequence and the visual feature vector sequence remain consistent.

[0075] The text-derived feature subset and the visual-derived feature subset are input into the embedding module of the feature encoding layer. The embedding module converts each feature into a vector, converting text-derived features such as semantic label features and association weight features into text feature vectors, and converting visual-derived features such as visual parameter features and matching strength features into visual feature vectors.

[0076] During the conversion process, ensure that the generated text feature vector sequence and visual feature vector sequence have the same dimensions to facilitate subsequent feature interaction processing. For example, both the text feature vector sequence and the visual feature vector sequence use vectors of the same length to represent each feature.

[0077] Step S1313: Call the multi-head self-attention mechanism to perform cross-attention calculation processing on the text feature vector sequence and the visual feature vector sequence to generate a text-visual interaction attention weight matrix, which is used to represent the correlation strength between text features and visual features.

[0078] The multi-head self-attention mechanism is called, which contains multiple parallel attention heads. Each attention head processes the text feature vector sequence and the visual feature vector sequence respectively, and calculates the cross-attention weight between the text features and the visual features.

[0079] Specifically, each attention head performs a linear transformation on the sequence of textual and visual feature vectors to obtain a query vector, a key vector, and a value vector. For the sequence of textual feature vectors, this linear transformation yields a textual query vector, a textual key vector, and a textual value vector; for the sequence of visual feature vectors, this linear transformation yields a visual query vector, a visual key vector, and a visual value vector.

[0080] Next, each attention head calculates the similarity between the text query vector and the visual key vector, and normalizes these similarities through the softmax function to obtain the cross-attention weight, which reflects the strength of the association between text features and visual features.

[0081] After the calculations of multiple attention heads are completed, the cross-attention weights obtained by each attention head are concatenated to generate a text-visual interaction attention weight matrix. Each element in this text-visual interaction attention weight matrix represents the strength of the association between a specific text feature and a specific visual feature.

[0082] Step S1314: performing weighted fusion processing on the text feature vector sequence and the visual feature vector sequence based on the text-visual interaction attention weight matrix to generate an interactive fusion feature vector sequence.

[0083] The text-visual interaction attention weight matrix is ​​used to perform a weighted fusion of the text feature vector sequence and the visual feature vector sequence. For each text feature vector, the corresponding visual feature vectors are weighted summed according to the strength of its association with each visual feature vector to obtain the corresponding fusion vector of the text feature vector. Similarly, for each visual feature vector, the corresponding text feature vectors are weighted summed according to the strength of its association with each text feature vector to obtain the corresponding fusion vector of the visual feature vector.

[0084] These fused vectors are arranged in the original sequence to generate an interactive fusion feature vector sequence. This interactive fusion feature vector sequence contains both text feature information and visual feature information, and reflects the interactive correlation between the two.

[0085] Step S1315: performing nonlinear transformation processing on the interactive fusion feature vector sequence through the feedforward neural network of the feature coding layer to enhance the expressive power of the features and generate an intermediate feature vector sequence.

[0086] The interactive fusion feature vector sequence is input into the feedforward neural network of the feature encoding layer. The feedforward neural network contains multiple fully connected layers, each of which uses a nonlinear activation function to process the input feature vector.

[0087] First, the first fully connected layer performs a linear transformation on the interactive fusion feature vector sequence to map the features to a higher-dimensional space; then, a nonlinear activation function is used to perform a nonlinear transformation on the transformed features, introducing nonlinear factors to enhance the expressive power of the features; then, the second fully connected layer maps the high-dimensional features back to the original dimensional space.

[0088] After being processed by the feedforward neural network, an intermediate feature vector sequence is generated. This intermediate feature vector sequence has stronger feature expression ability while retaining the original key information.

[0089] Step S1316: performing layer equalization processing on the intermediate feature vector sequence to obtain a coded feature vector with a unified representation form.

[0090] Each eigenvector in the intermediate eigenvector sequence may have differences in numerical range, distribution, etc. In order to make these eigenvectors have a unified representation, layer equalization processing is required.

[0091] Layer equalization uses a normalization operation to transform each eigenvector in the intermediate eigenvector sequence so that its mean is zero and its variance is one. Specifically, the mean and variance of the intermediate eigenvector sequence are calculated, and then the mean is subtracted from each eigenvector and divided by the square root of the variance to obtain the normalized eigenvector.

[0092] After layer equalization, an encoded feature vector with a unified representation is obtained, which can be better processed by the subsequent feature extraction layer.

[0093] Step S132: Receive the encoded feature vector through the element extraction layer of the specification generation model, perform the behavior subject definition operation, determine the subject scope and subject attribute characteristics to which the behavior specification is applicable, and generate a subject definition result.

[0094] After receiving the encoded feature vector, the feature extraction layer of the specification generation model first focuses on defining the subject of the behavior. By analyzing and extracting subject-related information from the encoded feature vector, it determines the scope of subjects to which the behavioral specification applies and the attributes and characteristics of these subjects, ultimately forming the subject definition results.

[0095] Step S1321: The encoded feature vector is received through the element extraction layer of the specification generation model, and matching processing is performed through a preset subject feature template to identify the feature components related to the behavior subject in the encoded feature vector to obtain the subject feature vector.

[0096] After receiving the encoded feature vector, the feature extraction layer calls the preset subject feature template. The subject feature template is pre-set based on a large amount of historical data and behavioral norms, and contains typical feature patterns related to the behavioral subject.

[0097] The coded feature vector is matched with the subject feature template, and the similarity between the two is calculated to identify the feature components related to the behavior subject in the coded feature vector. The above feature components together constitute the subject feature vector, which concentrates on reflecting the information related to the behavior subject.

[0098] Step S1322: Analyze the type attributes of the behavior subject based on the subject feature vector, determine the category identifier and category feature description to which the subject belongs, and generate a subject type definition result.

[0099] Conduct in-depth analysis of the subject feature vector to extract feature information that can reflect the subject type attributes. The above information may include the subject's composition form, functional role, etc.

[0100] Based on this characteristic information, the subject's category identifier is determined by referring to the preset subject type classification standards. For example, category identifiers may include "individual," "group," "organization," etc. At the same time, the characteristics of the category are described, such as the size of the group and the structural characteristics of the organization, to generate the subject type definition result.

[0101] Step S1323: Based on the subject type definition result, the constraint information related to the subject range in the coding feature vector is extracted, the quantity range, identity limitation and qualification conditions of the subject are determined, and the subject range definition result is generated.

[0102] Based on the results of the subject type definition, constraint information related to the subject scope is further extracted from the encoded feature vector. The above information may involve the number of subjects, identity requirements, required qualifications, etc.

[0103] Analyze this constraint information to determine the number of subjects, such as whether it is a single subject, multiple subjects, or a group of a specific size; clarify the subject's identity, such as a specific occupation or identity identifier; and determine the subject's qualifications, such as whether certain skills or certifications are required. This information is integrated to generate the subject scoping results.

[0104] Step S1324: Analyze the attribute feature information in the subject feature vector to identify the inherent attribute features and dynamic attribute features of the subject, wherein the inherent attribute features include the static feature parameters of the subject, and the dynamic attribute features include the state feature parameters of the subject.

[0105] The attribute feature information in the subject feature vector is analyzed to distinguish between inherent attribute features and dynamic attribute features. Intrinsic attribute features are relatively stable feature parameters inherent to the subject itself, such as the subject's physical features, structural features and other static feature parameters.

[0106] Dynamic attribute characteristics are characteristic parameters exhibited by the subject in different states, such as the subject's activity state and its adaptation to the environment. This distinction provides a more comprehensive understanding of the subject's attribute characteristics.

[0107] Step S1325: Integrate the subject type definition result, the subject scope definition result and the attribute characteristic information, and generate a subject definition result including subject identification, type description, scope limitation and attribute parameters according to the subject definition rules.

[0108] Integrate the results of subject type definition, subject scope definition and attribute feature information. Subject definition rules are pre-set and used to standardize the generation format and content requirements of subject definition results.

[0109] According to the subject definition rules, the integrated information is organized into a subject definition result including subject identification (used to uniquely identify the subject), type description (detailed description of the subject type), scope limitation (specific definition of the subject scope) and attribute parameters (specific parameters of the subject's inherent attributes and dynamic attributes).

[0110] Step S133: Based on the subject definition result, the temporal and spatial scope limitations of the behavior action are analyzed through the element extraction layer, the time interval limitation and spatial area limitation of the behavior are determined, and the behavior boundary conditions are generated.

[0111] Based on the subject definition results, the feature extraction layer further analyzes the temporal and spatial scope constraints of the behavior. By extracting and analyzing the time and space-related information in the encoded feature vector, the time interval and spatial region where the behavior occurred are determined, thereby generating the behavior boundary conditions.

[0112] Specifically, the time-related feature components in the encoded feature vector are analyzed to identify information such as the start time, end time, and duration of the behavior, and to determine the time interval limitation, such as "a specific period of time on weekdays" or "within several consecutive hours".

[0113] At the same time, the spatial-related feature components in the encoded feature vector are analyzed to identify the specific location and area where the behavior occurred, and to determine the spatial area limitation, such as "an area within a specific building" or "within a certain geographical range".

[0114] Integrate time interval limitations and spatial area limitations to generate behavioral boundary conditions and clarify the temporal and spatial scope of the behavior.

[0115] Step S134: combining compliance clues in the text-image collaborative features, identifying permitted and prohibited forms of the behavior, and constructing a compliance determination benchmark for the behavior, the compliance determination benchmark including a positive determination indicator and a negative determination indicator.

[0116] Compliance-related clues are extracted from the text-image collaborative feature set. These clues may include descriptions of permitted or prohibited behaviors in the text, examples of compliant or non-compliant behaviors shown in the images, etc.

[0117] Based on these compliance clues, the permitted forms of behavior are identified, that is, which behaviors are in compliance with regulatory requirements; at the same time, the prohibited forms of behavior are identified, that is, which behaviors are not in compliance with regulatory requirements.

[0118] Compliance determination criteria are established based on permitted and prohibited forms. Positive determination indicators are used to determine whether an action falls within permitted forms, including the specific manifestations and implementation methods of the action. Negative determination indicators are used to determine whether an action falls within prohibited forms, also including the specific manifestations and implementation methods of the action. These positive and negative determination indicators can clearly determine whether an action is compliant.

[0119] Step S135: Integrate the subject definition results, the behavioral boundary conditions and the compliance judgment criteria, and arrange and process them according to the preset standard text structure to generate a draft behavioral specification, which includes the clause number, clause content and corresponding schematic image index.

[0120] Integrate the subject definition results, behavioral boundary conditions, and compliance determination criteria. The preset regulatory text structure is based on the common behavioral code text format, including the order of clauses and content organization.

[0121] Arrange the integrated information according to the pre-set standard text structure. Assign each clause a unique clause number and clearly define the clause content, which should include information such as the subject, behavioral boundaries, and compliance determination. Also, add a corresponding diagram index to each clause, which can be used to link to diagrams related to the clause content.

[0122] After editing and processing, a draft code of conduct is generated, which fully presents all the elements of the code of conduct.

[0123] Step S140: performing consistency comparison processing on the draft code of conduct and a preset regulatory standard library to obtain a standard comparison result, which includes a conforming item identifier and a revised item identifier.

[0124] To ensure the accuracy and compliance of the draft code of conduct, it is necessary to compare it with the pre-set regulatory standards library. By comparing the content of the draft with the standards library, the consistency and differences between the two are identified and marked as conforming items and revised items respectively, forming the standard comparison results.

[0125] Step S141: Obtain a preset regulatory standard library, which contains historically valid behavioral standard texts and corresponding standard schematic diagrams. Each standard text contains standard clauses, scope of application and judgment basis.

[0126] Access a pre-defined regulatory standards library, which has been accumulated and collated over a long period of time and contains a large number of historically valid code of conduct standards. Each code of conduct standard specifies the standard clause, its scope of application, and the basis for determining whether an action complies with the clause.

[0127] At the same time, the standard library also contains a set of standard diagram pictures corresponding to the standard text. These pictures intuitively show the behavioral norms described in the standard clauses, helping to better understand the content of the standard.

[0128] Step S142: Perform clause splitting on the draft code of conduct to obtain multiple independent draft clause units. Each draft clause unit includes clause content, corresponding schematic pictures, and applicable conditions.

[0129] Perform clause splitting on the draft code of conduct. According to the logical relationship and content independence of the clauses, split the draft into multiple independent draft clause units.

[0130] Each draft clause unit contains complete information: the clause content is the specific regulation of the clause on the behavior; the corresponding schematic picture is the picture used to illustrate the clause content; the applicable conditions are the specific situations and premises to which the clause applies.

[0131] Step S143: Perform clause splitting on the standard text in the regulatory standard library to obtain multiple independent standard clause units. Each standard clause unit includes standard clause content, corresponding standard pictures, and standard applicable conditions.

[0132] Use the same method as splitting the draft code of conduct to perform clause splitting on the standard text in the regulatory standard library to obtain multiple independent standard clause units.

[0133] Each standard clause unit contains standard clause content, which is the specific regulation of the standard on the behavior; the corresponding standard picture, which is used to visually display the standard clause; the standard applicable conditions, which clarify the situations to which the standard clause applies.

[0134] Step S144: Perform semantic matching on the draft clause content of the draft clause unit and the standard clause content of the standard clause unit to generate a text similarity score.

[0135] Perform semantic matching on the draft clause content of the draft clause unit and the standard clause content of the standard clause unit. Through natural language processing technology, analyze the semantic similarity between the two, and generate a text similarity score. This text similarity score reflects the closeness of the meanings of the two text contents.

[0136] Step S1441: Perform word segmentation and词性标注处理(无法准确翻译该部分,可能是词性标注的意思) on the draft clause content of the draft clause unit, obtain an effective word sequence after removing stop words, and perform word vector conversion on each effective word to generate a draft word vector sequence.

[0137] Perform word segmentation on the draft clause content to split the text into individual words. Then, perform词性标注处理(无法准确翻译该部分,可能是词性标注的意思) on each word to clarify its词性(无法准确翻译该部分,可能是词性的意思), such as noun, verb, adjective, etc.

[0138] Remove stop words from the text. These words usually have no actual semantics or have little impact on semantics, such as "的", "在", "和", etc. The remaining words form an effective word sequence.

[0139] Each valid word in the valid word sequence is transformed into a word vector, converting the word into a corresponding vector representation. These vectors can reflect the semantic information of the word. Multiple word vectors are arranged in the original word order to generate a draft word vector sequence.

[0140] Step S1442: performing word segmentation and part-of-speech tagging on the standard clause content of the standard clause unit, removing stop words to obtain a standard word sequence, performing word vector conversion on each standard word, and generating a standard word vector sequence.

[0141] The same method as that used to process the content of the draft clauses is used to perform word segmentation and part-of-speech tagging on the content of the standard clauses, and the standard word sequence is obtained after removing stop words.

[0142] Perform word vector conversion on each standard word in the standard word sequence to generate a standard word vector sequence, which has the same vector dimension as the draft word vector sequence.

[0143] Step S1443: Calculate the cosine similarity between the draft word vector sequence and the word vectors at corresponding positions in the standard word vector sequence to obtain a word-level similarity matrix.

[0144] Calculate the cosine similarity between the draft word vector sequence and the corresponding word vector in the standard word vector sequence. Cosine similarity measures the directional similarity between two vectors. A larger cosine similarity indicates that the two vectors are oriented closer and the corresponding words are more semantically similar.

[0145] The cosine similarity values ​​of all corresponding position word vectors are arranged into a matrix form to obtain a word-level similarity matrix. Each element in the word-level similarity matrix represents the semantic similarity of the words at the corresponding position in the two word sequences.

[0146] Step S1444: performing an optimal path search process on the word-level similarity matrix, determining the best matching path between two word sequences, and generating a path matching score.

[0147] A dynamic programming algorithm is used to search for the optimal path in the word-level similarity matrix. The optimal path is the path from the upper left corner to the lower right corner of the word-level similarity matrix, where the sum of the cosine similarities is the largest, representing the most matching word correspondence between the two word sequences.

[0148] By calculating the sum of the cosine similarities on the optimal path and performing normalization, a path matching score is generated, which reflects the overall matching degree of the two word sequences.

[0149] Step S1445: extract the syntactic structure features of the draft clause content and the standard clause content, calculate the edit distance of the syntax tree, and obtain a structural similarity score.

[0150] Perform syntactic analysis on the draft clauses and standard clauses to generate their respective syntax trees. The syntax trees reflect the syntactic relationships between words in a sentence, such as subject-predicate relationships, verb-object relationships, etc.

[0151] Extract the structural features of the syntax tree, such as the depth of the tree, the number and type of nodes, the connection relationship between nodes, etc. Calculate the edit distance between two syntax trees. The edit distance refers to the minimum number of editing operations (such as inserting, deleting, and replacing nodes) required to transform one syntax tree into another.

[0152] According to the size of the edit distance, the conversion is used to obtain a structural similarity score. The smaller the edit distance, the higher the structural similarity score, which means that the syntactic structures of the two texts are more similar.

[0153] Step S1446: performing weighted sum processing on the path matching score and the structural similarity score according to preset weights to obtain preliminary text similarity.

[0154] The weights of the path matching score and the structural similarity score are preset according to their importance in semantic matching.

[0155] Multiply the path matching score by its corresponding weight, multiply the structural similarity score by its corresponding weight, and then add the two products to obtain the preliminary text similarity, which comprehensively reflects the similarity between the two texts in word semantics and syntactic structure.

[0156] Step S1447: performing interval mapping processing on the preliminary text similarity, mapping it into a preset numerical interval, and generating a text similarity score.

[0157] In this embodiment, a numerical range, such as between 0 and 1, is preset to indicate the degree of text similarity. The preliminary text similarity is then subjected to interval mapping, converting its value to a preset numerical range through methods such as linear transformation. The converted value serves as the text similarity score. A score closer to 1 indicates a greater semantic similarity between the draft clause and the standard clause; a score closer to 0 indicates a greater semantic difference between the two clauses.

[0158] Step S145: Perform visual matching processing on the schematic image corresponding to the draft clause unit and the standard image corresponding to the standard clause unit to generate an image similarity score.

[0159] Visual matching is performed between the schematic images corresponding to the draft clause units and the standard images corresponding to the standard clause units. Using image processing techniques, the degree of similarity in visual features between the two images is analyzed to generate an image similarity score, which reflects the degree of similarity between the two images in terms of content and visual expression.

[0160] Specifically, visual features of the two images can be extracted, such as color histograms, texture features, and shape features. The similarity between these visual features is then calculated, for example, by calculating the intersection distance of color histograms or the Euclidean distance of texture features. These similarity values ​​are then combined to produce an image similarity score.

[0161] Step S146: Calculate the comprehensive matching degree based on the text similarity score and the image similarity score. When the comprehensive matching degree reaches a preset threshold, mark the draft clause unit as a compliant item and add a compliant item identifier.

[0162] Set corresponding weights for text similarity scores and image similarity scores, and the size of the weights is determined according to the importance of text and image in matching.

[0163] Multiply the text similarity score by its weight, multiply the image similarity score by its weight, and then add the two results to get the overall match score.

[0164] A comprehensive matching degree threshold is preset. When the calculated comprehensive matching degree reaches the threshold, it indicates that the draft clause unit and the standard clause unit are relatively consistent in content and visuals. The draft clause unit is marked as a compliant item, and a compliant item identifier is added. The identifier may include the matching standard clause unit information.

[0165] Step S147: When the comprehensive matching degree does not reach the preset threshold, analyze the location and type of the difference point, mark the draft clause unit as a revision item and add a revision item identifier, which includes a difference type description and a reference standard index.

[0166] In this embodiment, if the comprehensive matching degree does not reach the preset threshold, it indicates that there are significant differences between the draft clause unit and the standard clause unit, and the differences need to be analyzed. First, the specific location of the differences is determined, whether they are differences in text content, differences in schematic images, or both.

[0167] For differences in text content, we can further analyze the type of differences. They may be deviations in semantic expression, such as inaccurate descriptions of behavioral actions; they may be different applicable conditions, such as discrepancies between the scope of application of draft clauses and standard clauses; they may also be differences in the basis for compliance determination, such as inconsistent indicators for determining behavioral compliance.

[0168] For differences in schematic pictures, the type of difference is also analyzed. It may be that the behavior shown in the picture does not match the standard picture, such as there is a deviation in the action posture; it may also be that the scene background of the picture is different from the standard picture, such as there is a difference in the details of the background environment.

[0169] After identifying the location and type of the difference, the draft clause unit is marked as a revision and a revision identifier is added. The difference type description in the revision identifier details the specific type and manifestation of the difference, and the reference standard index points to the corresponding standard clause unit in the regulatory standards library.

[0170] Step S148: Integrate the conforming item identifiers and revised item identifiers of all draft clause units to generate standard comparison results.

[0171] After completing the comparison of all draft clause units, the compliance or revision indicators for each unit are summarized. These indicators are combined in the order of the draft clause units to form a complete standard comparison result. This standard comparison result clearly shows which clauses in the draft code of conduct meet regulatory standards and which clauses require revision.

[0172] Step S150: iteratively adjust the draft code of conduct according to the correction item identifier in the code comparison result, generate a final code of conduct supervision instruction and send it to the supervision execution terminal.

[0173] Based on the correction items identified in the code comparison results, the draft code of conduct is modified and improved in a targeted manner. After multiple iterations of adjustments to ensure compliance with regulatory standards, it is converted into a command format recognizable by the regulatory execution terminal and sent to the terminal to achieve effective supervision of relevant behaviors.

[0174] Step S151: parse the amendment item identifier in the specification comparison result, and extract the draft clause unit, difference type description and reference standard index corresponding to each amendment item.

[0175] Parse the amendment item identifiers in the standard comparison results and extract the draft clause units corresponding to each amendment one by one to identify the specific clauses that need to be revised. Also, extract the difference type description to understand the specific type of difference between the clause and the standard clause, as well as the reference standard index to determine the standard clause unit to be used for revision.

[0176] For example, if a revision item mark shows that there is a semantic deviation between the behavioral action description of the draft clause unit and the standard clause, and the reference standard index points to a standard clause unit in the regulatory standard library, then after analysis it can be clearly seen that it is the behavioral action description part of the draft clause unit that needs to be modified, and the corresponding standard clause unit will be used as a reference.

[0177] Step S152: Retrieve the corresponding standard clause unit and standard schematic diagram from the regulatory standard library according to the reference standard index as a reference for revision.

[0178] Based on the reference standard index obtained through analysis, the corresponding standard clause unit and standard diagram image are accurately retrieved from the regulatory standards library. The standard clause content, standard application conditions, judgment basis, etc. in the standard clause unit, as well as the visual information displayed in the standard diagram image, together form the reference basis for the revision process, ensuring that the revised draft clause is consistent with the regulatory standards.

[0179] Step S153: For each amendment, compare the draft clause unit with the corresponding standard clause unit, identify the semantic differences and expression differences, and generate a difference analysis report.

[0180] For each amendment, the draft clause unit is carefully compared with the retrieved standard clause unit. In terms of text content, the clause content is compared word by word, identifying semantic differences, such as different meanings of the same behavior. At the same time, differences in expression are identified, such as differences in word usage and sentence structure.

[0181] For schematic images, we compare the corresponding schematic images of the draft clauses with the standard images to identify visual differences, such as the angle of action and background details. We then organize and categorize these differences, generating a detailed analysis report that clearly identifies the specific content, location, and manifestation of the differences.

[0182] Step S154: Based on the difference analysis report, the draft clause content of the draft clause unit is semantically revised in combination with the association mapping rule set, so that the revised draft clause content is consistent with the core semantics of the standard clause.

[0183] Based on the difference analysis report and the previously constructed association mapping rule set, the draft clause content of the draft clause unit is semantically revised. For the identified semantic differences, the core semantics of the standard clause are referenced and the wording of the draft clause is adjusted.

[0184] During the revision process, ensure that the revised content complies with the semantic consistency constraints in the association mapping rule set, that is, the semantics of the text description matches the corresponding visual elements of the image. For example, if the core semantics of the standard clause "Safety devices must be inspected before operating the equipment" emphasizes the necessity of inspection, while the draft clause states "Safety devices may be inspected while operating the equipment," the word "may" should be revised to "need" to align with the core semantics of the standard clause.

[0185] For example, step S1541: extracting semantic difference points from the difference analysis report, determining the expression parts in the content of the draft clause that conflict with the core semantics of the standard clause, and marking them as semantic segments to be revised.

[0186] Carefully review the difference analysis report and identify any semantic differences. Based on these differences, locate the specific portions of the draft clause that conflict with the core semantics of the standard clause and mark these portions as semantic fragments to be revised. These semantic fragments to be revised may be a single word, a phrase, or a complete sentence; their common characteristic is that they are inconsistent with the core semantics of the standard clause.

[0187] Step S1542: calling the semantic parsing model to perform deep semantic analysis on the semantic segment to be corrected, identifying the type and cause of the semantic conflict, and generating a semantic conflict analysis result.

[0188] The semantic parsing model is invoked, and the semantic segment to be corrected is input into the model for deep semantic analysis. By analyzing the syntactic structure, collocation of word meanings, and contextual associations of the segment, the model identifies the type of semantic conflict, such as conceptual confusion, incorrect scoping, and inverted logical relationships.

[0189] At the same time, analyzing the causes of semantic conflicts may be due to inaccurate understanding of the terms or insufficiently rigorous expressions.

[0190] Step S1543: Based on the semantic conflict analysis result and with reference to the corresponding standard clause content, a plurality of candidate revised semantic segments are generated.

[0191] Based on the results of the semantic conflict analysis and the corresponding standard clause content, we conceived a variety of possible revision schemes and generated multiple candidate revised semantic segments. Each candidate segment was adjusted to address the semantic conflict points, attempting to align the revised semantics with the core semantics of the standard clause.

[0192] For example, if the semantic segment to be revised has a semantic conflict due to conceptual confusion, multiple different expressions are generated as candidate revised semantic segments with reference to the accurate expression of the concept in the standard clause.

[0193] Step S1544: convert each candidate modified semantic segment into a corresponding visual feature description, check whether it matches the visual element of the schematic image corresponding to the clause, and filter out the candidate modified semantic segments that meet the association mapping rule set.

[0194] For each candidate modified semantic segment, it is converted into a corresponding visual feature description based on the correspondence in the association mapping rule set. Then, these visual feature descriptions are checked to see if they match the visual elements of the schematic image corresponding to the clause, such as whether the subject's outline features, action posture key points, and background area texture features are consistent with the visual feature description.

[0195] Filter out candidate revised semantic fragments that match the visual elements of the schematic image and comply with the association mapping rule set to ensure that the revised text content is consistent with the image information.

[0196] Step S1545: performing semantic fluency evaluation on the filtered candidate revised semantic segments, and calculating the grammatical correctness score and semantic coherence score of each segment.

[0197] After screening, candidate revised semantic segments are evaluated for semantic coherence. This evaluation covers both grammatical correctness and semantic coherence. Grammatical correctness primarily examines the completeness of sentence structure, appropriate wording, and correct punctuation. Semantic coherence primarily examines the logical flow and semantic coherence of the segment, both within the segment and within the context.

[0198] The grammatical correctness score and semantic coherence score of each candidate segment are calculated using preset evaluation indicators and algorithms.

[0199] Step S1546: Select the candidate revised semantic segment with the highest grammatical correctness score and semantic coherence score, and replace the semantic segment to be revised in the content of the draft clause to complete a single semantic revision.

[0200] Compare the grammatical correctness and semantic coherence scores of each candidate semantic segment for revision, select the candidate segment with the highest score, and use this candidate segment to replace the semantic segment to be revised in the draft clause, completing the single revision process for this semantic difference.

[0201] Step S155: Perform visual adjustment processing on the schematic image corresponding to the revised draft clause content so that the visual elements of the image match the semantic elements of the revised text and meet the collaborative constraint conditions.

[0202] After completing the text revisions, visual adjustments will be made to the corresponding illustrations of the draft clauses. Based on the semantic elements of the revised text, the image will be analyzed for areas that require adjustment, such as the subject's posture and background details.

[0203] Edit the image through image processing tools, such as adjusting the subject's action angle, modifying the distribution of objects in the background, etc., so that the visual elements of the image match the corrected text semantic elements and comply with the collaborative constraints in the association mapping rule set, such as semantic consistency constraints and spatiotemporal correlation constraints.

[0204] Step S156: The revised clause unit is re-compared with the regulatory standard library for consistency. If there are still revised items, the above revision process is repeated until all revised items are converted into compliant items.

[0205] The revised clause unit is then compared again with the corresponding standard clause unit in the regulatory standards library for consistency, and the overall matching degree is calculated. If the overall matching degree reaches the preset threshold, the clause unit has been converted to a compliant item; if it does not reach the threshold, it indicates that there are still discrepancies and the correction process of steps S151 to S155 needs to be repeated.

[0206] This process is repeated until all revised items are converted into compliant items, ensuring that all clauses in the draft code of conduct are consistent with the standards in the regulatory standards library.

[0207] Step S157: Integrate all clause units that have passed the comparison, rearrange them according to the logical order of the code of conduct, and generate the final code of conduct text and the corresponding schematic diagram set.

[0208] After all clause units have been compared, they are integrated and rearranged according to the inherent logical sequence of the code of conduct, such as the order of subject definition, behavioral boundaries, and compliance determination, to form a final code of conduct text with a clear structure and rigorous logic.

[0209] At the same time, the schematic diagrams corresponding to each clause unit are sorted out to form a schematic diagram set that matches the final code of conduct text. The picture set corresponds one-to-one with the clauses in the text, making it easier to understand and implement.

[0210] Step S158: Convert the final behavioral specification text into an instruction format recognizable by the regulatory execution terminal, add an execution time identifier and an execution scope identifier, and generate a final behavioral specification regulatory instruction.

[0211] According to the interface specifications and data format requirements of the regulatory execution terminal, the final behavior specification text is converted into an instruction format that can be recognized by the terminal, such as a specific XML format, JSON format, etc.

[0212] The converted instructions are marked with an execution time to clarify the effective and expiration dates of the regulatory instructions; and an execution scope mark is added to clarify the geographical scope and subject scope of the instructions. These marks ensure that the regulatory instructions can be executed at the correct time and scope.

[0213] Step S159: Send the final behavior specification supervision instruction to the supervision execution terminal, triggering the terminal's specification storage and execution preparation operations.

[0214] The final code of conduct supervision instructions are sent to the supervision execution terminal through the network communication protocol. After receiving the instructions, the terminal first verifies the instructions to check their integrity and legality.

[0215] After verification, the terminal triggers the standardized storage operation, storing the supervision instructions in the terminal's local database for subsequent query and call. At the same time, execution preparation operations are performed, such as loading the relevant execution program and initializing execution parameters, to prepare for subsequent behavior supervision.

[0216] Pre-training the norm generation model is a key step in ensuring that it can accurately extract behavioral norm elements. It specifically includes the following steps: Step S211: Collect a large amount of behavioral code sample data, including historical behavioral code documents, corresponding schematic diagrams, and related compliance judgment cases, etc., to build a model training data set.

[0217] In this embodiment, the collected data must cover different types of behaviors, subjects, and scenarios to ensure the generalization ability of the model. The collected data is cleaned and preprocessed, such as removing duplicate data, correcting erroneous information, segmenting and normalizing text, and resizing and converting images.

[0218] Step S212: Divide the training data set into a training set, a validation set, and a test set, wherein the training set is used to learn model parameters, the validation set is used to adjust parameters during model training, and the test set is used to evaluate the final performance of the model.

[0219] The division ratio can be set according to the amount of data, for example, a ratio of 7:2:1 can be used for division.

[0220] Step S213: Construct the network structure of the specification generation model, which includes a feature encoding layer and a factor extraction layer. The feature encoding layer adopts a Transformer structure with a multi-head self-attention mechanism, and the factor extraction layer adopts a structure combining a multi-layer perceptron and a conditional random field.

[0221] Set the model's hyperparameters, such as the dimension of the feature vector, the number of heads in the multi-head self-attention mechanism, the number of neurons in the hidden layer, the learning rate, the number of iterations, etc.

[0222] Step S214: Input the text-image collaborative features from the training set into the model for training. During the training process, the model parameters are continuously adjusted through the backpropagation algorithm to ensure that the subject definition results, behavioral boundary conditions, and compliance judgment criteria output by the model are as consistent as possible with the annotation results in the training data.

[0223] Use the validation set to monitor the training effect of the model. When the loss function value on the validation set no longer decreases, stop training to avoid overfitting of the model.

[0224] Step S215: Use the test set to perform performance evaluation on the trained model. The evaluation indicators include the accuracy of subject definition, the accuracy of behavior boundary condition extraction, the accuracy of compliance judgment benchmark generation, etc.

[0225] Based on the evaluation results, if the model performance does not meet the preset requirements, the model's network structure or hyperparameters are adjusted and retrained; if the requirements are met, the model parameters are saved and the pre-training of the standard generation model is completed.

[0226] The purpose of training the bidirectional attention mechanism model is to improve its accuracy in calculating the correlation between text semantic elements and image visual elements. The specific steps are as follows: Step S311: Collect paired sample data of text semantic elements and image visual elements. These sample data must contain known association labels to construct a model training data set.

[0227] Preprocess the data, such as converting text semantic elements into vector form, performing feature extraction and vector conversion on image visual elements, etc.

[0228] Step S312: Divide the training set, validation set and test set. The division method refers to step S212.

[0229] Step S313: Construct the network structure of the bidirectional attention mechanism model, which includes a text encoding subnetwork, an image encoding subnetwork, and a bidirectional attention calculation subnetwork.

[0230] Set the model's hyperparameters, such as the dimension of the encoding vector, the number of attention heads, the learning rate, etc.

[0231] Step S314: Input the text semantic elements and image visual elements in the training set into the model for training, and adjust the model parameters through the back propagation algorithm to minimize the error between the correlation degree output by the model and the annotation correlation degree of the sample.

[0232] Use the validation set to monitor the model training process and prevent overfitting.

[0233] Step S315: Use the test set to evaluate the model's performance, using the mean square error of the correlation calculation as the evaluation metric. Adjust the model based on the evaluation results until the model performance meets the requirements, and save the model parameters.

[0234] Figure 2 A schematic diagram illustrates exemplary hardware and software components of a system 100 for collaboratively defining behavioral norms using natural language and images, which can implement the concepts of the present invention, as provided in some embodiments of the present invention. For example, a processor 120 can be used in the system 100 for collaboratively defining behavioral norms using natural language and images, and for executing the functions described in the present invention.

[0235] The natural language and image collaborative definition behavior norm supervision system 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the natural language and image collaborative definition behavior norm supervision large model system of the present application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0236] For example, the natural language and picture collaborative definition behavior norms supervision system 100 may include a network port 110 connected to the network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the natural language and picture collaborative definition behavior norms supervision system 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The natural language and picture collaborative definition behavior norms supervision system 100 also includes an I / O interface 150 between the computer and other input and output devices.

[0237] For ease of explanation, only one processor is described in the natural language and picture collaborative definition of behavioral norms supervision system 100. However, it should be noted that the natural language and picture collaborative definition of behavioral norms supervision system 100 in this application can also include multiple processors, so the steps performed by one processor described in this application can also be performed jointly or individually by multiple processors. For example, if the processor of the natural language and picture collaborative definition of behavioral norms supervision system 100 executes step A and step B, it should be understood that step A and step B can also be performed jointly by two different processors or individually in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor execute steps A and B together.

[0238] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, a large-scale model system for behavioral norm supervision that is collaboratively defined by natural language and pictures as described above is implemented.

[0239] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.

Claims

1. A large-scale model system for behavioral regulation that is collaboratively defined by natural language and images, characterized by: A method for use in the supervisory system is included, the method comprising: Constructing an association mapping rule set between natural language descriptions and standard schematic images, wherein the association mapping rule set includes correspondences between text semantic elements and image visual elements and collaborative constraints; Receiving a behavior description text and a scene diagram image of a scene to be supervised, and performing a bidirectional conversion process on the behavior description text and the scene diagram image according to the association mapping rule set to obtain a text-image collaborative feature set; Based on the text-image collaborative feature set, a pre-trained norm generation model is called to perform a behavioral norm element extraction operation to generate a behavioral norm draft including behavioral boundary conditions and compliance determination benchmarks; Performing a consistency comparison between the draft code of conduct and a preset regulatory standard library to obtain a standard comparison result, wherein the standard comparison result includes a conforming item identifier and a revised item identifier; The draft code of conduct is iteratively adjusted according to the correction item identifier in the code comparison result, and a final code of conduct supervision instruction is generated and sent to the supervision execution terminal.

2. The large-scale model system for behavior regulation based on collaborative definition of natural language and images according to claim 1 is characterized in that: The step of constructing a set of association mapping rules between the natural language description and the standard schematic diagram includes: Collect natural language description samples and corresponding standard diagram image samples in historical behavior standard files to form a sample set, wherein the natural language description samples include behavior subjects, behavior actions, and scene qualifiers, and the standard diagram image samples include behavior subject images, action posture images, and scene background images; Performing semantic word segmentation on the natural language description sample, extracting the part-of-speech tag and semantic role labeling of each word segmentation unit, and obtaining text semantic elements, wherein the text semantic elements include a subject reference identifier, an action verb identifier, and a scene modification identifier; Performing visual feature extraction processing on the standard schematic picture sample to identify the subject contour area, action posture key points and scene background area in the picture to obtain picture visual elements, wherein the picture visual elements include subject contour features, posture key point coordinates and background area texture features; Invoking a bidirectional attention mechanism model to calculate the correlation between the text semantic elements and the image visual elements to generate a semantic visual correlation matrix, wherein the elements in the semantic visual correlation matrix represent the matching strength between the text semantic elements and the image visual elements; According to the semantic visual association matrix, the text visual element pairs whose matching strength meets the preset conditions are selected, and the corresponding relationship between the text semantic elements and the image visual elements is constructed in combination with the behavioral norm logical relationship; Based on the corresponding relationship, collaborative constraints for text description and picture display are set, and the collaborative constraints include semantic consistency constraints, spatiotemporal correlation constraints, and expression integrity constraints. The corresponding relationship and the collaborative constraints are integrated to form an association mapping rule set.

3. The large-scale model system for behavior regulation based on collaborative definition of natural language and images according to claim 1 is characterized in that: The receiving of a behavior description text and a scene diagram image of a scene to be supervised, and performing a bidirectional conversion process on the behavior description text and the scene diagram image according to the association mapping rule set to obtain a text-image collaborative feature set, including: Receiving a behavior description text of a scenario to be regulated, performing syntactic structure analysis on the behavior description text, identifying a subject-verb-object structure and modifying elements, and extracting semantic elements of the text to be processed, wherein the semantic elements of the text to be processed include an identifier of the subject to be regulated, a description of the action to be regulated, and a description of the scenario to be regulated; Receive a scene schematic image of a scene to be supervised, perform image segmentation processing on the scene schematic image to separate the subject area, the action area, and the background area, and extract visual elements of the image to be processed, wherein the visual elements of the image to be processed include the outline of the subject to be supervised, the posture of the action to be supervised, and the background features of the scene to be supervised; According to the corresponding relationship in the association mapping rule set, the semantic elements of the text to be processed are converted into corresponding visual feature descriptions to generate a text-to-visual feature set, wherein the text-to-visual feature set includes subject visualization parameters, action visualization parameters, and scene visualization parameters; According to the corresponding relationship in the association mapping rule set, the visual elements of the image to be processed are converted into corresponding semantic feature descriptions to generate a visual-to-text feature set, wherein the visual-to-text feature set includes a subject semantic label, an action semantic label, and a scene semantic label; Calling a feature fusion module to perform collaborative verification processing on the text-to-visual feature set and the visual-to-text feature set, eliminating conflicting features that violate collaborative constraints, and obtaining a preliminary collaborative feature set; The preliminary collaborative feature set is subjected to feature dimension adaptation processing so that the expression dimensions of the text-derived features and the visual-derived features are consistent, and a text-image collaborative feature set including associated weight parameters is generated.

4. The large-scale model system for behavior regulation based on collaborative definition of natural language and images according to claim 1 is characterized in that: The pre-trained specification generation model is called based on the text-image collaborative feature set to perform a behavior specification element extraction operation to generate a behavior specification draft containing behavior boundary conditions and compliance determination benchmarks, including: Input the text-image collaborative feature set into the feature encoding layer of the pre-trained canonical generation model, model the correlation between the features through the multi-head self-attention mechanism, and generate an encoded feature vector; The encoding feature vector is received through the element extraction layer of the specification generation model, and a behavior subject definition operation is performed to determine the subject scope and subject attribute characteristics to which the behavior specification is applicable, and generate a subject definition result; Based on the subject definition results, the temporal and spatial scope limitations of the behavior actions are analyzed through the element extraction layer to determine the time interval limitation and spatial area limitation of the behavior, and generate the behavior boundary conditions; Combining compliance clues in text-image collaborative features, identifying permitted and prohibited forms of behavioral actions, and constructing a compliance determination benchmark for the behavioral actions, the compliance determination benchmark includes positive determination indicators and negative determination indicators; Integrate the subject definition results, the behavioral boundary conditions and the compliance judgment criteria, arrange and process them according to the preset standard text structure, and generate a draft behavioral code, which includes the clause number, clause content and the corresponding schematic image index.

5. The large-scale model system for behavior regulation based on collaborative definition of natural language and images according to claim 4 is characterized in that: The text-image collaborative feature set is input into the feature encoding layer of the pre-trained canonical generation model, and the correlation relationship between the features is modeled through a multi-head self-attention mechanism to generate an encoded feature vector, including: Performing feature splitting processing on the text-image collaborative feature set to obtain a text-derived feature subset and a visual-derived feature subset, wherein the text-derived feature subset includes semantic label features and association weight features, and the visual-derived feature subset includes visual parameter features and matching strength features; Inputting the text-derived feature subset and the visual-derived feature subset into an embedding module of a feature coding layer, performing vector conversion processing on each of them, and generating a text feature vector sequence and a visual feature vector sequence, wherein the dimensions of the text feature vector sequence and the visual feature vector sequence are consistent; Calling a multi-head self-attention mechanism to perform cross-attention calculation processing on the text feature vector sequence and the visual feature vector sequence to generate a text-visual interaction attention weight matrix, where the text-visual interaction attention weight matrix is ​​used to represent the correlation strength between text features and visual features; Performing weighted fusion processing on the text feature vector sequence and the visual feature vector sequence based on the text-visual interaction attention weight matrix to generate an interactive fusion feature vector sequence; Performing nonlinear transformation processing on the interactive fusion feature vector sequence through a feedforward neural network of a feature coding layer to enhance the expressive power of the features and generate an intermediate feature vector sequence; Layer equalization is performed on the intermediate feature vector sequence to obtain a coded feature vector with a unified representation form.

6. The large-scale model system for behavior regulation based on collaborative definition of natural language and images according to claim 4 is characterized in that: The element extraction layer of the specification generation model receives the encoded feature vector, performs a behavior subject definition operation, determines the subject scope and subject attribute characteristics to which the behavior specification applies, and generates a subject definition result, including: The encoding feature vector is received through the element extraction layer of the specification generation model, and matching processing is performed through a preset subject feature template to identify the feature components related to the behavior subject in the encoding feature vector to obtain the subject feature vector; Analyzing the type attributes of the behavior subject based on the subject feature vector, determining the category identifier and category feature description to which the subject belongs, and generating a subject type definition result; Extracting constraint information related to the subject range in the coded feature vector based on the subject type definition result, determining the number range, identity limitation and qualification conditions of the subjects, and generating a subject range definition result; Analyzing the attribute feature information in the subject feature vector to identify the inherent attribute features and dynamic attribute features of the subject, wherein the inherent attribute features include static feature parameters of the subject, and the dynamic attribute features include state feature parameters of the subject; The subject type definition result, the subject scope definition result and the attribute characteristic information are integrated, and a subject definition result including a subject identification, type description, scope limitation and attribute parameters is generated according to the subject definition rules.

7. The large-scale model system for behavior regulation based on collaborative definition of natural language and images according to claim 1 is characterized in that: The draft code of conduct is compared with the preset regulatory standard library for consistency, and the standard comparison results are obtained, including: Obtain a preset regulatory standards library, which contains historically valid behavioral standard texts and corresponding standard diagrams. Each standard text contains standard clauses, scope of application, and judgment basis; Splitting the draft code of conduct into clauses to obtain multiple independent draft clause units, each of which includes clause content, corresponding schematic diagrams, and applicable conditions; The standard text in the regulatory standard library is split into clauses to obtain multiple independent standard clause units, each of which contains the content of the standard clause, the corresponding standard image and the applicable conditions of the standard; Performing semantic matching processing on the draft clause content of the draft clause unit and the standard clause content of the standard clause unit to generate a text similarity score; Performing visual matching processing on the schematic image corresponding to the draft clause unit and the standard image corresponding to the standard clause unit to generate an image similarity score; Calculating a comprehensive matching degree based on the text similarity score and the image similarity score; when the comprehensive matching degree reaches a preset threshold, marking the draft clause unit as a compliant item and adding a compliant item identifier; When the comprehensive matching degree does not reach the preset threshold, the location and type of the difference point are analyzed, the draft clause unit is marked as a revision item and a revision item identifier is added. The revision item identifier includes a description of the difference type and a reference standard index; Integrate the conforming item identifiers and revised item identifiers of all draft clause units to generate standard comparison results.

8. The large-scale model system for behavior regulation based on collaborative definition of natural language and images according to claim 7 is characterized in that: The performing semantic matching processing on the draft clause content of the draft clause unit and the standard clause content of the standard clause unit to generate a text similarity score includes: Performing word segmentation and part-of-speech tagging on the draft clause content of the draft clause unit, removing stop words to obtain a valid word sequence, performing word vector conversion on each valid word, and generating a draft word vector sequence; Performing word segmentation and part-of-speech tagging on the standard clause content of the standard clause unit, removing stop words to obtain a standard word sequence, performing word vector conversion on each standard word, and generating a standard word vector sequence; Calculating the cosine similarity between the draft word vector sequence and the word vectors at corresponding positions in the standard word vector sequence to obtain a word-level similarity matrix; Performing an optimal path search process on the word-level similarity matrix to determine the best matching path between two word sequences and generate a path matching score; Extracting syntactic structural features of the draft clause content and the standard clause content, calculating the edit distance of the syntax tree, and obtaining a structural similarity score; Performing weighted sum processing on the path matching score and the structure similarity score according to preset weights to obtain a preliminary text similarity; The preliminary text similarity is subjected to interval mapping processing, mapped into a preset numerical interval, and a text similarity score is generated.

9. The large-scale model system for behavior regulation based on collaborative definition of natural language and images according to claim 1 is characterized in that: The iterative adjustment of the draft code of conduct according to the revision item identifier in the code comparison result, generating a final code of conduct supervision instruction and sending it to the supervision execution terminal includes: Parsing the amendment item identifiers in the standard comparison results, extracting the draft clause unit, difference type description and reference standard index corresponding to each amendment item; Retrieving the corresponding standard clause unit and standard schematic diagram from the regulatory standard library according to the reference standard index as a reference for revision; For each amendment, compare the draft clause unit with the corresponding standard clause unit, identify the semantic differences and expression differences, and generate a difference analysis report; Based on the difference analysis report, semantically correct the content of the draft clauses in the draft clause unit in combination with the association mapping rule set, so that the revised content of the draft clauses is consistent with the core semantics of the standard clauses; Perform visual adjustments on the schematic images corresponding to the revised draft clauses so that the visual elements of the images match the semantic elements of the revised text and meet the coordination constraints; The revised clause unit is re-compared with the regulatory standard library for consistency. If there are still revised items, the above revision process is repeated until all revised items are converted into compliant items; Integrate all the clause units that have passed the comparison, rearrange them according to the logical order of the code of conduct, and generate the final code of conduct text and corresponding schematic diagram set; Convert the final behavioral specification text into an instruction format recognizable by the regulatory execution terminal, add an execution time identifier and an execution scope identifier, and generate a final behavioral specification regulatory instruction; The final behavior specification supervision instruction is sent to the supervision execution terminal, triggering the terminal's specification storage and execution preparation operations.

10. A large-scale model system for behavioral regulation that is collaboratively defined by natural language and images, characterized by: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement a large model system for behavioral norms supervision collaboratively defined by natural language and pictures as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Perimeter rule generation method and electronic equipment

    CN118552912A

  • Vision-text collaborative abstract generation method and system based on multi-modal learning

    CN119862861A

  • Research and development document processing method and device

    CN120087351A

  • Fabricated building component intelligent generation and real-time detection method and system based on multi-modal AI

    CN120449262A

  • Image description generation method and apparatus, device, medium, and product

    WO2023179308A1