A power material thinking chain soft label generation method and device

By combining multimodal neural network models with Bayesian networks and conditional random fields, the power material review process is quantified, and soft tags are generated. This solves the problems of inconsistent review standards and difficulty in tracing opinions, and makes the review process more explicit and accurate.

CN120689122BActive Publication Date: 2025-11-18JIANGSU ELECTRIC POWER INFORMATION TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511208540.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-18
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

In the process of reviewing power materials, the review standards are complex, the application of veto clauses is inconsistent, and the language style of review opinions is inconsistent, resulting in inaccurate review results and difficulty in tracing back to the source.

Method used

A multimodal neural network model is used to extract features from historical review data of power materials. Combined with the power material review reasoning rule base, Bayesian network, and conditional random field, expert review opinions are quantified, and soft tags for power material thinking chains are generated to make the review process explicit and traceable.

Benefits of technology

This process makes the review process explicit and structured, improving review efficiency and effectiveness, and ensuring the accuracy and traceability of review results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689122B_ABST
    Figure CN120689122B_ABST
Patent Text Reader

Abstract

The application discloses a kind of electric power material thinking chain soft label generation method and device, method includes: obtaining electric power material historical review data;Based on the pre-constructed multi-modal neural network model, the feature extraction and analysis of electric power material historical review data are carried out, and the first thinking chain is given;Combining electric power material review reasoning rule base, the first thinking chain is verified and corrected, and the second thinking chain is obtained;Fusion bayesian network and conditional random field, the dependency relationship between each step in the second thinking chain of electric power material to be reviewed data is quantified, and the basic similarity, dependency coefficient and uncertainty coefficient of each step are obtained, the label determination weight is determined in combination with preset, the soft label of each step is determined, and the electric power material thinking chain soft label is generated.The application is improved by thinking chain and soft label generation, quantitative review opinion and corresponding soft label are given, so that review process is traceable, and review effect is ensured to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of electric power material data analysis, and particularly relates to a power material thinking chain soft label generation method and device. BACKGROUND

[0002] Electric power materials are all technology-intensive products, involving complex technical fields such as high voltage, insulation, thermal stability, electromagnetic compatibility, and their quality standards are much higher than those of general industrial materials. The procurement of electric power materials needs to strictly control the review process under highly unified technical specifications. At present, in the actual bid review process in the field of electric power materials, there are problems such as "complex review standards and non-uniform application of rejection clauses". Due to the differences in document structure and inconsistent content of different batches and different bid documents, the evaluation experts often rely on personal experience in judgment, which is prone to situations such as unclear reference clauses of rejection opinions and misjudgment of application conditions. Moreover, the review opinions given by the evaluation experts in the review process are free writing, and the language style is not uniform, which increases the difficulty of reviewing the review opinions.

[0003] Patent application CN120219055A discloses an intelligent auxiliary bid review method and system for electric power bidding professionals, which includes: obtaining uploaded procurement information, and creating a procurement batch based on configured basic parameters and procurement review rules to generate a review task list with a unique batch number. Based on the review task list and the configured basic parameters, the context information of each review task is extracted to construct a context feature vector of the review task. The context feature vector is matched with the preset rejection rules to obtain a rejection suggestion list corresponding to each review task in the review task list. The rejection suggestion list is bound to the corresponding bid task, and the bid task bound with the rejection suggestion list is pushed to the user terminal. The user terminal receives the user's operation on each rejection suggestion to generate the final rejection opinion, and stores and archives the final rejection opinion, improves the review opinion generation efficiency, standardizes the rejection expression, and reduces the expert work intensity.

[0004] In the above-mentioned related prior art, by comparing a single bid document with the corresponding bid rules, the review result is not considered comprehensively, which makes the review result not accurate enough.

[0005] How to quantify the evaluation opinions of experts, make the evaluation process of electric power materials explicit, structured, and traceable, and ensure the evaluation effect and improve the evaluation efficiency is a problem to be solved at present. SUMMARY

[0006] In view of the defects in the prior art, the application provides a power material thinking chain soft label generation method and device, which comprises the following steps: obtaining power material historical review data; based on a pre-constructed multi-modal neural network model, performing feature extraction and analysis on the power material historical review data to give a first thinking chain; combining a power material review reasoning rule base, verifying and correcting the first thinking chain to obtain a second thinking chain; quantifying the dependency relationship between each step in the second thinking chain of the power material to-be-reviewed data by fusing a Bayesian network and a conditional random field to obtain a basic similarity, a dependency coefficient and an uncertainty coefficient of each step; determining a soft label of each step in the second thinking chain of the power material to-be-reviewed data according to a preset label determination weight, fusing the basic similarity, the uncertainty coefficient and the dependency coefficient, and generating a power material thinking chain soft label. Through the improvement of the dynamic thinking chain and the soft label generation, the power material bidding review scene is adapted, the "implicit thinking" of the expert review is converted into "explicit reasoning steps", the expert review opinion is quantified, the soft label of the corresponding step is given, the review process is made explicit and structured, and the review effect in the bidding process is ensured and the review efficiency is improved.

[0007] In a first aspect, the application provides a power material thinking chain soft label generation method, which specifically comprises the following steps:

[0008] Obtaining power material historical review data;

[0009] Based on a pre-constructed multi-modal neural network model, performing feature extraction and analysis on the power material historical review data to give a first thinking chain;

[0010] Combining a power material review reasoning rule base, verifying and correcting the first thinking chain to obtain a second thinking chain;

[0011] Fusing a Bayesian network and a conditional random field to quantify the dependency relationship between each step in the second thinking chain of the power material to-be-reviewed data to obtain a basic similarity, a dependency coefficient and an uncertainty coefficient of each step;

[0012] According to a preset label determination weight, fusing the basic similarity, the uncertainty coefficient and the dependency coefficient to determine a soft label of each step in the second thinking chain of the power material to-be-reviewed data, and generating a power material thinking chain soft label.

[0013] Further, the multi-modal neural network model comprises a multi-modal input channel, a shared encoder, a fusion layer and a decoder.

[0014] Based on a pre-constructed multi-modal neural network model, performing feature extraction and analysis on the power material historical review data to give a first thinking chain, which is obtained through the following steps:

[0015] The power material historical review data of the corresponding modal is respectively tensor-converted in the multi-modal input channel to form a plurality of tensors;

[0016] The shared encoder extracts semantic vectors of each group of tensors respectively, and transmits the semantic vectors to a fusion layer;

[0017] After splicing each group of semantic vectors, a fusion feature is formed through a plurality of cross-attention units;

[0018] Based on the fusion feature, the fusion feature is decoded in the decoder to obtain a natural language step and a target word group positioning, and a first thought chain is given.

[0019] Further, based on the fusion feature, the fusion feature is decoded in the decoder to obtain a natural language step and a target word group positioning, and a first thought chain is given, which specifically includes:

[0020] The fusion feature is linearly transformed to give an initial hidden state of the decoder;

[0021] The initial hidden state is used as the starting state of the decoder word group sequence generation to generate an initial word group and update the initial hidden state;

[0022] According to the current hidden state and the previously generated word group, a current word group is generated and the current hidden state is updated until a word group sequence containing all word groups is generated;

[0023] Based on the regular expression formed by the power material corpus, the word group sequence is matched and analyzed to give the target word group, obtain the natural language step and the target word group positioning, and give the first thought chain.

[0024] Further, the power material review reasoning rule base includes at least one of parameter and specification association rules, step order rules, threshold judgment rules, and step dependency rules;

[0025] In combination with the power material review reasoning rule base, the first thought chain is verified and corrected to obtain a second thought chain, which specifically includes:

[0026] According to each rule in the power material review reasoning rule base, each step in the first thought chain is verified, and if the current step does not conform to the current rule, a similar step to the current step is obtained for replacement to complete the correction of the current step;

[0027] When each step in the first thought chain is verified, a second thought chain is obtained.

[0028] Further, the dependence relationship between each step in the second thinking chain of the power material to be reviewed data is quantified by fusing the Bayesian network and the conditional random field, to obtain the basic similarity, the dependence coefficient and the uncertainty coefficient of each step, specifically including:

[0029] The similarity between each step in the second thinking chain of the power material to be reviewed data and the corresponding expert result vector is analyzed to give the basic similarity of each step.

[0030] According to the Bayesian network model, the dependence relationship of each step in the second thinking chain of the power material to be reviewed data is analyzed to give the uncertainty coefficient corresponding to each step.

[0031] Each step in the second thinking chain of the power material to be reviewed data is input into the conditional random field constructed in advance to obtain the transition probability from the previous step to the current step and determine the dependence coefficient of each step.

[0032] Further, the similarity between each step in the second thinking chain of the power material to be reviewed data and the corresponding expert result vector is analyzed to give the basic similarity of each step, specifically including:

[0033] Each step in the second thinking chain of the power material to be reviewed data is quantified to obtain a model result vector.

[0034] The expert result vector corresponding to the model result vector is searched from the expert intermediate result library to give at least one target result vector, wherein the expert intermediate result library is obtained by analyzing and quantifying the multi-modal power material historical review data.

[0035] The similarity of the model result vector and the target result vector is analyzed, and the highest similarity is selected as the basic similarity.

[0036] Further, the expert intermediate result library is obtained by the following steps:

[0037] Obtain the historical bidding documents of power materials and the technical specification book of power materials;

[0038] Perform entity recognition on the historical bidding documents of power materials to obtain structured file entity parameters;

[0039] Correlate the file entity parameters with the specification clauses in the technical specification book of power materials to obtain the power material specification knowledge graph, wherein the nodes of the power material specification knowledge graph are specification clauses or file entity parameters, and the edges of the power material specification knowledge graph are the correlation between the specification clauses and the file entity parameters.

[0040] Based on the modal types of historical review data of power materials, feature extraction is performed on the historical review data of power materials for each modality to obtain modal features corresponding to multiple modalities. Among them, the historical review data of power materials includes document entity parameters, specification clauses in technical specifications and review texts.

[0041] The modal features corresponding to multiple modalities are weighted and concatenated, and then integrated with the knowledge graph of power material specifications to obtain expert result vectors, forming an expert intermediate result library.

[0042] Furthermore, based on the Bayesian network model, the dependencies between each step in the second thought chain of the power materials to be reviewed are analyzed, and the uncertainty coefficients corresponding to each step are given, specifically including:

[0043] Construct a Bayesian network model, initialize the model parameters, and define the prior distribution;

[0044] Each step is input into the Bayesian network model. By combining the dependencies of each step with the prior distribution, the model prediction results are obtained and the posterior distribution of each step is given.

[0045] Sample from the posterior distribution and calculate the variance of the model predictions in the sample set;

[0046] The variance of the model prediction results in the sample set is normalized to obtain the uncertainty coefficient.

[0047] Furthermore, the various steps in the second thought chain of the power materials to be reviewed are input into a pre-constructed conditional random field to obtain the transition probability from the previous step to the current step and determine the dependency coefficient of each step, specifically including:

[0048] Input each step of the second thinking chain of the power materials to be reviewed into a pre-constructed conditional random field to obtain the transition probability from the previous step to the current step and determine it as the dependency coefficient of the current step.

[0049] The construction of conditional random fields specifically includes:

[0050] Define the observation sequence and the hidden state sequence, where the observation sequence includes each step and the hidden state sequence includes the dependencies between the steps;

[0051] Based on the conditional probability of conditional random fields, a feature function is set;

[0052] By combining historical review data of power materials from multiple modes, a conditional random field is trained, and the values ​​of the feature functions are extracted until convergence.

[0053] Furthermore, the label determination weights include a first weight, a second weight, and a third weight;

[0054] Based on the pre-defined labels, weights are determined, and basic similarity, uncertainty coefficient, and dependency coefficient are integrated to determine the soft labels for each step in the second thinking chain of the power materials to be reviewed, generating soft labels for the power materials thinking chain, specifically including:

[0055] Calculate the product of the basic similarity of the current step and the corresponding first weight to obtain the first parameter;

[0056] The second parameter is obtained by multiplying the uncertainty coefficient and the corresponding second weight.

[0057] Calculate the product of the dependency coefficient of the current step and the corresponding third weight to obtain the third parameter;

[0058] Summing the first, second, and third parameters of the current step yields the soft label for the current step in the second thinking chain of power material data to be reviewed, generating a soft label for the power material thinking chain.

[0059] Furthermore, the multimodal neural network model is obtained through the following steps:

[0060] Perturb a portion of the historical review data for power supplies to generate adversarial examples;

[0061] By fusing adversarial samples with historical review data of power materials, an adversarial training set is obtained;

[0062] The initial network model was trained adversarially using an adversarial training set, and the comprehensive loss function was analyzed. The comprehensive loss function includes an adversarial loss function and a raw loss function. The adversarial loss function is the loss generated by the adversarial samples during training, and the raw loss function is the loss generated by the historical review data of power materials during training. The initial network model was trained using the historical review data of power materials.

[0063] If the overall loss function converges or reaches the maximum number of iterations, give the multimodal neural network model.

[0064] Secondly, the present invention also provides a power material thinking chain soft tag generation device, employing any of the above-mentioned power material thinking chain soft tag generation methods, including:

[0065] The data acquisition module is used to acquire historical review data of power materials;

[0066] The thought chain generation module is used to extract and analyze features from historical review data of power materials based on a pre-built multimodal neural network model, and to provide the first thought chain.

[0067] The thinking chain correction module is used to combine the power material review reasoning rule base to verify and correct the first thinking chain to obtain the second thinking chain.

[0068] The step analysis module is used to integrate Bayesian networks and conditional random fields to quantify the dependencies between the steps in the second thinking chain of the power materials to be reviewed data, and to obtain the basic similarity, dependency coefficient and uncertainty coefficient of each step.

[0069] The tag generation module is used to determine the weights based on preset tags, integrate basic similarity, uncertainty coefficient and dependency coefficient, determine the soft tags for each step in the second thinking chain of power materials to be reviewed, and generate soft tags for the power materials thinking chain.

[0070] The present invention provides a method and apparatus for generating soft tags for power material thinking chains, which has at least the following beneficial effects:

[0071] (1) By using a multimodal neural network model to extract and analyze the features of historical review data of power materials, a first thinking chain is given, the expert review opinions are quantified and the "implicit thinking" of the expert review is transformed into "explicit reasoning steps"; then, by integrating Bayesian network and conditional random field, the dependency relationship between each step in the second thinking chain of power materials to be reviewed is quantified, the accuracy of each step is evaluated from different aspects, and the weight is determined by combining the preset labels, and the soft labels of each step in the second thinking chain of power materials to be reviewed are determined. By improving the dynamic thinking chain and soft label generation, it is adapted to the bidding review scenario of power materials, so that the review process is traceable and the review effect in the bidding process is guaranteed.

[0072] (2) By using a multimodal neural network model, the historical review data of power materials in various modes are processed and the features in different modes are extracted. Multimodal data provides more comprehensive review information, which helps to improve the accuracy of the review results. The features of different modes are fused together, and the first thinking chain is constructed using the fused multimodal features. Integrating the features of different modes can reduce the bias that may be caused by a single mode and enhance the robustness of the model. At the same time, the first thinking chain can explicitly show the review logic, making the review logic traceable while improving the accuracy and reliability of the review.

[0073] (3) By using conditional random fields, the dependencies between steps are explicitly simulated and the rationality and coherence of these steps are evaluated, thereby improving the coherence between each step. Bayesian networks are used to evaluate the uncertainty and causal relationship in each step, and probabilistic reasoning results are obtained, which helps to understand the causal logic of the reasoning steps. At the same time, the structure and parameters of Bayesian networks provide the interpretability of the model, which helps to understand the rationality between each step. Attached Figure Description

[0074] Figure 1 A flowchart of the method for generating soft tags for the power material mind chain provided in an embodiment of the present invention;

[0075] Figure 2 A flowchart for determining the first thought chain provided in an embodiment of the present invention;

[0076] Figure 3 An architecture diagram of the multimodal neural network model provided in an embodiment of the present invention;

[0077] Figure 4 A flowchart for quantifying the dependencies between steps provided in an embodiment of the present invention;

[0078] Figure 5 A flowchart for determining the soft tags in each step of an embodiment of the present invention;

[0079] Figure 6 This is a structural block diagram of the power material thinking chain soft tag generation device provided in an embodiment of the present invention.

[0080] Among them, 201 is the data acquisition module; 202 is the thought chain generation module; 203 is the thought chain correction module; 204 is the step analysis module; and 205 is the tag generation module. Detailed Implementation

[0081] To better understand the above technical solutions, a detailed description of the solutions will be provided below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0082] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms, and “multiple” generally includes at least two unless the context clearly indicates otherwise.

[0083] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that an article or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device that includes said element.

[0084] A tender document is a written document issued by the tendering party to potential bidders. It aims to provide bidders with the information needed to prepare their tender documents and to clarify the tender requirements. It is the core document of the entire tendering process, specifying many key aspects such as the technical requirements, commercial terms, evaluation criteria, and contract terms of the tendered project.

[0085] From a technical perspective, the technical specifications and requirements in the tender documents need to be reasonable. For example, when procuring a complex set of industrial equipment, the technical parameters should accurately reflect the tenderer's needs for equipment performance, compatibility, and safety. If the technical requirements are too high or too low, the tender may fail. Requirements that are too high may result in no suitable supplier being able to meet them, while requirements that are too low may lead the tenderer to purchase a product that does not meet actual usage needs.

[0086] A review panel composed of professionals from within the tendering entity (such as technical experts, business experts, and legal experts) reviews each section of the tender documents based on their expertise and experience. For example, technical experts will focus on whether the technical specifications meet industry standards, business experts will check whether the commercial terms are reasonable, and legal experts will review whether the tender documents comply with laws and regulations.

[0087] With the development of information technology, some bidding management software has acquired basic bidding document review capabilities. This software can use preset rules to check whether the bidding documents contain key clauses (such as evaluation criteria and bid validity periods), and whether there are obvious typographical errors or formatting issues. However, software review mainly plays a supporting role; for complex technical and legal issues, further human judgment is still required.

[0088] Different reviewers may have differing interpretations of the tender documents. For example, when judging the reasonableness of technical requirements, one technical expert might consider a certain parameter necessary, while another might deem it overly demanding. This subjective difference can lead to inconsistent review results, affecting the quality of the tender documents. Furthermore, tender documents involve knowledge from multiple fields, including technical, commercial, and legal aspects. If reviewers lack knowledge in a particular area, it may result in review errors. For instance, in tenders for emerging technology fields, reviewers may have limited understanding of the technical specifications and market conditions of new technologies, making it difficult to accurately determine the reasonableness of the technical requirements in the tender documents.

[0089] Currently, the quantitative standards for reviewing tender documents are not comprehensive enough. For example, when assessing the reasonableness of commercial terms, it is difficult to use a fixed value to measure whether the bid bond amount is appropriate. Different project sizes, industry characteristics, and other factors make it difficult to formulate unified quantitative standards, forcing reviewers to rely on experience and subjective judgment for evaluation.

[0090] This invention provides a method and apparatus for generating soft tags for the thinking chain of power materials. The method includes: acquiring historical review data of power materials; extracting and analyzing features from the historical review data of power materials based on a pre-constructed multimodal neural network model to generate a first thinking chain; verifying and correcting the first thinking chain by combining a power material review reasoning rule base to obtain a second thinking chain; quantifying the dependencies between each step in the second thinking chain of the power material data to be reviewed by integrating Bayesian networks and conditional random fields to obtain the basic similarity, dependency coefficient, and uncertainty coefficient of each step; determining weights according to preset tags, integrating the basic similarity, uncertainty coefficient, and dependency coefficient to determine the soft tags for each step in the second thinking chain of the power material data to be reviewed, and generating soft tags for the power material thinking chain. This invention addresses the core pain points of "implicit expert review logic and difficulty in quantifying and tracing conclusions" in the intelligent review scenario of bidding and procurement by improving the dynamic thinking chain and soft tag generation, thus adapting it to the bidding and review scenario of power materials. By transforming the "implicit thinking" of expert review into "explicit reasoning steps," quantifying expert review opinions, and mapping key nodes in the reasoning process into interpretable soft labels, the review process becomes traceable, achieving the goal of "traceable review logic, quantifiable results, and decision-making assistance," thus ensuring the effectiveness of the review to a certain extent.

[0091] like Figure 1 As shown in the figure, this embodiment of the invention provides a method for generating soft tags for the thinking chain of power materials. The specific steps are as follows:

[0092] S101: Obtain historical review data for power materials.

[0093] Among them, the historical review data of power materials includes various multimodal data related to the review opinions during the bidding process, such as the voice review records of review experts, expert discussion segments in video conferences, charts (such as equipment parameter tables, engineering drawings, etc.) and pictures (such as equipment appearance drawings, installation site drawings, etc.) in the tender documents.

[0094] S102: Based on a pre-built multimodal neural network model, feature extraction and analysis are performed on historical review data of power materials, and the first thought chain is given.

[0095] Furthermore, the multimodal neural network model includes multimodal input channels, a shared encoder, a fusion layer, and a decoder;

[0096] Based on a pre-built multimodal neural network model, feature extraction and analysis are performed on historical review data of power materials, providing the first thought chain, and referencing... Figure 2 Specifically, it includes:

[0097] In the multimodal input channel, the historical review data of power materials for the corresponding modes are processed by tensor transformation to form multiple sets of tensors;

[0098] The shared encoder extracts the semantic vectors of each group of tensors and transmits the semantic vectors to the fusion layer;

[0099] After the semantic vectors of each group are concatenated, they are combined into a fused feature through multiple layers of cross-attention units;

[0100] Based on the fusion features, the fusion features are decoded in the decoder to obtain the natural language steps and target word location, and the first thought chain is given.

[0101] The architecture of the above multimodal neural network model is as follows: Figure 3 As shown, historical review data of power materials in multiple modalities are input into the multimodal input channel. Tensor transformation is performed on the historical review data of power materials in each modality. The transformed results are entered into the shared encoder to extract the corresponding semantic vectors. Multiple semantic vectors are concatenated and fused in the fusion layer to obtain fused features. Finally, the fused features are decoded by the decoder to obtain the corresponding first thought chain.

[0102] In one specific implementation, historical review data of power materials, including text modalities, text table modalities, image modalities, and metadata modalities, is input into the multimodal input channel. In a specific example, the historical review data of power materials corresponding to the text modal includes portable document format (PDF) / optical character recognition (OCR) data such as technical specifications, tender documents, and commercial terms response forms; the historical review data of power materials corresponding to the text table modal includes comma-separated values ​​(CSV) / Excel data such as structured quotation forms, warranty period forms, and delivery schedule forms; the historical review data of power materials corresponding to the image modal includes image data such as scanned copies of qualification certificates, photos of product nameplates, and factory production site diagrams; and the historical review data of power materials corresponding to the metadata modal includes one-hot encoding / embedded vector format data such as timestamps, voltage levels, material categories, and supplier numbers.

[0103] The shared encoders include text encoders, table encoders, image encoders, and metadata encoders. The text encoder extracts sentence vectors or local n-tuples. In a specific example, the text encoder is a bidirectional encoder base model + bidirectional long short-term memory-convolutional neural network (BERT-base + Bi-LSTM-CNN). The table encoder converts tabular data into a format suitable for model input. In a specific example, the table encoder is a tabular neural network (TabNet) or a Transformer-Encoder. The image encoder removes fully connected layers (FC), retaining 2048-dimensional feature maps. In a specific example, the image encoder is a residual network with 50 layers (ResNet50). The metadata encoder processes and transforms metadata. In a specific example, the metadata encoder is a 2-layer feedforward neural network (multilayer perceptron, MLP).

[0104] The fusion layer includes Early-Fusion and Cross-Attention. Early-Fusion is used to concatenate vectors from multiple modalities obtained from the shared encoder. The Cross-Attention unit enables Transformer-Cross to query each modality, outputting a 512-dimensional fused feature.

[0105] The task header of the multimodal neural network model outputs multiple tasks, including a classification of completeness, with results for complete, partially missing, and misaligned binding. When the classification result is partially missing, the exact number of missing pages is given. The task header also includes providing a thought chain, which is decoded using two layers of gated recurrent units (GRUs) to generate a natural language thought chain token sequence, i.e., the first thought chain.

[0106] In a specific example, taking "the first step in the first thought chain is document integrity check" as an example, the entire process of feature extraction → analysis → generation is given. First, multimodal data is input, including text modal data of the "10kV Ring Main Unit Technical Specification" obtained by scanning OCR, text table modal data of the bid quotation table with 27 rows and 9 columns (including delivery date, warranty period, and unit price), image modal data of 3 scanned copies of the manufacturer's qualification certificate, and metadata modal data of "material category = ring main unit, voltage level = 10kV, bidding time = 2024-05-20".

[0107] Then, feature extraction is performed on the data from each modality. For text modality data, BERT is used, outputting z_T∈R^768 [classification token (CLS)]. For text table modality data, TabNet is used, outputting z_Tab∈R^128, for example, column attention weights highlight the "warranty period" column. For image modality data, ResNet50 is used, outputting z_I∈R^2048, and the final feature map is then subjected to global average pooling (GAP). For metadata modality data, MLP is used, outputting z_M∈R^64.

[0108] Next, the features obtained from each modality are fused and analyzed, and input into the fusion layer to obtain z_f=CrossAttn([z_T,z_Tab,z_I,z_M])∈R^512, where z_T is the semantic vector corresponding to the text modality, z_Tab is the semantic vector corresponding to the text table modality, z_I is the semantic vector corresponding to the image modality, and z_M is the semantic vector corresponding to the metadata modality. The softmax output of the integrity classifier is [0.87,0.11,0.02]. The probability of the classification result being complete is 0.87, the probability of the classification result being partially missing is 0.11, and the probability of the classification result being misaligned is 0.02. Therefore, the classification result is complete.

[0109] Finally, the first thought chain is generated and decoded using a chain-head mechanism. Inputting z_f and the start symbol <|start|>, the fused features are decoded step-by-step to obtain:

[0110] Step-1 word: "First"; Step-2 word: "Check"; Step-3 word: "Cover"; ...; The final sentence is: First, check the cover, table of contents, tender letter, technical specifications, and commercial terms response form, a total of 5 key items; then compare the page numbers of the table of contents from 1 to 46 to ensure they are consecutive and without missing pages; next, check the clarity of 3 scanned copies of the qualification certificates, which are greater than 300 DPI; based on the overall judgment, the document is complete and without any missing parts. Therefore, the final sentence is the first thought chain automatically generated by the multimodal neural network model.

[0111] Furthermore, based on the fusion features, the fusion features are decoded in the decoder to obtain the natural language steps and target phrase localization, providing the first thought chain, which specifically includes:

[0112] A linear transformation is applied to the fused features to provide the initial hidden state of the decoder;

[0113] The initial hidden state is used as the starting state for generating the decoder phrase sequence. Initial phrases are generated and the initial hidden state is updated.

[0114] Based on the current hidden state and the previously generated phrase, generate the current phrase and update the current hidden state until a phrase sequence containing all phrases is generated;

[0115] Based on regular expressions formed from power material corpora, the word sequence is matched and analyzed to identify the target word, obtain the natural language steps and target word location, and provide the first thought chain.

[0116] The process of decoding the fused features specifically includes:

[0117] First, the fused features are linearly transformed to obtain the initial hidden state, that is, the fused features z_f are linearly mapped to obtain the initial hidden state. Where φ() is the activation function, W_h is the weight matrix used to learn the relationship between the fused features and the initial hidden state, and b_h is the bias vector. Then, a joint vocabulary is constructed: vocabulary V = {ordinary words} ∪ {target words} ∪ {<page_i> ,<col_j> ,<img_k> Next, GRU is used for step-by-step decoding, specifically as follows:

[0118] ;

[0119] Among them, \(h_t\) is the hidden state at the current time step \(t\), \(x_t\) is the input feature at the current time step \(t\), \(h_{t - 1}\) is the hidden state at time step \(t - 1\), \(y_t\) is the output distribution at the current time step \(t\), which is a probability distribution representing the probability of each possible output. The softmax function is used to convert a vector into a probability distribution. \(W_{out}\) is the weight matrix of the output layer, used to map the hidden state to the output space, \(b_{out}\) is the bias vector of the output layer, and the argmax function is used to find the index of the element with the maximum probability in the probability distribution. \(x_{t + 1}\) is the input feature at time step \(t + 1\). Then, regular extraction and localization are performed, using multiple regular expressions for extraction and localization, specifically including regular expressions for determining page numbers, column indexes, and image indexes. For example, Regular Expression 1: <page_(\d+)(-(\d+))?> → page number; Regular Expression 2: <col_(\d+)> → column index; Regular Expression 3: <img_(\d+)> → image index. In other embodiments, other regular expressions can also be configured according to the actual situation, and no limitation is made in this regard. Finally, natural language explanations are used for splicing. Generate the corresponding word sequence for the words, and according to different word extraction and localization, finally splice the words after extraction and localization into a "step + reason + location" triple. For example, the output is: S = "Step description, because [reason], see details in [location]."

[0120] In a specific example, first, the data of each modality is fused to obtain the fused feature, that is, the four types of information of text, table, picture, and metadata are all compressed into a 512-dimensional vector \(z_f\). Then, GRU is used to process the fused feature \(z_f\) to obtain the corresponding initial hidden state \(h0\). In order to accurately insert the corresponding content in the chain of thought, multiple regular expressions are used to locate the page number <page_i>, column index <col_j>, and image index <img_k>. For example, <page_23> is "page 23", <col_3> is "Excel column 3", and <img_1> is "scan image 1". Finally, characters are output sequentially using \(h0\), and each time a character is output, \(h0\) is updated once, and so on in a loop until the end symbol pops up to obtain the output result. After obtaining the output result, the regular expressions in the output result need to be translated completely. For example, when seeing <page_number>, translate it into "page number", when seeing <col_number>, translate it into "column number", and when seeing <img_number>, translate it into "image number". The content corresponding to the regular expressions is highlighted on the displayed PDF or table. Finally, adjust the translated output result to make it a complete review opinion. For example, "Checked pages 1 - 22 of the technical specification, found that page 23 of the bid letter is missing, and the clarity of scan image 1 of the certificate meets the requirements, so it is determined that the document is partially missing."

[0121] In other embodiments, the multimodal neural network model is further obtained through the following steps:

[0122] Perturb a portion of the historical review data for power supplies to generate adversarial examples;

[0123] By fusing adversarial samples with historical review data of power materials, an adversarial training set is obtained;

[0124] The initial network model was trained adversarially using an adversarial training set, and the comprehensive loss function was analyzed. The comprehensive loss function includes an adversarial loss function and a raw loss function. The adversarial loss function is the loss generated by the adversarial samples during training, and the raw loss function is the loss generated by the historical review data of power materials during training. The initial network model was trained using the historical review data of power materials.

[0125] If the overall loss function converges or reaches the maximum number of iterations, give the multimodal neural network model.

[0126] S103: Combining the power material review reasoning rule base, the first thinking chain is verified and corrected to obtain the second thinking chain.

[0127] Furthermore, the power material review reasoning rule base includes at least one of the following: parameter and specification association rules, step sequence rules, threshold judgment rules, and step dependency rules.

[0128] By combining the power material review reasoning rule base, the first thinking chain is verified and corrected to obtain the second thinking chain, which specifically includes:

[0129] Based on the rules in the power material review reasoning rule base, each step in the first thinking chain is verified. If the current step does not conform to the current rule, a similar step is obtained and replaced to complete the correction of the current step.

[0130] Once all steps in the first thought chain have been verified, the second thought chain is obtained.

[0131] Among them, the parameter-specification association rule indicates that a certain parameter or parameter type is associated with a certain specification clause, that is, when the parameter or parameter type appears, a corresponding specification clause must exist. This is expressed by a logical expression as follows:

[0132]

[0133] The step order rule indicates the sequential relationship between one step and another. For example, parameter extraction step A must precede canonical matching step B, which can be expressed as a logical expression:

[0134]

[0135] The threshold judgment rule determines whether a parameter value exceeds or falls below a corresponding threshold. If so, conflict analysis is triggered. For example, if parameter value D exceeds the corresponding threshold T, conflict analysis is triggered. This can be expressed as a logical expression:

[0136]

[0137] Step dependency rules indicate that a certain step must depend on one or more other steps. For example, the compliance determination step must follow the parameter extraction step and the specification matching step. This can be expressed as a logical expression:

[0138]

[0139] Iterate through all rules in the power material review reasoning rule base, judge the first thinking chain obtained by the multimodal neural network model, and check whether the first thinking chain satisfies the rules. If it does, determine the first thinking chain as the second thinking chain. If it does not, modify the first thinking chain and give the modified second thinking chain. The second thinking chain satisfies all rules in the power material review reasoning rule base.

[0140] In a specific example, there are parameter extraction and compliance judgment steps. The relationship between these two steps conforms to the standard, but a step between them is missing. In this case, according to the step dependency rules, a standard matching step is inserted, resulting in a first logical chain that conforms to the standard: parameter extraction step → standard matching step → compliance judgment step. If there is an error in the step order, the order of the steps is adjusted according to the step order rules. If there is a mismatch between parameters and standard clauses, the corresponding standard clauses are queried from the power materials standard knowledge graph and replaced. If a parameter value in a step exceeds or falls below the corresponding threshold, the parameter value is reread and re-judged according to the threshold judgment rules. If the judgment result still exceeds or falls below the corresponding threshold, a contradiction analysis is triggered, and the parameter value is marked or the parameter value or corresponding threshold is manually adjusted.

[0141] The generated first thought chain is logically verified and corrected according to the preset rules in the power material review reasoning rule base to ensure that the first thought chain meets the logical requirements of power material review. For example, the order of steps is checked to ensure it is reasonable and the matching relationship between parameters and specifications is accurate. If logical errors are found in a step (such as matching parameters with irrelevant specifications), the thought chain steps are corrected to ensure that they meet the logical requirements of power material review.

[0142] S104: By integrating Bayesian networks and conditional random fields, the dependencies between each step in the second thinking chain of power material data to be reviewed are quantified, and the basic similarity, dependency coefficient and uncertainty coefficient of each step are obtained.

[0143] Furthermore, by integrating Bayesian networks and conditional random fields, the dependencies between each step in the second thought chain of the power materials to be reviewed are quantified, yielding the basic similarity, dependency coefficient, and uncertainty coefficient of each step. Figure 4 Specifically, it includes:

[0144] Analyze the similarity between each step in the second thinking chain of power material data awaiting review and the corresponding expert result vector, and give the basic similarity of each step;

[0145] Based on the Bayesian network model, the dependencies of each step in the second thinking chain of the power materials to be reviewed are analyzed, and the uncertainty coefficients corresponding to each step are given.

[0146] By inputting each step of the second thought chain of the power materials to be reviewed into a pre-constructed conditional random field, the transition probability from the previous step to the current step is obtained and the dependency coefficient of each step is determined.

[0147] Understandably, the data on power materials awaiting review refers to relevant data that needs to be reviewed, including bidding documents, technical specifications, and other related content. Based on the second thought chain obtained from the aforementioned content, the data on power materials awaiting review is fed into the corresponding multimodal neural network model, and the output results are corrected. This yields the steps in the second thought chain for the data on power materials awaiting review. Next, the data on power materials awaiting review needs to be reviewed, and corresponding soft labels are provided.

[0148] Furthermore, the similarity between each step in the second thought chain of the power materials to be reviewed and the corresponding expert result vector is analyzed, and the basic similarity of each step is given, specifically including:

[0149] Quantify each step in the second thinking chain of the power materials to be reviewed to obtain the model result vector;

[0150] Search for the expert result vector corresponding to the model result vector from the expert intermediate result library, and give at least one target result vector. The expert intermediate result library is obtained by analyzing and quantifying the historical review data of multimodal power materials.

[0151] Analyze the similarity between the model result vector and the target result vector, and select the one with the highest similarity as the basic similarity.

[0152] In one specific implementation, the power material review data is substituted into the obtained second thinking chain to obtain the corresponding thinking review data. Each step in the thinking review data is quantified to obtain the model result vector. This quantization process can be implemented using the multimodal input channels and shared encoder of a multimodal neural network model, or it can be implemented using a bidirectional encoder, ensuring that the model result vector and the target result vector have the same format. For example, if the model result vector contains material types, a target result vector with the same material type as the model result vector is searched from the expert intermediate result library. There can be one or more target result vectors, and the search can be conducted based on material type in the expert intermediate result library or from other aspects; there are no limitations on this. Then, the similarity between the model result vector and the target result vector is calculated. In this example, cosine similarity is used to measure the similarity between the model result vector and the target result vector, specifically expressed as:

[0153] ;

[0154] in, For cosine similarity, For the model result vector, is the expert result vector, and n is the dimension of the model result vector and the target result vector.

[0155] For example, in the specification matching step, the parameter "rated voltage deviation (value +5.1%)" is matched with "specification DL / T 5432-2021 53.2.1 (threshold ≤ ±5%)". The model result vector uses BERT to encode the above text, resulting in a 300-dimensional vector P. T300 The vector P that retrieves the "standard matching step" and the material type "transformer" from the expert intermediate results database is the expert result vector. Z300 (The expert result vector represents the records of similar steps by experts in the historical review data of power materials). Calculate P. T300 and P Z300 The basic similarity.

[0156] Furthermore, the expert intermediate results database is obtained through the following steps:

[0157] Obtain historical bidding documents and technical specifications for power equipment;

[0158] Entity identification is performed on historical bidding documents for power materials to obtain structured document entity parameters;

[0159] By associating document entity parameters with the specification clauses in the power materials technical specification book, a power materials specification knowledge graph is obtained. The nodes of the power materials specification knowledge graph are specification clauses or document entity parameters, and the edges of the power materials specification knowledge graph are the association relationships between specification clauses and document entity parameters.

[0160] Based on the modal types of historical review data of power materials, feature extraction is performed on the historical review data of power materials for each modality to obtain modal features corresponding to multiple modalities. Among them, the historical review data of power materials includes document entity parameters, specification clauses in technical specifications, and review texts.

[0161] The modal features corresponding to multiple modalities are weighted and concatenated, and then integrated with the knowledge graph of power material specifications to obtain expert result vectors, forming an expert intermediate result library.

[0162] In one specific implementation, the BERT-NER model is used to perform entity recognition on historical bidding documents, extracting entity parameters (such as "rated voltage deviation," "temperature rise test value," "delivery date," and "warranty period") from the bidding documents to obtain structured document entity parameters including bid document number, material type, parameter type, parameter name, and parameter value. The BERT-NER model is first pre-trained on a corpus of power-related language containing technical documents for power equipment and power engineering standards and specifications, allowing the model to learn specific vocabulary, semantics, and named entity features in the power field.

[0163] By associating the extracted document entity parameters with the specification clauses in the technical specifications, a power materials specification knowledge graph is obtained. For example, "rated voltage deviation" is associated with Clause 3.2.1 of DL / T 5432-2021 (requirement: ≤±5%), and "delivery period" is associated with Clause 4.1.2 of the State Grid Distribution Network Engineering Procurement Specification (requirement: ≤60 days). The nodes in the power materials specification knowledge graph are specification clauses (attribute: clause number / threshold) and entity parameters (attribute: parameter name). The edges of the power materials specification knowledge graph represent the associations between specification clauses and document entity parameters. The relationship between specification clauses and document entity parameters in the technical specifications is analyzed, and association edges are established for the corresponding nodes (e.g., "rated voltage deviation" is associated with "DL / T 5432-2021 §3.2.1").

[0164] Feature extraction is performed on the multimodal data of historical review data of power materials, and finally feature fusion is performed to obtain the expert result vector. The specific feature extraction and feature fusion process is implemented through the multimodal input channel of the multimodal neural network model and the shared encoder, or it can be implemented through a bidirectional encoder, which will not be elaborated here.

[0165] It's important to understand that historical review data for power equipment includes the used regulatory clauses and various document entity parameters. By leveraging the modal features corresponding to multiple modalities, relevant parameters and / or regulatory clauses can be retrieved from a power equipment regulatory knowledge graph. Constructing a power equipment regulatory knowledge graph facilitates the establishment of relationships between historical review data for power equipment.

[0166] For example, modal features include document number, material type, parameter name, parameter type, and parameter value. Each feature dimension in the modal features is associated with a node in the power material specifications knowledge graph. For instance, when the parameter name is "rated voltage deviation," the power material specifications knowledge graph can quickly retrieve the associated specification clause DL / T 5432-2021, Clause 3.2.1 (requirement: ≤±5%). Similarly, when the parameter name is "delivery period," the power material specifications knowledge graph can quickly retrieve the associated specification clause, Clause 4.1.2 of the State Grid Distribution Network Engineering Procurement Specification (requirement: ≤60 days). Furthermore, when the material type is a transformer, the power material specifications knowledge graph can retrieve related specification clauses or parameter names.

[0167] Furthermore, based on the Bayesian network model, the dependencies between each step in the second thought chain of the power materials to be reviewed are analyzed, and the uncertainty coefficients corresponding to each step are given, specifically including:

[0168] Construct a Bayesian network model, initialize the model parameters, and define the prior distribution;

[0169] Each step is input into the Bayesian network model. By combining the dependencies of each step with the prior distribution, the model prediction results are obtained and the posterior distribution of each step is given.

[0170] Sample from the posterior distribution and calculate the variance of the model predictions in the sample set;

[0171] The variance of the model prediction results in the sample set is normalized to obtain the uncertainty coefficient.

[0172] In one specific implementation, Bayesian networks are used to model the uncertainties in the thought chain steps (such as input fuzziness and model parameter uncertainty).

[0173] To construct a Bayesian network, first define the nodes in the model and initialize the model parameters. These parameters include the node's parameter values, confidence level, correlation coefficient, weight matrix, rule threshold, and step correctness probability. Then, define the prior distribution and initialize the node distribution based on historical power material review data; for example, the model parameters may follow a Gaussian distribution.

[0174] By inputting each step of the second thought chain of the power materials to be reviewed data into the constructed Bayesian network model, the model prediction results and posterior distributions for each step can be obtained. Monte Carlo sampling is performed from the posterior distribution, and the variance of the model prediction results is taken as the uncertainty coefficient for the current step, specifically expressed as follows:

[0175] ;

[0176] Among them, U i Let y be the uncertainty coefficient of the i-th step in the second thought chain of the power materials to be reviewed data. k It is the value sampled from the posterior distribution of the i-th step, where K is the number of Monte Carlo samplings and μ is the mean of the K sampled values.

[0177] A Bayesian network model is used to model the uncertainty of each step. The inputs, model parameters, and outputs of each step are treated as nodes in the Bayesian network. The prior distribution of each step is determined based on prior knowledge (such as the model's performance on historical power material review data and expert reliability assessments of the step). After each step in the thought chain is executed, the posterior distribution of each node is updated using Bayes' theorem based on the actual output results. The variance of each step's output is calculated using the belief propagation algorithm, thereby quantifying the uncertainty of the current step and obtaining the uncertainty coefficient.

[0178] Furthermore, the various steps in the second thought chain of the power materials to be reviewed are input into a pre-constructed conditional random field to obtain the transition probability from the previous step to the current step and determine the dependency coefficient of each step, specifically including:

[0179] Input each step of the second thinking chain of the power materials to be reviewed into a pre-constructed conditional random field to obtain the transition probability from the previous step to the current step and determine it as the dependency coefficient of the current step.

[0180] The construction of conditional random fields specifically includes:

[0181] Define the observation sequence and the hidden state sequence, where the observation sequence includes each step and the hidden state sequence includes the dependencies between the steps;

[0182] Based on the conditional probability of conditional random fields, a feature function is set;

[0183] By combining historical review data of power materials from multiple modes, a conditional random field is trained, and the values ​​of the feature functions are extracted until convergence.

[0184] In one specific implementation, a conditional random field (CRF) is used to model the dependencies between steps (e.g., one step affects the next).

[0185] A Conditional Random Field (CRF) is used to model the dependencies between steps in the second thought chain of the power materials to be reviewed data. The sequence of steps in the second thought chain is used as the observation sequence, and the dependencies between steps (such as the influence of the output of step i on the input of step i+1) are used as the hidden state sequence. Then, the CRF is trained using historical power materials review data to learn the transition features and state features between each step. Combined with the feature function, the values ​​of the feature function are extracted until convergence.

[0186] By training a CRF (using historical thought chain step sequences and corresponding expert-annotated step dependencies as training samples), the transition probabilities between steps (such as the transition probability from the parameter extraction step to the specification matching step) are learned, and then the dependency patterns between steps are analyzed. For example, it was found that the compliance determination step usually depends on the results of the parameter extraction step and the specification matching step.

[0187] The probability formula for CRF is specifically expressed as follows:

[0188] ;

[0189] Where x is the observation sequence, y is the hidden state sequence, and f k It is a feature function used to capture the state and transition features between each step, λ k Z(x) represents the feature weights, and Z(x) represents the normalization factor.

[0190] In a specific example, the step sequence is parameter extraction step → canonical matching step → compliance determination step. The dependency relationship is that the parameter extraction step is the input of the canonical matching step, and the output of the canonical matching step is the input of the compliance determination step. Using a conditional random field, the transition probability of the parameter extraction step → canonical matching step is calculated to be 0.9, and the transition probability of the canonical matching step → compliance determination step is 0.8. Therefore, the dependency coefficient of the current step, i.e., the canonical matching step, is the average of the forward transition probability (0.9) and the backward transition probability (0.8), which is 0.85. In other examples, the product of the forward transition probability and the backward transition probability can also be used; there is no restriction on this.

[0191] S105: Determine the weights based on the preset labels, integrate the basic similarity, uncertainty coefficient and dependency coefficient, determine the soft labels for each step in the second thinking chain of the power materials to be reviewed, and generate soft labels for the power materials thinking chain.

[0192] Furthermore, the label determination weights include a first weight, a second weight, and a third weight;

[0193] Weights are determined based on preset labels. By integrating basic similarity, uncertainty coefficient, and dependency coefficient, soft labels are assigned to each step of the second thinking chain of the power materials to be reviewed, generating soft labels for the power materials thinking chain. Figure 5 Specifically, it includes:

[0194] Calculate the product of the basic similarity of the current step and the corresponding first weight to obtain the first parameter;

[0195] The second parameter is obtained by multiplying the uncertainty coefficient and the corresponding second weight.

[0196] Calculate the product of the dependency coefficient of the current step and the corresponding third weight to obtain the third parameter;

[0197] Summing the first, second, and third parameters of the current step yields the soft label for the current step in the second thinking chain of power material data to be reviewed, generating a soft label for the power material thinking chain.

[0198] By comprehensively considering basic similarity, uncertainty coefficient, and dependency coefficient, and combining them with labels to determine weights, soft labels are calculated for each step in the second thinking chain of the power materials to be reviewed. The value of the soft label reflects the degree of certainty regarding the correctness of the step. The soft label of each step is used to trace the issues raised in the review comments. For example, a low soft label for the "compliance determination step" indicates that the result of this step is problematic and needs to be reviewed, or the corresponding review comments should not be considered.

[0199] The soft label in step i i Specifically, it is expressed as:

[0200]

[0201] Where α is the first weight, S i U represents the base similarity at step i, β is the second weight, and U... i Let γ be the uncertainty coefficient for the current step i, γ be the third weight, and T be the uncertainty coefficient for the current step i. i Let be the dependency coefficient for the current step i.

[0202] The first, second, and third weights can be set according to the actual application scenario and the degree of importance. For example, the first weight corresponding to the basic similarity is 0.4, the second weight corresponding to the uncertainty coefficient is 0.3, and the third weight corresponding to the dependency coefficient is 0.3.

[0203] Reference Figure 6 This invention provides a device for generating soft tags for the thinking chain of power materials, comprising:

[0204] Data acquisition module 201 is used to acquire historical review data of power materials;

[0205] The thought chain generation module 202 is used to extract and analyze features from historical review data of power materials based on a pre-built multimodal neural network model, and to provide the first thought chain.

[0206] The thinking chain correction module 203 is used to combine the power material review reasoning rule base to verify and correct the first thinking chain to obtain the second thinking chain.

[0207] Step analysis module 204 is used to integrate Bayesian network and conditional random field to quantify the dependency relationship between each step in the second thinking chain of power material review data, and obtain the basic similarity, dependency coefficient and uncertainty coefficient of each step.

[0208] The tag generation module 205 is used to determine the weights based on preset tags, integrate basic similarity, uncertainty coefficient and dependency coefficient, determine the soft tags of each step in the second thinking chain of the power materials to be reviewed, and generate soft tags for the power materials thinking chain.

[0209] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0210] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and variations of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and variations.

Claims

1. A method for generating soft tags for power material thinking chains, characterized in that, include: Obtain historical review data for power equipment; Based on a pre-built multimodal neural network model, feature extraction and analysis are performed on historical review data of power materials, providing a first thought chain. The multimodal neural network model includes a multimodal input channel, a shared encoder, a fusion layer, and a decoder. The first thought chain is obtained through the following steps: Tensor transformation is performed on the historical review data of power materials for each modality in the multimodal input channel to form multiple sets of tensors; the shared encoder extracts the semantic vectors of each set of tensors and transmits the semantic vectors to the fusion layer; after concatenation of the semantic vectors, a fusion feature is formed through multiple layers of cross-attention units. Based on the fusion features, the fusion features are decoded in the decoder to obtain the natural language steps and target word location, and the first thought chain is given. By combining the power material review reasoning rule base, the first thinking chain is verified and corrected to obtain the second thinking chain; This paper analyzes the similarity between each step in the second thought chain of power material review data and the corresponding expert result vector, and gives the basic similarity of each step; constructs a Bayesian network model, initializes the model parameters and defines the prior distribution; inputs each step into the Bayesian network model, combines the dependencies of each step and the prior distribution to obtain the model prediction results and give the posterior distribution of each step; samples are taken from the posterior distribution and the variance of the model prediction results in the sample set is calculated. The variance of the model prediction results in the sample set is normalized to obtain the uncertainty coefficient; each step in the second thinking chain of the power materials to be reviewed is input into the pre-constructed conditional random field to obtain the transition probability from the previous step to the current step and determine the dependence coefficient of each step. Based on the preset labels, the weights are determined, and the basic similarity, uncertainty coefficient, and dependency coefficient are integrated to determine the soft labels for each step in the second thinking chain of the power materials to be reviewed, thus generating soft labels for the power materials thinking chain.

2. The method for generating soft tags for the power material mind chain as described in claim 1, characterized in that, Based on the fusion features, the fusion features are decoded in the decoder to obtain the natural language steps and target phrase localization, providing the first thought chain, which specifically includes: A linear transformation is applied to the fused features to provide the initial hidden state of the decoder; The initial hidden state is used as the starting state for generating the decoder phrase sequence. Initial phrases are generated and the initial hidden state is updated. Based on the current hidden state and the previously generated phrase, generate the current phrase and update the current hidden state until a phrase sequence containing all phrases is generated; Based on regular expressions formed from power material corpora, the word sequence is matched and analyzed to identify the target word, obtain the natural language steps and target word location, and provide the first thought chain.

3. The method for generating soft tags for the power material mind chain as described in claim 1, characterized in that, The power material review reasoning rule base includes at least one of the following: parameter and specification association rules, step sequence rules, threshold judgment rules, and step dependency rules. By combining the power material review reasoning rule base, the first thinking chain is verified and corrected to obtain the second thinking chain, which specifically includes: Based on the rules in the power material review reasoning rule base, each step in the first thinking chain is verified. If the current step does not conform to the current rule, a similar step is obtained and replaced to complete the correction of the current step. Once all steps in the first thought chain have been verified, the second thought chain is obtained.

4. The method for generating soft tags for the power material mind chain as described in claim 1, characterized in that, Analyze the similarity between each step in the second thought chain of the power materials to be reviewed and the corresponding expert result vector, and give the basic similarity of each step, specifically including: Quantify each step in the second thinking chain of the power materials to be reviewed to obtain the model result vector; Search for the expert result vector corresponding to the model result vector from the expert intermediate result library, and give at least one target result vector. The expert intermediate result library is obtained by analyzing and quantifying the historical review data of multimodal power materials. Analyze the similarity between the model result vector and the target result vector, and select the one with the highest similarity as the basic similarity.

5. The method for generating soft tags for the power material mind chain as described in claim 4, characterized in that, The expert intermediate results database is obtained through the following steps: Obtain historical bidding documents and technical specifications for power equipment; Entity identification is performed on historical bidding documents for power materials to obtain structured document entity parameters; By associating document entity parameters with the specification clauses in the power materials technical specification book, a power materials specification knowledge graph is obtained. The nodes of the power materials specification knowledge graph are specification clauses or document entity parameters, and the edges of the power materials specification knowledge graph are the association relationships between specification clauses and document entity parameters. Based on the modal types of historical review data of power materials, feature extraction is performed on the historical review data of power materials for each modality to obtain modal features corresponding to multiple modalities. Among them, the historical review data of power materials includes document entity parameters, specification clauses in technical specifications and review texts. The modal features corresponding to multiple modalities are weighted and concatenated, and then integrated with the knowledge graph of power material specifications to obtain expert result vectors, forming an expert intermediate result library.

6. The method for generating soft tags for the power material mind chain as described in claim 1, characterized in that, The label weighting process includes a first weight, a second weight, and a third weight. Based on the pre-defined labels, weights are determined, and basic similarity, uncertainty coefficient, and dependency coefficient are integrated to determine the soft labels for each step in the second thinking chain of the power materials to be reviewed, generating soft labels for the power materials thinking chain, specifically including: Calculate the product of the basic similarity of the current step and the corresponding first weight to obtain the first parameter; The second parameter is obtained by multiplying the uncertainty coefficient and the corresponding second weight. Calculate the product of the dependency coefficient of the current step and the corresponding third weight to obtain the third parameter; Summing the first, second, and third parameters of the current step yields the soft label for the current step in the second thinking chain of power material data to be reviewed, generating a soft label for the power material thinking chain.

7. A device for generating soft tags for power material thinking chains, characterized in that, The method for generating soft tags for the power material mind chain as described in any one of claims 1-6 includes: The data acquisition module is used to acquire historical review data of power materials; The thought chain generation module is used to extract and analyze features from historical review data of power materials based on a pre-built multimodal neural network model, and generate a first thought chain. The multimodal neural network model includes a multimodal input channel, a shared encoder, a fusion layer, and a decoder. The first thought chain is generated through the following steps: Tensor transformation is performed on the historical review data of power materials in the corresponding modality in the multimodal input channel to form multiple sets of tensors; the shared encoder extracts the semantic vectors of each set of tensors and transmits them to the fusion layer; the semantic vectors are concatenated and then processed through multiple layers of cross-attention units to form a fusion feature; based on the fusion feature, the decoder decodes the fusion feature to obtain the natural language steps and target phrase localization, thus generating the first thought chain. The thinking chain correction module is used to combine the power material review reasoning rule base to verify and correct the first thinking chain to obtain the second thinking chain. The step analysis module is used to analyze the similarity between each step in the second thinking chain of the power materials to be reviewed data and the corresponding expert result vector, and to give the basic similarity of each step; construct a Bayesian network model, initialize the model parameters and define the prior distribution; input each step into the Bayesian network model, combine the dependencies of each step and the prior distribution to obtain the model prediction results and give the posterior distribution of each step; sample from the posterior distribution and calculate the variance of the model prediction results in the sample set; normalize the variance of the model prediction results in the sample set to obtain the uncertainty coefficient; input each step in the second thinking chain of the power materials to be reviewed data into a pre-constructed conditional random field to obtain the transition probability from the previous step to the current step and determine the dependency coefficient of each step; The tag generation module is used to determine the weights based on preset tags, integrate basic similarity, uncertainty coefficient and dependency coefficient, determine the soft tags for each step in the second thinking chain of power materials to be reviewed, and generate soft tags for the power materials thinking chain.

Citation Information

Patent Citations

  • Intelligent auxiliary bid evaluation method and system for electric power bid invitation major

    CN120219055A

  • Bid evaluation system, terminal and method based on quantization rule and assistance

    CN114841730A

  • Construction method and device for domain-oriented knowledge graph

    CN119831022A