Purchase file verification method and device based on multi-mode and rule optimization

By combining multimodal feature extraction and semantic-visual collaborative processing with a procurement review decision model based on a differentiable reward function, the system addresses the shortcomings of multimodal data processing and rigid rule adaptation in traditional power procurement document review systems, thereby improving the accuracy and efficiency of the review process.

CN120806825AActive Publication Date: 2025-10-17JIANGSU ELECTRIC POWER INFORMATION TECH

Patent Information

Application Number
CN202511308317.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-17
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

Traditional power procurement document review systems have difficulty in fully capturing multimodal data, and their rule adaptation is rigid and inefficient, resulting in insufficient review accuracy and efficiency.

Method used

By extracting multimodal features from power procurement documents, semantic and visual collaborative feature extraction is performed. The results are then analyzed and judged using a procurement review decision model with a differentiable reward function, and a verification report is generated.

Benefits of technology

It enables efficient collaborative processing of multimodal data, improves the accuracy and robustness of the review process, enhances the transparency and operability of the review, and dynamically adjusts strategies to improve efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806825A_ABST
    Figure CN120806825A_ABST
Patent Text Reader

Abstract

The invention discloses a purchase file verification method and device based on multi-mode and rule optimization. The purchase file verification method comprises the steps of obtaining an electric power purchase file to be verified; extracting image features and non-image features in the to-be-verified power purchase file, and aligning the image features and the non-image features; performing semantic coding on the non-image features to obtain purchase semantic features; splicing the purchasing semantic features and the image features through a semantic and visual collaborative feature extraction strategy to obtain a purchasing vector; and based on a pre-constructed purchase audit decision model, in combination with the differentiable reward function, analyzing and judging the purchase vector, giving a purchase file verification result, marking the power purchase to-be-verified file, and generating a purchase file verification report. Efficient cooperative processing of multi-modal data in purchase file verification is achieved by extracting and aligning multi-modal features, semantic and visual association of purchase files is captured by using a semantic and visual cooperative feature extraction strategy, and the accuracy and robustness of auditing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent review, and particularly relates to a procurement document checking method and device based on multi-modal and rule optimization. BACKGROUND

[0002] Electric power procurement document review is that electric power material procurement experts review procurement technical specification documents according to issued procurement review points before procurement project bidding, and procurement personnel correct the technical specification documents according to the review results of the experts.

[0003] With the expansion of the scale of electric power material procurement and the improvement of equipment complexity, procurement review faces the dual challenges of multi-modal data integration and dynamic rule adaptation. Traditional review systems mostly rely on shallow analysis of single-modal data, and are difficult to fully capture key features in electrical diagrams and wiring diagrams. At the same time, rules in the field of equipment procurement are complex and frequently updated, and hard constraints and implicit business logic are intertwined, resulting in low efficiency and errors in manual review.

[0004] Patent application CN112446649A discloses a review method and device for material procurement plan, which includes: obtaining a material procurement plan to be reviewed, wherein the material procurement plan to be reviewed at least includes: description information of the material to be purchased, and procurement mode; determining the target review points corresponding to the description information of the material to be purchased according to the corresponding relationship between the pre-established description information of the material and the review points; reviewing the material procurement plan to be reviewed according to the pre-set review rules corresponding to the target review points to obtain the review result, wherein the review result is used to indicate whether the material procurement plan to be reviewed is qualified, and / or to mark the error information existing in the material procurement plan to be reviewed; and outputting the review result of the material procurement plan to be reviewed, solving the problems of low efficiency and accuracy of material procurement plan review.

[0005] The above related technology only compares and reviews the description information of the material and the review points, without considering other information in material procurement, and does not comprehensively review the points in material procurement. How to process and review the multi-modal information in material procurement to improve the review accuracy and efficiency of material procurement is a problem to be solved at present. SUMMARY

[0006] In view of the defects in the prior art, the application provides a procurement document verification method and device based on multi-modal and rule optimization, which comprises the following steps: obtaining a power procurement document to be verified; extracting image features and non-image features in the power procurement document to be verified to form multi-modal features, and aligning the multi-modal features; performing semantic coding on the non-image features to obtain procurement semantic features; splicing the procurement semantic features and the image features by a semantic and visual collaborative feature extraction strategy to obtain a procurement vector; analyzing and judging the procurement vector based on a pre-constructed procurement review decision model combined with a differentiable reward function to give a procurement document verification result, wherein the differentiable reward function is obtained by compiling hard constraints of procurement review; and labeling the power procurement document to be verified according to the procurement document verification result and generating a procurement document verification report.

[0007] By extracting multi-modal features in the power procurement document to be verified and aligning the multi-modal features, efficient collaborative processing of multi-modal data in procurement document verification is realized; by splicing the procurement semantic features and the image features by the semantic and visual collaborative feature extraction strategy to obtain the procurement vector, the limitation of single modal data in traditional review is broken, the semantic and visual correlation of the procurement document can be fully captured, and the accuracy and robustness of the review are improved; the procurement vector is analyzed and judged by the procurement review decision model, the procurement document verification result is given, the power procurement document to be verified is labeled, and the procurement document verification report is generated, which is convenient for subsequent tracing and rectification, and improves the transparency and operability of the review work.

[0008] In a first aspect, the application provides a procurement document verification method based on multi-modal and rule optimization, which specifically comprises the following steps: obtaining a power procurement document to be verified; extracting image features and non-image features in the power procurement document to be verified to form multi-modal features, and aligning the multi-modal features; performing semantic coding on the non-image features to obtain procurement semantic features; splicing the procurement semantic features and the image features by a semantic and visual collaborative feature extraction strategy to obtain a procurement vector; analyzing and judging the procurement vector based on a pre-constructed procurement review decision model combined with a differentiable reward function to give a procurement document verification result, wherein the differentiable reward function is obtained by compiling hard constraints of procurement review; labeling the power procurement document to be verified according to the procurement document verification result and generating a procurement document verification report.

[0009] Further, the image features are obtained by the following steps: obtaining an original procurement image matrix in the power procurement document to be verified; According to the visual attention allocation function, the procurement key area is extracted from the original procurement image matrix in combination with the visual learning parameter, and a procurement visual vector is obtained; According to the text semantic mapping function, in combination with the text learning parameter, a procurement image semantic vector related to the original procurement image matrix is extracted from the pre-constructed procurement corpus; Based on the vector conversion function, the procurement visual vector and the procurement image semantic vector are weighted and fused to obtain an image feature.

[0010] Further, the image feature is specifically represented as:

[0011] Wherein, V caption is a fixed-dimensional image feature used to describe the main content of the image in the power procurement to-be-verified file; F image is the original procurement image matrix containing pixel-level visual information; Ψ(·,Θ Ψ ) is a visual attention allocation function, Θ Ψ is a visual learning parameter used to filter procurement key areas from the original procurement image matrix; C text is a pre-constructed procurement corpus; Ω(·,Θ ω ) is a text semantic mapping function, Θ ω is a text learning parameter used to generate a procurement image semantic vector related to the original procurement image matrix from the procurement corpus; ⊙ is a weighted fusion operation of the procurement visual vector and the procurement semantic vector; Φ(·) is a vector conversion function that maps the fused vector to a fixed-dimensional image feature.

[0012] Further, through the semantic and visual collaborative feature extraction strategy, the procurement semantic feature and the image feature are spliced to obtain a procurement vector, which specifically includes: The procurement semantic feature and the image feature are respectively preprocessed to obtain a standard procurement semantic feature and a standard image feature; The similarity of each standard procurement semantic feature and standard image feature is analyzed to give a first procurement correlation degree; Based on the first procurement correlation degree, a procurement feature pair of the standard procurement semantic feature and the standard image feature is constructed, and the standard procurement semantic feature and the standard image feature in the procurement feature pair are cross-modal learned to adjust the standard procurement semantic feature and the standard image feature, and obtain a semantic learning feature and an image learning feature; The semantic learning feature and the image learning feature are processed for dimension consistency and spliced to obtain a procurement vector.

[0013] Further, the first procurement correlation degree is specifically represented as:

[0014] wherein, Sim(F text , F image ) is the similarity of the standard procurement semantic feature and the standard image feature, (F text , F image ) is the procurement feature pair, F text is the standard procurement semantic feature, F image is the standard image feature, ||F text || is the length of the standard procurement semantic feature, and ||F image || is the length of the standard image feature.

[0015] Further, based on the pre-constructed procurement review decision model, in combination with the differentiable reward function, the procurement vector is analyzed and judged to give the procurement file verification result, specifically including: deep semantic understanding and feature conversion are performed on the procurement vector to obtain a model procurement vector; based on the differentiable reward function, the degree of fit of each hard constraint in the procurement review and the model procurement vector is analyzed and fused to obtain a procurement reward value corresponding to the model procurement vector; different values of the procurement file verification result corresponding to the model procurement vector are searched, and in combination with the procurement reward value corresponding to the model procurement vector, selection probabilities and evaluation results corresponding to different values of the procurement file verification result are obtained.

[0016] Further, the differentiable reward function is obtained through the following steps: a matching degree function is constructed based on the matching degree of the hard constraint of the procurement review and the procurement vector; a constraint association function is constructed based on the semantic similarity of the hard constraint of the procurement review and the procurement vector; in combination with a preset balance coefficient, the contribution proportions of the matching degree function and the constraint association function are balanced and fused to give a fit function of each hard constraint; based on a preset constraint coefficient, the fit functions of each hard constraint are weighted and summed and mapped to obtain the differentiable reward function.

[0017] Further, the differentiable reward function is specifically represented as: wherein, Rdiff(R i , S) is the reward value of the differentiable reward function corresponding to the i th hard constraint R i and the procurement vector S, σ(·) is a Sigmoid activation function for mapping the result to the range of [0, 1], λ i is the constraint coefficient of the i th hard constraint, δ(R i , S) is the matching degree function, and φ(R iS) is a constraint correlation function, γ is a balance coefficient, γ ∈ (0, 1), and n is the total number of hard constraints.

[0018] Further, the procurement review decision model comprises a multi-layer self-attention mechanism, a feedforward neural network, and a multi-layer fully connected network comprising a procurement output network and a procurement evaluation network. The feedforward neural network is connected with the multi-layer self-attention mechanism and receives a semantic understanding vector output by the multi-layer self-attention mechanism for feature conversion. The multi-layer fully connected network is connected with the feedforward neural network and outputs different values of the procurement document verification result and a probability distribution corresponding to the different values. The procurement output network comprises a first input layer, a first hidden layer, and a first output layer. The first input layer receives a feature conversion vector output by the feedforward neural network and transmits the feature conversion vector to the first output layer through the first hidden layer to obtain different values of an intermediate review result and selection probabilities corresponding to the different values. The procurement evaluation network comprises a second input layer, a second hidden layer, and a second output layer. The second input layer receives the different values of the intermediate review result and the selection probabilities corresponding to the different values and transmits the different values and the selection probabilities to the second output layer through the second hidden layer to adjust the selection probabilities of the different values of the intermediate review result and feed back the adjusted selection probabilities to the first hidden layer.

[0019] Further, the network parameters of the multi-layer fully connected network are determined by the following steps: All parameters of the multi-layer self-attention mechanism and the feedforward neural network are set to a frozen state. The input end of the multi-layer fully connected network is connected with the feedforward neural network, and the output end of the multi-layer fully connected network outputs different values of the procurement document verification result and a probability distribution corresponding to the different values. The feature conversion vector output by the feedforward neural network is used as input data, and a historical review result label set is used as supervision data, wherein the historical review result label set comprises review labels for historical power procurement documents. The small-batch gradient descent method is used to adjust the network parameters of the multi-layer fully connected network, and the cross-entropy loss between the probability distribution output by the multi-layer fully connected network and the historical review result label set is calculated. Based on the cross-entropy loss, the network parameters of the multi-layer fully connected network are updated through back propagation, and the iteration is repeated until the cross-entropy loss converges.

[0020] Further, the first hidden layer comprises a shared encoder, and the first output layer comprises a main classification head, an auxiliary regression head, and an auxiliary binary classification head. The shared encoder captures core features related to the intermediate review result in the feature conversion vector input by the feedforward neural network. The main classification head is used to give the intermediate review result according to the core features output by the shared encoder. The auxiliary regression head is configured to analyze a deviation range of the central review result based on the core features, give a deviation result, and feed back to the second input layer. The auxiliary binary classification head is configured to analyze a bid risk of the central review result based on the core features, give a risk result, and feed back to the second input layer.

[0021] In a second aspect, the present application further provides a procurement document verification device based on multi-modal and rule optimization, which adopts the procurement document verification method based on multi-modal and rule optimization according to any one of the above, comprising: The data acquisition module is configured to acquire the power procurement document to be verified. The feature extraction module is configured to extract image features and non-image features in the power procurement document to be verified, form multi-modal features, and align the multi-modal features. The feature encoding module is configured to perform semantic encoding on the non-image features to obtain procurement semantic features. The feature fusion module is configured to splice the procurement semantic features and the image features by a semantic and visual collaborative feature extraction strategy to obtain a procurement vector. The feature analysis module is configured to analyze and judge the procurement vector based on a pre-constructed procurement review decision model in combination with a differentiable reward function, and give a procurement document verification result, wherein the differentiable reward function is obtained by compiling hard constraints of the procurement review. The report output module is configured to label the power procurement document to be verified according to the procurement document verification result, and generate a procurement document verification report.

[0022] The procurement document verification method and device based on multi-modal and rule optimization provided by the present application have at least the following beneficial effects: (1) The multi-modal features in the power procurement document to be verified are extracted and aligned, efficient collaborative processing of multi-modal data in the procurement document verification is realized, the procurement semantic features and the image features are spliced by the semantic and visual collaborative feature extraction strategy to obtain a procurement vector, the limitation of single modal data in traditional review is broken, the semantic and visual correlation of the procurement document can be fully captured, the accuracy and robustness of the review are improved, the procurement document verification result is given by analyzing and judging the procurement vector by using the procurement review decision model, the power procurement document to be verified is labeled, and the procurement document verification report is generated, which is convenient for subsequent tracing and rectification, and improves the transparency and operability of the review work.

[0023] (2) The hard constraints of the procurement review are compiled into a differentiable reward function, the procurement review decision model can dynamically adjust the strategy and select different values of the procurement document verification result by searching the strategy, the intelligentization of the review process is realized, and the review efficiency is improved.

[0024] (3) By encoding the non-image features semantically, the text information is converted into a feature vector containing semantics, the high-dimensional data is converted into a low-dimensional vector, which is convenient for processing and analysis; at the same time, the low-dimensional vector can reduce the consumption of computing resources, improve the processing and analysis efficiency, and through semantic encoding, data of different modalities (such as text, table) is uniformly represented as a vector, which is convenient for fusion processing. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 The flowchart of the procurement document verification method based on multi-modal and rule optimization provided in the embodiment of the present application is shown in the figure; Figure 2 The flowchart of determining the image features provided in the embodiment of the present application is shown in the figure; Figure 3 The flowchart of obtaining the procurement vector provided in the embodiment of the present application is shown in the figure; Figure 4 The flowchart of giving the procurement document verification result provided in the embodiment of the present application is shown in the figure; Figure 5 The flowchart of constructing the differentiable reward function provided in the embodiment of the present application is shown in the figure; Figure 6 The architecture diagram of the procurement audit decision model provided in the embodiment of the present application is shown in the figure; Figure 7 The structural block diagram of the procurement document verification device based on multi-modal and rule optimization provided in the embodiment of the present application is shown in the figure.

[0026] Among them, 201, data acquisition module; 202, feature extraction module; 203, feature encoding module; 204, feature fusion module; 205, feature analysis module; 206, report output module. DETAILED DESCRIPTION

[0027] In order to better understand the above technical solutions, the above technical solutions will be described in detail in conjunction with the drawings of the specification and specific embodiments. Obviously, the described embodiments are only part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor belong to the scope of protection of the present application.

[0028] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Multiple" generally includes at least two.

[0029] It is also to be understood that the terminology "include," "includes," or any other variation thereof is not limiting, such that a device or method that comprises a list of elements is not necessarily limited to only those elements but can include other elements not expressly listed or inherent to such device or method. An element proceeded by "comprises a... " does not, without more constraints, exclude the presence of additional identical elements in the device or method that includes the element.

[0030] The traditional procurement document verification technology has many defects. One is that the multi-modal data processing capability is insufficient. The traditional system can only process structured text or table data, and the analysis of electrical diagrams and wiring diagrams relies on manual annotation, resulting in incomplete and low-efficiency information extraction. Two is that the rule adaptation is rigid. The hard constraints are realized by a static rule engine, which cannot dynamically adjust the priority or handle implicit business logic, and frequent manual intervention is required for rule library updating. Three is that the model iteration is lagging. The existing audit model is trained based on a closed data set, which is difficult to integrate new rules or device types, resulting in poor adaptability to new business scenarios. In addition, the traditional system lacks interpretability, and the returned decision only provides simple results, which cannot locate specific field or graph element level errors, increasing the review cost.

[0031] Although the related art attempts to introduce a machine learning model, the problems of multi-modal data alignment difficulty and rule expression rigidity still restrict the audit intelligence level. The present application provides a procurement document verification method and device based on multi-modal and rule optimization. The method comprises: obtaining an electric power procurement document to be verified; extracting image features and non-image features in the electric power procurement document to be verified to form multi-modal features, and aligning the multi-modal features; performing semantic coding on the non-image features to obtain procurement semantic features; performing splicing on the procurement semantic features and the image features by a semantic and visual collaborative feature extraction strategy to obtain a procurement vector; analyzing and judging the procurement vector based on a pre-constructed procurement audit decision model combined with a differentiable reward function to give a procurement document verification result, wherein the differentiable reward function is obtained by compiling hard constraints of the procurement audit; and labeling the electric power procurement document to be verified according to the procurement document verification result and generating a procurement document verification report.

[0032] The multi-modal features in the power procurement to-be-verified file are extracted, and the multi-modal features are aligned, so that efficient collaborative processing of multi-modal data in procurement file verification is realized; through the semantic and visual collaborative feature extraction strategy, the procurement semantic features and image features are spliced to obtain the procurement vector, breaking the limitation of single modal data in traditional auditing, which can fully capture the semantic and visual correlation of the procurement file, and improve the accuracy and robustness of the audit; the procurement auditing decision model is used to analyze and judge the procurement vector, give the procurement file verification result, and label the power procurement to-be-verified file, generate the procurement file verification report, which is convenient for subsequent tracing and rectification, and improves the transparency and operability of the auditing work.

[0033] As shown in Figure 1 The embodiment of the application provides a procurement file verification method based on multi-modal and rule optimization, which specifically comprises the following steps: S101: obtaining a power procurement to-be-verified file.

[0034] Specifically, the power procurement to-be-verified file is a power equipment procurement plan file submitted by the procurement department, including a power equipment procurement plan table and related drawings of each power equipment, for example, an electrical wiring diagram, the plan table including device model, device quantity, device technical parameters and other information.

[0035] S102: extracting image features and non-image features in the power procurement to-be-verified file, forming multi-modal features, and aligning the multi-modal features.

[0036] Among them, the multi-modal features include image features and non-image features. Further, referring to Figure 2 The image features in the power procurement to-be-verified file are extracted through the following steps: Obtaining an original procurement image matrix in the power procurement to-be-verified file; According to the visual attention allocation function, combining the visual learning parameters, the procurement key area is extracted from the original procurement image matrix to obtain the procurement visual vector; According to the text semantic mapping function, combining the text learning parameters, the procurement image semantic vector related to the original procurement image matrix is extracted from the pre-constructed procurement corpus; Based on the vector conversion function, the procurement visual vector and the procurement image semantic vector are weighted and fused to obtain the image features.

[0037] In one specific implementation, key information from text, tables, and related drawings, including electrical wiring diagrams, is first extracted from the power procurement documents to be verified. This generates structured records, transforming the previously scattered, unstructured information into organized, structured data, paving the way for subsequent processing. Non-image features are obtained by extracting text and tables from the power procurement documents and performing structured processing. For image features, an image description generation algorithm is applied to the related drawings to generate vectors describing the key content of the corresponding drawings, i.e., image features. For example, an image description generation algorithm can be applied to an electrical wiring diagram to generate vectors describing the key content of the diagram, i.e., corresponding image features. Image features convert the complex image information in the related drawings into understandable vector form, resolving the issue of image information being difficult to directly incorporate into analysis and calculations. Finally, the text, tables, and image features are aligned to ensure accurate correlations between different types of data, avoid audit biases caused by data association errors, and ensure consistency of multimodal data in subsequent processing.

[0038] Furthermore, the image features are specifically expressed as:

[0039] Among them, V caption is a fixed-dimensional image feature used to describe the main content of the image in the power procurement verification document; F image is the original purchased image matrix, containing pixel-level visual information; Ψ(·,Θ Ψ ) is the visual attention allocation function, Θ Ψ is the visual learning parameter used to filter the key procurement areas from the original procurement image matrix; C text is a pre-built procurement corpus; Ω(·,Θ ω ) is the text semantic mapping function, Θ ω is a text learning parameter used to generate a purchase image semantic vector related to the original purchase image matrix from the purchase corpus; ⊙ is the weighted fusion operation of the purchase visual vector and the purchase semantic vector; Φ(·) is a vector conversion function that maps the fused vector to an image feature of fixed dimension.

[0040] In a specific example, the relevant figure is an electrical wiring diagram. Substitute the electrical wiring diagram into the calculation formula of the image description generation algorithm, that is, the electrical wiring diagram is the original procurement image matrix F image , substitute the electrical wiring diagram into the visual attention allocation function Ψ(·,Θ Ψ ), according to the adjustment of visual learning parameters, the key areas of procurement are extracted and the procurement visual vector is obtained. At the same time, the electrical wiring diagram is substituted into the text semantic mapping function Ω(·,Θ ωIn the text learning parameter adjustment, the purchase image semantic vector related to the original purchase image matrix is extracted from the pre-constructed purchase corpus. The visual learning parameter and the text learning parameter can be defined according to the demand or experience, or determined through model training, and the limitation is not made. Finally, the obtained purchase visual vector and purchase image semantic vector are substituted into the vector conversion function Φ(·), and weighted fusion is performed to obtain the image feature.

[0041] S103: The non-image features are semantically encoded to obtain the purchase semantic features.

[0042] In a specific embodiment, the non-image features are data cleaned to remove noise data, duplicate data and invalid data in the structured data. For example, deleting null values, correcting incorrect formats, etc. After completing the data cleaning, the structured data is converted into a unified format, and the data standardization is completed.

[0043] After standardization, for structured data containing text, it is processed by word segmentation, part-of-speech tagging, named entity recognition, etc. Then, the context information in the structured data is analyzed to understand the relationship between the data. For example, in a purchase plan table, the relationship between “product name” and “purchase quantity” is understood. Then the structured data is encoded, and the encoding method is selected according to the actual demand, for example, using word embedding for semantic encoding: mapping words or phrases in text data to high-dimensional vector space. For example, using Word2Vec or GloVe to map “transformer” to a vector. Sentence embedding can also be used for semantic encoding: mapping sentences to vectors. For example, using BERT or Sentence-BERT to map “The device must have overload protection function and can run continuously for 1 hour at 1.2 times rated current” to a vector. According to the demand, the encoded vector is spliced or converted to obtain the purchase semantic feature.

[0044] In a specific example, the power equipment purchase plan file includes the sentence “Need to purchase high-voltage switch cabinet with model ABB-1234, rated voltage 12kV, rated current 1250A.”. The sentence is decomposed into “need to purchase”, “model”, “ABB-1234”, “high-voltage switch cabinet”, “rated voltage”, “12kV”, “rated current”, “1250A” and other words or phrases. Extract key information such as “ABB-1234”, “12kV”, “1250A”. Use word embedding technology (such as Word2Vec or BERT) to convert each word or phrase into a vector and combine into a comprehensive vector.

[0045] For example, in the power equipment procurement plan document, there is a sentence "The equipment must have overload protection function and can run continuously for 1 hour at 1.2 times rated current." The sentence is decomposed into words or phrases such as "equipment", "must have", "overload protection function", "can run continuously for", "1.2 times rated current", "1 hour". Extract key information such as "overload protection function", "1.2 times rated current", "1 hour". Convert the entire sentence into a vector using sentence embedding technology (such as Sentence-BERT).

[0046] By semantically encoding non-image features, text information is converted into semantic feature vectors, high-dimensional data is converted into low-dimensional vectors, which facilitates processing and analysis. At the same time, low-dimensional vectors can reduce the consumption of computing resources and improve processing efficiency. And through semantic encoding, data of different modalities (such as text, table) are uniformly represented as vectors, which facilitates fusion processing.

[0047] S104: Obtain the procurement vector by splicing the procurement semantic features and image features through the semantic and visual collaborative feature extraction strategy.

[0048] The semantic and visual collaborative feature extraction strategy is a method for comprehensive processing of multi-modal data, which extracts more representative and semantically consistent features by fusing semantic information (such as text description, structured data, etc.) and visual information (such as images, videos, etc.), thereby providing more effective input for the downstream procurement vector review task.

[0049] The semantic and visual collaborative feature extraction strategy is a feature extraction method that cooperatively processes semantic features and visual features to mine the internal correlation and complementarity between them, thereby generating a fusion feature vector that can reflect both semantic information and visual information. Its core is to align semantic features and visual features in the same feature space through cross-modal learning, and to improve the representation ability and semantic consistency of features through collaborative optimization.

[0050] The semantic and visual collaborative feature extraction strategy can effectively mine the complementarity of semantic and visual information, providing stronger feature support for the procurement vector review task.

[0051] Further, referring to Figure 3 , the procurement semantic features and image features are spliced through the semantic and visual collaborative feature extraction strategy to obtain the procurement vector, which specifically includes: Preprocess the procurement semantic features and image features to obtain standard procurement semantic features and standard image features; Analyze the similarity of each standard procurement semantic feature and standard image feature to give a first procurement correlation degree; based on the first procurement correlation degree, a procurement feature pair of the standard procurement semantic feature and the standard image feature is constructed, and cross-modal learning is performed on the standard procurement semantic feature and the standard image feature in the procurement feature pair to adjust the standard procurement semantic feature and the standard image feature, so as to obtain semantic learning features and image learning features; The semantic learning features and the image learning features are subjected to dimension consistency processing and splicing to obtain a procurement vector.

[0052] Further, the first procurement correlation degree is specifically represented as:

[0053] , wherein Sim(F text , F image ) is the similarity of the standard procurement semantic feature and the standard image feature, (F text , F image ) is the procurement feature pair, F text is the standard procurement semantic feature, F image is the standard image feature, ||F text || is the length of the standard procurement semantic feature, and ||F image || is the length of the standard image feature.

[0054] In a specific embodiment, the procurement semantic features are preprocessed, including removing redundant information and performing standardization format operation to obtain the standard procurement semantic features. Meanwhile, the image features are preprocessed, including adjusting the size and performing normalization operation to obtain the standard image features. Then, the similarity between the standard procurement semantic features and the standard image features is calculated and analyzed to measure the semantic correlation degree and give the first procurement correlation degree. Based on the first procurement correlation degree, a procurement feature pair of the standard procurement semantic feature and the standard image feature is constructed, for example, taking the standard procurement semantic feature A as an example, the first procurement correlation degrees with each standard image feature are a1, a2, a3, …, an, the first procurement correlation degrees are sorted in descending order, and it is judged whether the maximum value in the first procurement correlation degrees reaches a correlation threshold value. If yes, the procurement feature pair is established, otherwise, the procurement feature pair is not established. Cross-modal learning is performed on the standard procurement semantic feature and the standard image feature in the procurement feature pair to adjust the standard procurement semantic feature and the standard image feature, so as to obtain semantic learning features and image learning features. Based on the construction result of the procurement feature pair, cross-modal contrast learning is performed to make the standard procurement semantic features and the standard image features related in semantics more close in the feature space. The semantic learning features and the image learning features after the cross-modal contrast learning are subjected to dimension consistency processing, and the processed semantic learning features and the image learning features are spliced to form a procurement vector.

[0055] By using the semantic and visual collaborative feature extraction strategy, the standard procurement semantic features and the standard image features are compared across modalities, so that the standard procurement semantic features and the standard image features related to semantics are closer in the feature space. This process enhances the relevance between text and image and enables a more comprehensive understanding of the overall content of the power equipment procurement plan document. Finally, the processed features are spliced into a procurement vector, which integrates multi-modal information and provides a comprehensive and comprehensive feature representation for the subsequent audit decision module, reducing the limitations of single modal information.

[0056] S105: Based on the pre-constructed procurement audit decision model, the procurement vector is analyzed and judged in combination with the differentiable reward function, and the procurement file verification result is given.

[0057] The differentiable reward function is obtained by compiling the hard constraints of procurement audit, and the hard constraints of procurement audit are the hard constraints of enterprise internal power equipment procurement including equipment specification standards and procurement budget restrictions.

[0058] Further, with reference to Figure 4 , based on the pre-constructed procurement audit decision model, the procurement vector is analyzed and judged in combination with the differentiable reward function, and the procurement file verification result is given, specifically including: Deep semantic understanding and feature conversion are performed on the procurement vector to obtain a model procurement vector; Based on the differentiable reward function, the degree of fit of each hard constraint in the procurement audit and the model procurement vector is analyzed and fused to obtain a procurement reward value corresponding to the model procurement vector; Search for different values of the procurement file verification result corresponding to the model procurement vector, combine the procurement reward value corresponding to the model procurement vector, and obtain the selection probability and evaluation result corresponding to the different values of the procurement file verification result.

[0059] Further, with reference to Figure 5 , the differentiable reward function is obtained by the following steps: Based on the matching degree of the hard constraints of procurement audit and the procurement vector, a matching function is constructed; Based on the semantic similarity of the hard constraints of procurement audit and the procurement vector, a constraint association function is constructed; Combine the preset balance coefficient to balance the contribution proportion of the matching function and the constraint association function and fuse to give the fit function of each hard constraint; Based on the preset constraint coefficient, the fit functions of each hard constraint are weighted and summed and mapped to obtain the differentiable reward function.

[0060] Further, the differentiable reward function is specifically represented as: Where Rdiff(Ri Rdiff(R i ,S) is the reward value of the differentiable reward function corresponding to the procurement vector S, σ(·) is a Sigmoid activation function for mapping the result to the range of [0, 1], λ i is the constraint coefficient of the i th hard constraint, δ(R i ,S) is a matching degree function, φ(R i ,S) is a constraint association function, γ is a balance coefficient, γ ∈ (0, 1), and n is the total number of hard constraints.

[0061] In a specific embodiment, the hard constraints of procurement review are compiled to obtain a differentiable reward function. Wherein, Rdiff(R i ,S) ∈ [0, 1], that is, the reward value is mapped to between 0 and 1, λ i is used to reflect the priority of the hard constraint, which is determined by statistical analysis of historical power procurement file review data. When S completely matches R i , the matching degree function δ(R i ,S) outputs 1, and outputs -1 when it does not match at all, and outputs a continuous value between -1 and 1 when it partially matches. By calculating the semantic similarity between S and R i , a continuous value between 0 and 1 is output, the balance coefficient is set according to the actual business requirements and rule characteristics, and is used to adjust the contribution proportion of the matching degree function and the constraint association function, and is used to adjust the contribution proportion of the rule matching degree and the association degree. σ(·) is a Sigmoid activation function for mapping the final reward value to between 0 and 1 to form a differentiable reward value, i is the serial number of the hard constraint of the procurement review, and n is the total number of hard constraints of the procurement review.

[0062] By constructing a differentiable reward function, the rigid hard constraint can be involved in the analysis and judgment of the procurement vector in a calculable and optimized form, so that the procurement file verification result is more in line with the actual business rules.

[0063] Further, referring to Figure 6 , the procurement review decision model includes a multi-layer self-attention mechanism, a feedforward neural network, and a multi-layer fully connected network including a procurement output network and a procurement evaluation network; The feedforward neural network is connected with the multi-layer self-attention mechanism and receives the semantic understanding vector output by the multi-layer self-attention mechanism for feature conversion; The multi-layer fully connected network is connected with the feedforward neural network, and outputs different values of the procurement file verification result and the probability distribution corresponding to the different values; The procurement output network includes a first input layer, a first hidden layer, and a first output layer. The first input layer receives the feature transformation vector output by the feedforward neural network and transmits the feature transformation vector to the first output layer through the first hidden layer to obtain different values of the intermediate audit result and selection probabilities corresponding to the different values. The procurement evaluation network includes a second input layer, a second hidden layer, and a second output layer. The second input layer receives the different values of the intermediate audit result and the selection probabilities corresponding to the different values and transmits the different values of the intermediate audit result and the selection probabilities corresponding to the different values to the second output layer through the second hidden layer to adjust the selection probabilities of the different values of the intermediate audit result and feed back the selection probabilities to the first hidden layer.

[0064] In a specific embodiment, the procurement output network is an Actor network, and the procurement evaluation network is a Critic network. The procurement output network is used to learn an optimal strategy, that is, to select an action a according to the feature transformation vector output by the feedforward neural network to maximize the cumulative reward of the differentiable reward function. The first input layer receives the feature transformation vector output by the feedforward neural network. The first hidden layer is composed of multiple fully connected layers, each of which is followed by a nonlinear activation function. The output of the first output layer is a probability distribution of the action, that is, the output layer is a Softmax layer that outputs the probability of each action. The feature transformation vector output by the feedforward neural network is input into the first input layer, and the output of the first hidden layer is calculated through forward propagation of the first hidden layer. The action is generated according to the output of the first output layer.

[0065] The procurement evaluation network is used to evaluate the quality of the action selected by the procurement evaluation network. The second input layer is based on a state-action value function and represents the features of the current state and action. The second hidden layer is composed of multiple fully connected layers, each of which is followed by a nonlinear activation function. The second output layer outputs the state-action value and feeds back the state-action value to the first hidden layer.

[0066] In a specific example, a provincial information center is preparing to purchase 120 AI servers with a budget of 6 million yuan. Based on the power procurement file to be verified, a procurement vector is obtained by encoding the procurement demand, technical indicators, budget, historical prices, policy provisions, and supplier qualifications into a 512-dimensional dense vector.

[0067] The procurement vector is input into a multi-layer self-attention mechanism including 6 layers of self-attention modules for semantic understanding. For example, the 1st layer associates "120 servers" and "budget of 6 million" to find that "unit price of 50,000 per unit" falls within the market price range of 45,000-52,000, and preliminarily marks it as "normal price". For another example, the 6th layer compares "supplier qualification" with "historical bid-winning record" to find that one supplier has won bids 4 times in 3 years, but each time just won the bid at the lowest price of 0.3%, thus generating a hidden state of "suspected bid rigging". Through processing of the procurement vector by the multi-layer self-attention mechanism, the semantic understanding results of each layer of self-attention modules are fused to output a 256-dimensional semantic understanding vector. It can be understood that the multi-layer self-attention mechanism can include multiple layers of self-attention modules, each layer of self-attention module analyzes different content, and the number of self-attention modules is configured according to actual needs, which is not limited.

[0068] Then the feedforward neural network performs feature conversion on the 256-dimensional semantic understanding vector output by the multi-layer self-attention mechanism to obtain a 128-dimensional feature conversion vector, which maps the "semantic level" features to "decision level" features, for example, "normal price" is converted to a continuous value of 0.12 (close to 0 indicates no risk). The feature conversion vector is input into the procurement output network, and after passing through the first input layer, the first hidden layer and the first output layer, an intermediate audit result and its probability are obtained, for example, the probability of "pass" is 15%, the probability of "return" is 25%, and the probability of "modification" is 60%. The procurement evaluation network analyzes and evaluates the intermediate audit result and its probability, splices the three one-hot vectors of "pass / return / modify" with their respective probabilities, and gives an expected risk through the second hidden layer and the second output layer, for example, the score corresponding to "pass" is -82 points, the score corresponding to "return" is -10 points, and the score corresponding to "modification" is -5 points. Based on this, a 3-dimensional "correction signal" c = [-0.20, +0.05, +0.15] is output to increase the probability of "modification" and reduce the probability of "pass". The correction signal c is fed back to the first hidden layer of the procurement output network to adjust the probability.

[0069] Based on the foregoing, it can be understood that the procurement audit decision model constructed by the present application is not a conventional single neural network model, but a "hybrid architecture with external evaluation-feedback loop".

[0070] Further, the first hidden layer includes a shared encoder, and the first output layer includes a main classification head, an auxiliary regression head and an auxiliary binary classification head; The shared encoder captures the core features related to the intermediate audit result in the input feature conversion vector of the feedforward neural network; The main classification head is used to give the intermediate audit result according to the core features output by the shared encoder; The auxiliary regression head is configured to analyze the deviation range of the central review result based on the core features, give a deviation result, and feed back to the second input layer. The auxiliary binary classification head is configured to analyze the risk of surrounding bidding of the central review result based on the core features, give a risk result, and feed back to the second input layer.

[0071] In the procurement output network, the shared encoder belongs to a hidden layer and is responsible for converting input features into an intermediate representation suitable for subsequent classification and regression tasks. The main classification head, the auxiliary regression head, and the auxiliary binary classification head are all modules that further process the output of the shared encoder. Their relationship with the shared encoder is as follows: The shared encoder converts the feature vector input by the feedforward neural network into a higher-level feature representation that can capture information related to classification and regression tasks in the input data.

[0072] The main classification head predicts the review result (pass, return, or modify) of the procurement plan based on the output of the shared encoder. The main classification head directly receives the output of the shared encoder and converts it into the output of the classification task.

[0073] The auxiliary regression head predicts the budget deviation degree based on the output of the shared encoder. The auxiliary regression head directly receives the output of the shared encoder and converts it into the output of the regression task.

[0074] The auxiliary binary classification head predicts whether there is a risk of surrounding bidding (a binary classification task) based on the output of the shared encoder. The auxiliary binary classification head includes a fully connected layer with an activation function that maps the output of the shared encoder to the probability of surrounding bidding risk. The auxiliary binary classification head directly receives the output of the shared encoder and converts it into the output of the binary classification task.

[0075] In a specific example, the procurement vector is a 256-dimensional vector representing various features of the procurement plan (such as budget, quantity, technical indicators, etc.). The input dimension of the shared encoder is 256, and the output is a 64-dimensional feature vector. The main classification head maps the 64-dimensional feature vector output by the shared encoder to output the probability distribution of the review result: 15% pass, 25% return, and 60% modify. The auxiliary regression head maps the 64-dimensional feature vector output by the shared encoder to output the budget deviation degree: the budget deviation degree is 0.05. The auxiliary binary classification head maps the 64-dimensional feature vector output by the shared encoder to output the probability of surrounding bidding risk: the surrounding bidding risk is 0.87.

[0076] The shared encoder is a hidden layer responsible for feature extraction. The main classification head, the auxiliary regression head, and the auxiliary binary classification head are all output layers responsible for different tasks but further process the output of the shared encoder.

[0077] Further, the network parameters of the multi-layer fully connected network are determined by the following steps: setting all parameters of the multi-layer self-attention mechanism and the feedforward neural network to a frozen state; defining different values of the input end of the multi-layer fully connected network connected with the feedforward neural network and the output end of the multi-layer fully connected network outputting the procurement document verification result and the probability distribution corresponding to the different values; converting the feature vector output by the feedforward neural network as input data, and taking a historical audit result label set as supervision data, wherein the historical audit result label set includes an audit label for historical power procurement document auditing; using a small batch gradient descent method to adjust the network parameters of the multi-layer fully connected network, and calculating the cross-entropy loss between the probability distribution output by the multi-layer fully connected network and the historical audit result label set; updating the network parameters of the multi-layer fully connected network based on the cross-entropy loss through back propagation, and repeating iteration until the cross-entropy loss converges.

[0078] At the same time, the results of manual auditing are extracted from the historical auditing database to generate a label set for model training, providing a historical experience reference for auditing decisions, so that the model can learn the logic and standards of manual auditing, and improve the rationality and accuracy of auditing.

[0079] In a specific embodiment, a procurement auditing decision model is used as the basis, the subject of the procurement auditing decision model adopts a Transformer architecture, performs deep semantic understanding and feature conversion on the input procurement vector, and can capture complex semantic relationships and potential information in the power procurement document to be verified. The output of the multi-layer fully connected network is the probability distribution of the three auditing actions of "pass", "return" and "modify". At the same time, the procurement output network and the procurement evaluation network are mounted, in this example, the procurement output network is an actor network, and the procurement evaluation network is a critic network, the actor network is used to output the selection probability of different values of the procurement document verification result, the critic network is used to evaluate the value of the current procurement document verification result value, and the different values of the procurement document verification result are selected by the ε-greedy strategy, balancing the exploration of new actions and the use of known effective actions. According to the differentiable reward function, the corresponding procurement reward value is obtained to obtain reward feedback, and the network parameters of the procurement auditing decision model are updated to output the selection probability and evaluation result corresponding to the different values of the final procurement document verification result, so that the procurement auditing decision model can optimize the decision, improve the accuracy and adaptability of the auditing.

[0080] The procurement review decision model collects feedback data in the daily review process, provides the latest business data support for procurement review decision model updating, periodically re-trains the procurement review decision model with appropriate learning rate using incremental learning technology, so that the model can learn new knowledge without forgetting old knowledge, and adapt to the dynamic changes of business; at the same time, through the generation of edge error samples by the generative adversarial network, the robustness test is carried out to test the performance of the procurement review decision model in extreme or special cases, to ensure that the procurement review decision model can evolve dynamically with new procurement rules and equipment types, and continuously improve the reliability and stability of the review, wherein the generator of the generative adversarial network adopts the U-Net structure, takes the historical error samples as input, and generates variation samples by random noise injection; the discriminator is a convolutional neural network, which outputs the distribution difference value of the generated samples and the real samples, and selects the samples with a difference value less than 10% for robustness test.

[0081] S106: According to the procurement document verification result, the power procurement to be verified document is labeled, and a procurement document verification report is generated.

[0082] It can be understood that when the review result is passed, the result is clear and no further explanation of the review process is needed, but when the procurement document verification result is returned or modified, the field and graphic element level highlight explanation is generated for the returned or modified power procurement to be verified document, which points out the part that does not meet the hard constraint or review rule, so that the procurement department can intuitively understand the problem and reduce the communication cost. At the same time, these explanations are written into the review report to form a standard and formal feedback document, which is convenient for subsequent tracing and rectification, and improves the transparency and operability of the review work.

[0083] In one specific example, there is an audit task of an audit plan, and the headquarters of a convenience store formulates a beverage procurement plan for three new stores A, B and C in July, which needs to be audited.

[0084] First, the purchase to be verified file is obtained. For example, the demand table of each store, including the store area, the surrounding customer portrait, the historical sales, the promotion period, the demand of various drinks, etc. Then the image features and non-image features are extracted and aligned. Among them, the non-image features are obtained by structuring the various information in the demand table. The non-image features are semantically encoded to obtain the purchase semantic features. For example, "customer portrait = young white-collar" is converted into a one-hot vector, and "historical sales 1200 pieces" is standardized by Z-Score, that is, the numerical field is standardized, the category field is encoded, and the missing value is filled in. Finally, 30-dimensional standard purchase semantic features are obtained. The image features can be the street view and the real photo of each store. There are cars, trees, and pedestrians in the street view. First, use Mask-RCNN to remove irrelevant pixels, then adjust the brightness to 0.5 average, image denoising, size uniformity, and color normalization. Finally, 2048-dimensional standard image features are obtained.

[0085] By the semantic and visual collaborative feature extraction strategy, the purchase semantic features and the image features are spliced to obtain the purchase vector. The first purchase correlation degree of the purchase semantic features and the image features is analyzed. For example, it is found that "store A semantic says 'dense office buildings'" and "store A street view appears a large number of coffee cup logos" have a cosine similarity of 0.87 in the feature space, that is, the similarity matrix of the semantic features vs. the image features is calculated by the cosine distance, and the first purchase correlation degree (numerical value 0~1) is obtained. "Store A semantic" and "store A image" are bound into a pair; low correlation (such as store C semantic and store A image) is discarded. Based on the first correlation degree threshold (>0.7), the purchase feature pairs are constructed. Use the cross-modal Transformer to make the semantic features and the image features notice each other: the "customer" dimension in the semantic features is activated by the "school uniform" in the image, and the "coffee logo" in the image features is activated by the "office building" in the semantic. Output 30-dimensional semantic learning features + 2048-dimensional image learning features, which are aligned to the same latent space. The semantic learning features are 30-dimensional, the image learning features are 2048-dimensional, the image is reduced to 64-dimensional by 1x1 convolution, and the semantic is increased to 64-dimensional, that is, the linear projection ensures the consistency of the dimensions. The 64-dimensional semantic + 64-dimensional image is directly concatenated into a 128-dimensional vector to obtain a 128-dimensional purchase vector.

[0086] Based on the pre-constructed procurement review decision model, in combination with the differentiable reward function, the procurement vector is analyzed and judged, the procurement file verification result is given, and the power procurement to-be-verified file is labeled, and the procurement file verification report is generated. For example, input the 128-dimensional vector into the procurement review decision model, and output "pass" 75%, that is, the probability of passing the to-be-verified file is 75%, indicating that the procurement plan in the to-be-verified file passes the review, and the procurement can be carried out accordingly. For another example, input the 128-dimensional vector into the procurement review decision model, and output "return" 70%, that is, the probability of returning the to-be-verified file is 70%, and the to-be-verified file needs to be modified, and the data in the procurement plan is judged in turn, for example, the number of A drinks in the procurement plan is 22.5 boxes, the number of B drinks is 5000 boxes, and the number of C drinks is 200 boxes, wherein the number of A drinks cannot be a decimal number, the procurement plan of A drinks is labeled, the number of B drinks exceeds the procurement threshold, and the procurement plan of B drinks is highlighted, and the number of C drinks meets the requirements and does not need to be labeled. The procurement file verification report is generated in combination with the labeled procurement plan.

[0087] In summary, based on the procurement file verification method based on multi-modal and rule optimization, in the power equipment procurement scene, the structured and aligned multi-modal data is realized through cross-modal learning and feature splicing, the procurement vector is generated, the differentiable reward function is constructed, the hard constraint of procurement review is integrated, the procurement file verification result is output by the procurement review decision model, and the procurement file verification report is generated based on the labeling of the power procurement to-be-verified file. The auditing efficiency, accuracy and dynamic adaptability are greatly improved, and the auditing demand of complex procurement plan in the power industry is met.

[0088] Reference Figure 7 The embodiment of the application provides a procurement file verification device based on multi-modal and rule optimization, comprising: The data acquisition module 201 is used for acquiring the power procurement to-be-verified file; The feature extraction module 202 is used for extracting the image features and non-image features in the power procurement to-be-verified file, forming multi-modal features, and aligning the multi-modal features; The feature encoding module 203 is used for performing semantic encoding on the non-image features to obtain procurement semantic features; The feature fusion module 204 is used for splicing the procurement semantic features and the image features to obtain a procurement vector through a semantic and visual collaborative feature extraction strategy; The feature analysis module 205 is used for analyzing and judging the procurement vector based on the pre-constructed procurement review decision model in combination with the differentiable reward function, and giving the procurement file verification result, wherein the differentiable reward function is obtained by compiling the hard constraint of the procurement review; The report output module 206 is configured to mark the power purchase file to be verified according to the verification result of the purchase file, and generate a purchase file verification report.

[0089] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described modules can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0090] Although preferred embodiments of the present application have been described, those skilled in the art, once they know the basic creative concept, can make additional changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application. Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application also intends to include these modifications and variations.

Claims

1. A procurement document verification method based on multimodality and rule optimization, characterized in that: include: Obtain documents to be verified for power procurement; Extract image features and non-image features from power procurement verification documents to form multimodal features, and align the multimodal features; Semantically encode non-image features to obtain procurement semantic features; Through the semantic and visual collaborative feature extraction strategy, the procurement semantic features and image features are spliced ​​together to obtain the procurement vector; Based on a pre-built procurement review decision model and a differentiable reward function, the procurement vector is analyzed and judged, and the procurement document verification result is given. The differentiable reward function is obtained by compiling the hard constraints of the procurement review. Based on the procurement document verification results, the power procurement documents to be verified are marked and a procurement document verification report is generated.

2. The procurement document verification method based on multimodality and rule optimization according to claim 1 is characterized in that: Extract the image features of the power procurement verification file through the following steps: Obtain the original procurement image matrix in the power procurement file to be verified; According to the visual attention allocation function and combined with the visual learning parameters, the procurement key area is extracted from the original procurement image matrix to obtain the procurement visual vector; According to the text semantic mapping function and combined with the text learning parameters, the procurement image semantic vector related to the original procurement image matrix is ​​extracted from the pre-built procurement corpus; Based on the vector conversion function, the procurement visual vector and the procurement image semantic vector are weightedly fused to obtain the image features.

3. The procurement document verification method based on multimodality and rule optimization according to claim 2 is characterized in that: Image features are specifically expressed as: ; Among them, V caption is the image feature of fixed dimension; F image is the original purchased image matrix; Ψ(·,Θ Ψ ) is the visual attention allocation function, Θ Ψ is the visual learning parameter; C text is a pre-built procurement corpus; Ω(·,Θ ω ) is the text semantic mapping function, Θ ω is the text learning parameter; ⊙ is the weighted fusion operation of the purchase visual vector and the purchase semantic vector; Φ(·) is the vector conversion function.

4. The procurement document verification method based on multimodality and rule optimization according to claim 1 is characterized in that: Through the semantic and visual collaborative feature extraction strategy, the procurement semantic features and image features are spliced ​​together to obtain the procurement vector, which specifically includes: Preprocessing the procurement semantic features and image features respectively to obtain standard procurement semantic features and standard image features; Analyze the similarity between each standard procurement semantic feature and the standard image feature, and give the first procurement relevance; Based on the first procurement relevance, constructing procurement feature pairs of standard procurement semantic features and standard image features, performing cross-modal learning on the standard procurement semantic features and standard image features in the procurement feature pairs, adjusting the standard procurement semantic features and standard image features, and obtaining semantic learning features and image learning features; The semantic learning features and image learning features are dimensionally consistent and concatenated to obtain the procurement vector.

5. The procurement document verification method based on multimodality and rule optimization according to claim 1 is characterized in that: Based on the pre-built procurement review decision model and the differentiable reward function, the procurement vector is analyzed and judged, and the procurement document verification results are given, including: Perform deep semantic understanding and feature conversion on the purchase vector to obtain the model purchase vector; Based on a differentiable reward function, the degree of fit between the various hard constraints in the procurement review and the model procurement vector is analyzed and integrated to obtain the procurement reward value corresponding to the model procurement vector; The search model procurement vector corresponds to different values ​​of the procurement document verification results, and combined with the procurement reward value corresponding to the model procurement vector, the selection probability and evaluation results corresponding to different values ​​of the procurement document verification results are obtained.

6. The procurement document verification method based on multimodality and rule optimization according to claim 5 is characterized in that: The differentiable reward function is obtained by the following steps: Based on the matching degree between the hard constraints of procurement review and the procurement vector, a matching function is constructed; Based on the semantic similarity between the hard constraints of procurement review and procurement vectors, a constraint association function is constructed; Combined with the preset balance coefficient, the contribution ratio of the balance matching function and the constraint association function is integrated and given the matching function of each hard constraint; Based on the preset constraint coefficients, the compliance functions of each hard constraint are weighted summed and mapped to obtain a differentiable reward function.

7. The procurement document verification method based on multimodality and rule optimization according to claim 5 is characterized in that: The procurement review decision model includes a multi-layer self-attention mechanism, a feedforward neural network, and a multi-layer fully connected network including a procurement output network and a procurement evaluation network; The feedforward neural network is connected to the multi-layer self-attention mechanism, and receives the semantic understanding vector output by the multi-layer self-attention mechanism for feature conversion; The multi-layer fully connected network is connected to the feedforward neural network to output different values ​​of the procurement document verification results and the probability distribution corresponding to different values; The procurement output network includes a first input layer, a first hidden layer, and a first output layer. The first input layer receives the feature conversion vector output by the feedforward neural network and propagates it to the first output layer through the first hidden layer to obtain different values ​​of the intermediate audit results and the corresponding selection probabilities of different values. The procurement evaluation network includes a second input layer, a second hidden layer and a second output layer. The second input layer receives different values ​​of the intermediate audit results and the selection probabilities corresponding to different values, and propagates them to the second output layer through the second hidden layer, and adjusts the selection probabilities of different values ​​of the intermediate audit results to feed back to the first hidden layer.

8. The procurement document verification method based on multimodality and rule optimization according to claim 7 is characterized in that: The network parameters of the multi-layer fully connected network are determined by the following steps: Set all parameters of the multi-layer self-attention mechanism and the feed-forward neural network to a frozen state; Define the connection between the input end of the multi-layer fully connected network and the feedforward neural network, and the output end of the multi-layer fully connected network outputs different values ​​of the procurement document verification result and the probability distribution corresponding to the different values; The feature transformation vector output by the feedforward neural network is used as input data, and the historical audit result label set is used as supervision data, wherein the historical audit result label set includes the audit labels of the historical power procurement document audits; The mini-batch gradient descent method is used to adjust the network parameters of the multi-layer fully connected network and calculate the cross entropy loss between the probability distribution of the multi-layer fully connected network output and the historical audit result label set; The network parameters of the multi-layer fully connected network are updated by back propagation based on the cross entropy loss, and the iterations are repeated until the cross entropy loss converges.

9. The procurement document verification method based on multimodality and rule optimization according to claim 7, characterized in that: The first hidden layer includes a shared encoder, and the first output layer includes a main classification head, an auxiliary regression head, and an auxiliary binary classification head; The shared encoder captures the core features related to the intermediate audit results in the input feature transformation vector of the feedforward neural network; The main classification head is used to provide intermediate audit results based on the core features output by the shared encoder; The auxiliary regression head is used to analyze the deviation range of the central audit results based on the core features, provide the deviation results, and feed them back to the second input layer; The auxiliary binary classification head is used to analyze the bid-rigging risk of the central audit results based on core features, provide risk results, and feed back to the second input layer.

10. A procurement document verification device based on multimodality and rule optimization, characterized in that: The procurement document verification method based on multimodality and rule optimization according to any one of claims 1 to 9 is adopted, comprising: A data acquisition module, used to obtain the power procurement files to be verified; A feature extraction module is used to extract image features and non-image features from the power procurement verification file to form multimodal features and align the multimodal features; Feature encoding module, used to semantically encode non-image features to obtain procurement semantic features; The feature fusion module is used to combine procurement semantic features and image features through a semantic and visual collaborative feature extraction strategy to obtain a procurement vector; The feature analysis module analyzes and judges the procurement vector based on a pre-built procurement review decision model and a differentiable reward function, generating procurement document verification results. The differentiable reward function is compiled from hard constraints on procurement review. The report output module is used to mark the power procurement documents to be verified based on the procurement document verification results and generate a procurement document verification report.

Citation Information

Patent Citations

  • Material purchasing plan reviewing method and device

    CN112446649A

  • Purchase demand intelligent distribution method and device, computer equipment and storage medium

    CN115187107A

  • Intelligent purchasing method and device

    CN117114813A

  • Bidding and tendering data processing method and system for tendering and tendering review

    CN118863823A

  • Multi-modal document information processing method, device and equipment based on large model agent and storage medium

    CN119623650A

Cited By

  • Series fault arc detection method for new energy automobile power battery based on neural network

    CN121878502A