A multi-modal supervised learning-based inspection auditing large model and an intelligent work order auditing method and system

The large-scale audit model developed through multimodal supervised learning solves the problems of small model size and weak generalization ability in power marketing audits, and achieves efficient and accurate anomaly detection of audit work orders, thereby improving the level of intelligence in audit work.

CN120806883BActive Publication Date: 2026-01-02STATE GRID JIANGSU ELECTRIC POWER CO LTD MARKETING SERVICE CENT +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511307938.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-01-02
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

Existing technologies for electricity marketing audits suffer from problems such as small model size, overfitting, inability to fully utilize the complexity of multimodal data, weak generalization ability, and poor robustness, resulting in low efficiency and insufficient accuracy in audit work.

Method used

We adopt a large-scale inspection model based on multimodal supervised learning. We extract the sequence features of the inspection domain instruction dataset through a multi-layer Transformer architecture, introduce domain expert knowledge, and combine improved adaptive low-rank adaptation technology and PPO algorithm to fine-tune the model. We also provide API interface services for anomaly detection.

Benefits of technology

It achieves efficient and accurate anomaly detection in audit work orders, improves the model's generalization ability and robustness, supports lightweight deployment of large models, and enhances the intelligence and accuracy of audit work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806883B_ABST
    Figure CN120806883B_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal supervised learning-based inspection auditing large model and an intelligent work order auditing method, characterized in that: collected electric power marketing inspection data is subjected to data cleaning and data labeling to form a marketing inspection field instruction data set; a multi-modal coding is used to extract sequence features of the marketing inspection field instruction data set, and the extracted sequence features are input into a base model with a multi-layer Transformer model as an architecture for deep feature extraction; a marketing inspection large model is constructed through a retrieval enhancement generation method; the constructed marketing inspection large model is subjected to pre-training and instruction fine-tuning, different types of data are subjected to unified modeling, understanding and generation tasks, and a PPO algorithm based on a heterogeneous strategy optimization is used to adjust instruction fine-tuning model parameters; based on an improved adaptive low-rank adaptation technology, model part layer weight matrices are approximated as products of low-rank matrices, and API interface services are provided to perform abnormality detection on input marketing inspection industry work order data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and more particularly, to a marketing inspection large model based on multi-modal and supervised learning algorithm and an intelligent work order auditing method and system. BACKGROUND

[0002] With the advancement of digital transformation, as the all-around supervision and control of marketing business, the inspection work urgently needs to rely on informationization and intelligent means to promote its high-quality development and adapt to the needs of the new era.

[0003] In recent years, with the rapid development of power marketing business, the types of business involved and the amount of data have increased dramatically, and the policy specifications for inspection are also evolving. Inspection personnel not only need to be familiar with the latest inspection specifications, but also need to master the constantly increasing, iterative marketing business policies, industry standards and regulations. The traditional marketing inspection work mainly carries out inspection work through online distribution of inspection work orders. Due to the large amount of work orders and extensive tasks, manual auditing has problems such as low efficiency and insufficient accuracy. In addition, individual differences, lack of experience and subjective judgment of inspection personnel may cause deviations in accuracy and fairness during the auditing process. In order to solve the above problems, improve work efficiency and reduce potential errors, it is urgent to introduce informationization and intelligent means to promote the efficient and accurate development of power marketing inspection.

[0004] In recent years, machine learning technology has made significant progress in marketing inspection. For example, existing technology proposes the necessity of applying power big data technology to power marketing audit, especially for anomaly detection in power data, which can effectively improve the quality of power marketing audit. Existing technology suggests using big data algorithms to analyze customer illegal electricity use behavior in power marketing inspection. In addition, existing technology designs an intelligent power marketing audit model based on knowledge graph, which solves the limitations of traditional audit models. Existing technology focuses on marketing audit label library technology based on multi-data fusion, emphasizing the importance of data integration and mining for efficient audit work. Although existing research has made some progress in the intelligent auditing technology of power marketing inspection work orders, there are still some challenges. The current model generally has the problem of small size and overfitting, and cannot fully utilize the complexity of multi-modal data. Small models can only handle local information and are difficult to cope with complex and variable power marketing environments, and their generalization ability is weak, which can easily lead to performance degradation when dealing with heterogeneous data. In addition, the existing small model has poor robustness in data noise and deformation, which limits its universality and application effect in actual scenarios.

[0005] With the rapid development of the power industry, marketing inspection work has gradually accumulated massive data containing rich information. Massive information promotes the scale of artificial intelligence technology model from quantitative change to qualitative change, and promotes the transformation of artificial intelligence technology in the power field from small model technology to large model. Large language model (LLM) refers to a neural network model with more than 100 million parameters based on the core architecture of Transformer, which is trained on massive text data in an unsupervised manner. It has strong multi-modal learning ability and 100 million data fitting ability in processing natural language tasks. Large models have significant advantages in natural language processing, especially in zero-shot learning, few-shot learning, and noisy data processing. They can effectively improve the generalization ability of models and solve the limitations of small models in handling complex data. Large models have shown great potential in multiple fields and cross-scenario applications. For example, the prior art proposes a cloud ERP community domain problem classification method based on BERT-TextCNN, which inputs cloud community problem text vectors into a pre-trained model to extract deep features. However, this method lacks pre-training of power professional knowledge, and has problems such as insufficient model generalization and poor domain adaptability. The prior art developed the first financial large model in the industry, which can provide customized financial professional services for different enterprises and promote the development of financial intelligence. The prior art developed a customized Chinese legal large model LawLLM, which provides intelligent services in deep application scenarios such as legal information extraction, judgment prediction, and super-long judgment processing for super-long text processing in the legal field. In the power grid field, small model artificial intelligence represented by deep learning and reinforcement learning has been widely applied, but large model technology is still in its infancy and needs to be developed and popularized.

[0006] To solve the above problems, there is an urgent need for a marketing inspection large model based on multi-modal and supervised learning algorithm and an intelligent work order auditing method and system. SUMMARY

[0007] To solve the problems in the prior art, the present application provides a marketing inspection large model based on multi-modal and supervised learning algorithm and an intelligent work order auditing method and system.

[0008] The present application adopts the following technical solutions.

[0009] The first aspect of the application relates to a multi-modal supervised learning-based inspection audit large model and an intelligent work order audit method, the method comprising the following steps: collecting power marketing inspection data by a marketing business support system, performing data cleaning and data labeling on the collected power marketing inspection data, constructing an inspection field instruction data set, and inputting a model pre-training and fine-tuning module; fully extracting sequence features of the inspection field instruction data set using multi-modal encoding, inputting the extracted sequence features into a base model with a multi-layer Transformer model as the architecture for deep feature extraction, introducing domain expert knowledge through a retrieval enhancement generation method to construct a marketing inspection large model; pre-training and instruction fine-tuning the constructed marketing inspection large model using the inspection field instruction data set, performing unified modeling, understanding and generation tasks on text, time sequence and other data types in the power marketing inspection scene, and adjusting the instruction fine-tuning model parameters using a PPO algorithm based on heterogeneous strategy optimization; based on the improved adaptive low-rank adaptation technology, the model part layer weight matrix is approximated as the product of a low-rank matrix, and an API interface service is provided to detect abnormalities in the input marketing inspection industry work order data.

[0010] The marketing business support system collects power marketing inspection data, performs data cleaning and data labeling on the collected power marketing inspection data, and constructs an inspection field instruction data set, which is input into a model pre-training and fine-tuning module, including: obtaining marketing inspection industry global data measured by power detection equipment, the marketing inspection industry global data being multi-source multi-modal information including inspection work orders, regulations and systems, policy reports, and daily work order quantities; preprocessing the collected global data, and obtaining power marketing inspection instruction fine-tuning data set using data labeling.

[0011] Fully extracting sequence features of the inspection field instruction data set using multi-modal encoding, including: extracting key features, aligning features and reconstructing features from multi-modal data including marketing inspection industry work order data, regulations and systems, policy reports, and daily work order quantities through a multi-modal encoding-decoding structure.

[0012] Key feature extraction, feature alignment and feature reconstruction are performed on multi-modal data including marketing inspection industry work order data, regulations, policy reports, and daily work order quantity through a multi-modal encoding-decoding structure, including: using a tokenizer of a large model to perform word segmentation processing on original text of unstructured text data including inspection work orders and regulations, constructing a numerical vector representation of each word according to a word table, and extracting text features through a text encoder; structured time series data including daily work order quantity and monthly work order quantity are divided into multiple patches, time series features are converted into text features through a multi-head attention mechanism, and a linear mapping layer is used to output time series representation with the same dimension as the text vector; different modal data in the unified feature space is mapped to the text embedding space of the large model through a multi-modal encoder and a fully connected layer.

[0013] The extracted sequence features are input into a base model with a multi-layer Transformer model architecture for deep feature extraction, and domain expert knowledge is introduced through a retrieval enhancement generation method to build a marketing inspection large model, including: using a multi-layer Transformer architecture to perform deep non-linear operations on multi-modal features to achieve deep feature representation of marketing inspection global data; a multi-modal decoder is constructed to reconstruct text and time series data of the marketing inspection industry, and multi-modal output is obtained; a retrieval enhancement generation technique is used to externally expand the knowledge base of the marketing inspection large model, and without updating the parameters of the large model, the prompt words matched with the problem are automatically expanded.

[0014] The marketing inspection large model is pre-trained and fine-tuned using inspection field instruction data sets to perform unified modeling, understanding and generation tasks on text, time series and other data types in the power marketing inspection scenario, including: using the base model to perform deep feature representation on unified multi-modal features, and reconstructing the original input through the trained text decoder and time series decoder.

[0015] Based on the improved adaptive low-rank adaptation technology, the weight matrix of part of the layers of the model is approximated as the product of low-rank matrices, including: based on the improved adaptive low-rank adaptation technology based on singular value decomposition, the parameters of the rest of the large model are frozen, and only the bias parameters or linear layer parameters of the Transformer network are updated.

[0016] Based on the improved adaptive low-rank adaptation technology, the weight matrix of part of the layers of the model is approximated as the product of low-rank matrices, and an API interface service is provided to perform anomaly detection on input marketing inspection industry work order data, including: calling the marketing inspection large model API to generate inspection work order data, using the power marketing inspection work order data to test the model, and performing intelligent research and analysis on abnormal work orders.

[0017] The second aspect of the application relates to a multi-modal supervised learning-based inspection and review large model and an intelligent work order review system using the method in the first aspect of the application, the system comprising a multi-modal instruction data set construction module, an inspection and review large model research and development module, a model pre-training and fine-tuning module, and a lightweight deployment and application module; the multi-modal instruction data set construction module collects power marketing inspection data from a marketing business support system, performs data cleaning and data labeling on the collected power marketing inspection data, and forms an inspection field instruction data set, which is input into the model pre-training and fine-tuning module; the inspection and review large model research and development module fully extracts sequence features of the inspection field instruction data set using multi-modal coding, inputs the extracted sequence features into a base model with a multi-layer Transformer model as the architecture for deep feature extraction, introduces field expert knowledge through a retrieval enhancement generation method, and constructs a marketing inspection large model; the model pre-training and fine-tuning module pre-trains and fine-tunes the constructed marketing inspection large model using the inspection field instruction data set, uniformly models, understands and generates tasks for text, time sequence and other data types in the power marketing inspection scene, and adjusts instruction fine-tuning model parameters using a PPO algorithm based on heterogeneous strategy optimization; the lightweight deployment and application module approximates part of the layer weight matrix of the model to the product of a low-rank matrix based on an improved adaptive low-rank adaptation technology, provides API interface services, and performs abnormality detection on input marketing inspection industry work order data.

[0018] The third aspect of the application relates to a terminal comprising a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to perform the steps of the method in the first aspect of the application.

[0019] The fourth aspect of the application relates to a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method in the first aspect of the application.

[0020] The beneficial effects of the present application are that, compared with the prior art, the marketing inspection large model based on multi-modal and supervised learning algorithm and intelligent work order auditing method and system in the present application fuse the global data of the power marketing inspection industry through multi-modal data, and perform data labeling to make instruction fine-tuning data set, provide large-scale professional field data support for model pre-training and fine-tuning, adopt multi-modal encoder to extract key features of different modal data, perform feature alignment to convert into unified sequence form, further construct a base model with multi-layer Transformer architecture as the core, introduce expert knowledge into the model performance through the retrieval augmented generation (RAG) module, use the pre-training-fine-tuning paradigm of the large model, adopt autoregressive prediction optimization target on the multi-modal feature sequence, combine supervised instruction fine-tuning and reinforcement learning algorithm to enhance the ability of the model to understand and follow human instructions in specific scenarios, align the output to human values and preferences through human feedback, improve the quality and reliability of the model answer, compress the parameter update amount in the model inference process through the low-rank adaptation (LoRA) technology based on singular value decomposition improvement, realize the lightweight deployment of the large model, and provide API interface service to realize the calling of the inspection auditing large model and intelligent work order auditing.

[0021] The beneficial effects of the present application also include:

[0022] 1. The inspection auditing large model based on multi-modal supervised learning and the intelligent work order auditing method provided by the present application solve the technical problems of marketing inspection work order abnormality checking in the power industry, realize effective detection and identification of abnormal work orders, have high operation efficiency and detection precision, can fully capture work order text features, and have other advantages.

[0023] 2. The base model of the Transformer architecture based on the attention mechanism is pre-trained and fine-tuned on the global data set in the field of power marketing inspection to realize the research and development of the marketing inspection large model, solve the problems of too large difference between general data distribution and multi-modal data in the inspection field and low professional degree, and at the same time, introduce professional power knowledge through the retrieval augmented generation, use the perplexity score of the large language model as a supervision signal to fine-tune the retriever parameters, solve the dynamic adaptation problem of the large model in the marketing inspection field, improve the inference ability of the large model in complex downstream tasks, and enhance the generalization of the model.

[0024] 3. Through multi-modal encoding-decoding operation on unstructured text such as inspection work order, policy report and structured time series data set such as daily work order quantity, the multi-head attention mechanism is used to convert the time series characteristics into text characteristics, enhance the cross-modal and reasoning ability of the model on time series data, and make work order data multi-dimensional evolution trend prediction in time domain scale to meet the needs of marketing policy changes and business sustainable development, help continuous optimization of inspection business process, improve the intelligent level of marketing inspection, and develop more accurate and effective inspection strategies.

[0025] 4. By introducing a low-rank adaptation method based on singular value decomposition, the problem of model performance and computational efficiency decline caused by fixed rank of parameter matrix of each layer of LoRA is solved. In addition, the redundant LoRA rank will cause the degradation of model performance and efficiency, and the importance of weights will differ in different layers of the Transformer model during fine-tuning. The parameter update matrix is singular value decomposed, and the rank of each layer is dynamically adjusted through singular value pruning during training, so that the expression ability of different layers is flexibly optimized according to the actual task requirements. This method effectively improves the parameter efficiency and task adaptability of model fine-tuning, reduces the model storage occupation and computational overhead without sacrificing the overall performance of the model. BRIEF DESCRIPTION OF DRAWINGS

[0026] Figure 1 A schematic diagram of the inspection and audit large model based on multi-modal supervised learning and the intelligent work order audit method of the present application. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical scheme and advantages of the present application clearer and more accurate, the technical scheme of the present application is described in detail below through multiple specific embodiments. The embodiments used by the present application are only used to explain the present application and are not used to limit the content of the present application.

[0028] In the first aspect of the present application, an inspection and audit large model based on multi-modal supervised learning and an intelligent work order audit method are provided, and the method comprises the following steps:

[0029] Step 1: Collecting power marketing inspection data from the marketing business support system, performing data cleaning and data labeling on the collected power marketing inspection data, constructing an inspection field instruction data set, and inputting the model pre-training and fine-tuning module.

[0030] The marketing inspection large model based on multi-modal supervised learning and the intelligent work order audit method need to collect data information related to industry work orders, mainly including the following data: rules and regulations, inspection work order, policy report, time series data. For example Figure 1The multi-modal supervised learning-based marketing inspection large model and the intelligent work order auditing method shown comprise industry inspection global data acquisition, data preprocessing, model development, model expansion, model pre-training, model fine-tuning, model deployment, and abnormal work order identification and analysis.

[0031] Industry inspection global data acquisition needs to pull non-structured text data related to the industry, including rules and regulations, inspection work orders, policy reports, and structured time series data such as daily work order volume; data preprocessing includes low-quality filtering, redundancy removal, privacy elimination, and data labeling.

[0032] First, collect industry-wide data for marketing inspection, including inspection work orders, rules and regulations, policy reports, and daily work order volume, then perform a series of preprocessing steps on the collected unstructured text information, including low-quality filtering, redundancy removal, and privacy elimination, and finally provide a scenario-rich inspection field instruction dataset for downstream tasks through data labeling. Low-quality filtering refers to deleting low-quality content from the collected data. By using a linear classifier based on feature hashing to evaluate the text, the classifier is trained on high-quality text to identify and remove low-quality data, achieving data set quality optimization and providing more reliable input for subsequent model training. Redundancy removal aims to identify and remove duplicate content at different granularities (such as sentences, paragraphs, and documents). The presence of duplicate information may affect the diversity of training data, further increasing instability during training. Therefore, the de-duplication operation is crucial to ensure the richness of the dataset and the stability of the model. Privacy elimination mainly involves removing any data that may contain personal sensitive information from the dataset, such as user names, addresses, and phone numbers. To reduce the risk of privacy leakage, these sensitive data will be deleted to ensure that the data processing process complies with privacy protection regulations and ensures data security. Data labeling aims to build an instruction fine-tuning dataset and fine-tune the model on this dataset to adapt to downstream tasks. Typically, instruction instances consist of instructions (task descriptions), input-output pairs, questions, and answers. By using templates to convert labeled natural language datasets into instruction format <input, output> pairs, structured data is provided for model fine-tuning.

[0033] Model development uses a multi-modal encoder to extract key features from input data and align multi-modal features, converting them into a unified sequence representation. The model base is a Transformer with an attention mechanism at its core.

[0034] Step 2, use multi-modal encoding to fully extract sequence features of the inspection field instruction dataset, and input the extracted sequence features into a base model with a multi-layer Transformer model architecture for deep feature extraction. By introducing domain expert knowledge through the retrieval enhancement generation method, the marketing inspection large model is constructed.

[0035] The power inspection audit large model is constructed to provide a unified and general solution for multi-modal and multi-scenario task processing in the field of power marketing inspection.

[0036] The process mainly involves converting data of different modalities into a unified feature space, and then implementing deep extraction and fusion of features through a Transformer-based architecture. Specifically, first, various forms of data such as text, time series data, etc. are converted into a unified sequence representation through a multi-modal encoder, and the features of these data are aligned. Then, based on these multi-modal data, a Transformer architecture based on a self-attention mechanism is used for deep modeling to extract key features from various data. Finally, the multi-modal decoder is used for feature reconstruction to ensure that text and time series data can be efficiently processed and obtain accurate output.

[0037] The multi-modal encoder is constructed to extract features from the input multi-modal data and align the data of different modalities into a unified feature space, providing a unified input representation for subsequent deep modeling. For input text data, a WordPiece-based tokenizer is used to split it into token sequences, and a word embedding matrix of a large model is used to map it to a numerical vector representation. Each text token is mapped to an embedding vector , constituting the text embedding , where M is the total number of tokens. Time series data usually has strong local stationarity, and sliding window technology is used to divide it into several patches, and each patch is treated as a time series token for processing. These time series features are converted into time series embeddings , where P is the length of each window, , and the embedding space dimension of the large model. Since there are modal differences between text data and time series data, the multi-head attention mechanism is used to convert time series features into text features to enhance the cross-modal understanding and reasoning ability of the large model.

[0038] The multi-head self-attention mechanism (MHSA) is used to fuse cross-modal features. The query matrix of the time series features is defined as , the key matrix , and the value matrix , where E is the embedding matrix of the large model. The attention score matrix S is calculated to fuse time series information and text information:

[0039]

[0040] Further, a fully connected layer The fused cross-modal features are mapped to the text embedding space of the larger model. Through the multimodal feature extraction and alignment steps described above, the model can effectively process data from different modalities such as text and time series, and map them uniformly into the same feature space.

[0041] Deep feature extraction based on GPT-2 uses the GPT-2 model based on the Transformer architecture to perform deep feature extraction on multimodal features. The GPT-2 model is composed of multiple stacked Transformer decoders, where each decoder layer consists of a multi-head self-attention mechanism (MHSA), a feedforward neural network (FFN), and layer normalization (LN).

[0042] First, define the input representation of the model as... This is a concatenated representation of text and temporal embeddings. Through each layer of the model, the input... The input representation is processed sequentially through a multi-head self-attention module, layer normalization, and a feedforward neural network to generate the next layer's input representation. The mathematical expression for this process is:

[0043]

[0044]

[0045] By progressively stacking multiple Transformer decoders, the model is able to extract rich contextual information from the input multimodal features, forming a deep feature representation.

[0046] Finally, after N layers of processing, the model's output representation This will be used for subsequent decoding tasks, and its formal expression is as follows: ,in It is an intermediate state after passing through the multi-head self-attention mechanism and the first normalization layer.

[0047] The model reconstructs features using a multimodal decoder, ensuring that the output text and temporal data can be restored to a format consistent with the input data. The decoding process consists of two main modules: a text decoder and a temporal decoder. Each decoder recovers the original data by mapping its input features back to the target data space. The text decoder is based on a Transformer decoder structure and generates text features stepwise. We define the input to the text decoder as... , No. The formula for generating the step is: ,in This is the output from the previous time step. The text decoder uses a masked multi-head self-attention mechanism to ensure that each generated token can only depend on previous tokens. This process can be represented as:

[0048]

[0049] ( )

[0050]

[0051] wherein denotes the hidden layer state of the step, denotes the intermediate state after the multi-head self-attention layer and the first normalization layer of the step, denotes a vocabulary mapping matrix, denotes the size of the vocabulary, denotes a bias term. Considering the local stationarity and long-term dependency of time series, the time series decoder combines causal convolution operations to capture temporal dependencies. Specifically, the input of the time series decoder is

[0052] , and its output is processed by dilated causal convolution:

[0053] ,

[0054] wherein the convolution operation captures local temporal patterns through a convolution kernel of size 7. To enhance the expressive power of the time series decoder, the local pattern is fused with the output of the time series encoder by incorporating a gating mechanism:

[0055]

[0056] wherein , denotes the Hadamard product, denotes the concatenation operation, is a sigmoid gating function, denotes the local contextual features extracted by the time series decoder at the current time step t, denotes the global contextual representation output by the time series encoder, which is finally mapped back to the original time series through a linear layer: .

[0057] Based on the retrieval-enhanced generation of expert knowledge injection, the retrieval-enhanced generation technology is introduced, which expands the knowledge boundary of the model by an external knowledge base, enhances the professionalism and timeliness of the generated model without modifying the parameters of the large model. The retriever encodes the query and retrieves the relevant documents from the external document library, and uses cosine similarity to evaluate the similarity between the query and the document:​ where is the query embedding, d is the document embedding. The retriever is fine-tuned so that it can retrieve documents that can effectively reduce the perplexity of the large model. Specifically, the optimization goal of the retriever is to minimize the following KL divergence loss:

[0058]

[0059]

[0060]

[0061] where, is the document distribution generated by the retriever, is the probability distribution output by the large model, and the minimization goal of the KL divergence is to make the retrieved documents and the output generated by the model more consistent. and are model parameters.

[0062] In the retrieval stage, the retriever selects the top documents most relevant to the query and inputs them to the large model along with the query. The final generated prediction probability is the weighted average of all documents and queries:

[0063]

[0064] where is the probability output by the large model given the query and the document.

[0065] Model expansion integrates domain expert knowledge by integrating retrieval enhancement generation modules and fine-tunes retriever parameters using perplexity scores, freezes large model parameters to expand large model knowledge boundaries and alleviate large model hallucination.

[0066] Step 3, use the inspection field instruction data set to pre-train and fine-tune the constructed marketing inspection large model, and use the PPO algorithm based on heterogeneous strategy optimization to adjust the instruction fine-tuning model parameters. The text, time sequence and other data types in the power marketing inspection scene are uniformly modeled, understood and generated. Task, and adopt the PPO algorithm based on heterogeneous strategy optimization to adjust the instruction fine-tuning model parameters.

[0067] The process of pre-training the power multi-modal inspection review large model on the global data set in the field of power inspection aims to realize unified modeling, understanding and generation of multi-modal data (including text, time series, etc.) in the power inspection scene. Specifically, this step includes two main sub-steps: training a multi-modal tokenization encoder and reconstruction module, and pre-training a power inspection review large model. Through the above two steps, the model can process data from different modalities and extract deeper feature representations from the data through the large model, achieving deep semantic understanding of work order text.

[0068] Training a multi-modal tokenization encoder and reconstruction module first preprocesses multi-modal data (including text and time series data), converts text data into word embeddings in a high-dimensional vector space using tokenization techniques and embedding methods, and segments time series data. The training process of the encoder can be formally represented as follows: wherein is the input data token, is the feature vector after encoder processing. The role of the reconstruction module is to restore the extracted feature vector to the original sequence. The training process of the decoder is as follows: wherein is the reconstructed data sequence. The reconstruction module is trained by minimizing the reconstruction error, where the loss function combines the cross-entropy loss of text data and the mean square error loss of time series data:

[0069]

[0070] is the error loss weight, is the model prediction.

[0071] Through this collaborative optimization process, it is ensured that the model not only effectively extracts multi-modal features, but also maintains the structure and information of the data in the reconstruction process, thereby improving the performance of the model in downstream tasks.

[0072] Model pre-training first trains the tokenization encoder and decoder on the input inspection global data set, encodes multi-modal data into continuous feature vectors, and uses the decoder to reconstruct them; then pre-train the marketing inspection large model, apply a unified autoregressive prediction optimization objective to the multi-modal interleaved sequence, and realize the understanding and generation tasks of unified multi-modal data.

[0073] The pre-trained power inspection audit large model is unsupervisedly trained on a large-scale high-quality power marketing inspection global dataset. Specifically, after the unified feature sequence is obtained by encoding the input multi-modal data, the GPT-2 model is used for deep feature extraction to perform complex nonlinear changes and extract higher-level abstract features. Then, the feature sequence is reconstructed by the multi-modal decoder trained above. The model uses unified autoregressive prediction for optimization to achieve unified multi-modal data understanding and generation tasks.

[0074] The model is fine-tuned on the instruction fine-tuning dataset using the instruction fine-tuning and improved PPO algorithm, aiming to enhance the pre-trained marketing inspection large model's understanding and execution ability of natural language processing task instructions by combining instruction fine-tuning and improved PPO algorithm, while ensuring that the model's output conforms to human values and preferences. This process combines supervised fine-tuning, reward model training, and PPO fine-tuning based on hetero-strategy optimization to optimize the model's behavior, making it perform more accurately and reliably in real-world tasks in the field of power marketing inspection.

[0075] The pre-trained model is fine-tuned by first fine-tuning the pre-trained marketing inspection large model on the artificially annotated inspection domain instruction dataset. This dataset contains specific task instructions and their correct outputs, and the model learns the mapping between these instructions and outputs to improve its adaptability to specific tasks. During the fine-tuning process, the model adjusts its parameters by minimizing the loss between the instruction output and the true label, which is mathematically expressed as:

[0076]

[0077] where, is the predicted output generated by the model at time step t, is the model's predicted probability based on the input instruction and the model's parameters T is the length of the output sequence;

[0078] The instruction fine-tuning model trains a reward model to evaluate the quality and compliance of the model's generated output. The reward model's output value reflects the degree of matching between the generated output and human expectations, with higher rewards indicating that the model's output better aligns with human preferences and values. The reward model is trained based on the artificially annotated dataset using a loss function improved by imitation learning, which calculates the difference between the generated output and the expected output and introduces an autoregressive language model loss in the output for optimization. The loss function used in the reward model training is:

[0079]

[0080] where denotes the empirical distribution of the training data set, and in the model, denotes the likelihood probability under the given input prompt and the preferred output . and is the loss weight, is the rejection output, is the Sigmoid function.

[0081] PPO fine-tuning instructions based on heterogeneous strategy optimization fine-tune the marketing inspection large model, so that the model generates output that meets human preferences and values when given task instructions, thereby improving the reliability and generalization ability of the model. PPO algorithm is a reinforcement learning method based on heterogeneous strategy optimization, which limits the amplitude of each strategy update to maintain the stability of training, and uses importance sampling to evaluate the difference between the current strategy and the old strategy. The basic form of the strategy gradient of the PPO algorithm is as follows:

[0082]

[0083]

[0084] Model fine-tuning combines supervised instruction fine-tuning and PPO algorithm based on heterogeneous strategy optimization to enhance the model's ability to understand and follow human instructions in specific scenarios, while aligning the output to human values and preferences through human feedback to improve the quality and reliability of the model's answers.

[0085] Step 4, based on the improved adaptive low-rank adaptation technology, the model part layer weight matrix is approximated as the product of low-rank matrices, and API interface services are provided to detect abnormalities in input marketing inspection industry work order data.

[0086] The low-rank adaptation method improved based on singular value decomposition compresses the parameter quantity updated when the model adapts to downstream tasks, and realizes the strong quantization deployment and application of the inspection and audit large model by calling the API interface.

[0087] The low-rank adaptation method improved based on singular value decomposition compresses the parameters of the model to reduce the storage and computing requirements of the model. Specifically, the weight matrix is decomposed using singular value decomposition, and then the rank of the LoRA module in different layers is dynamically adjusted according to the gradient information, and its parameter update process can be formalized as:

[0088]

[0089] where, and are orthogonal, is a diagonal matrix containing singular values of. In the actual training process, the importance of singular values is adjusted according to the gradient change of the weight. By calculating the contribution degree of each singular value and pruning the singular values according to the size of the contribution degree, dynamic rank adjustment is realized. The goal of this process is to minimize redundant parameter updates while maintaining the effectiveness and robustness of the model.

[0090] The inspection work order data is applied to downstream tasks by calling the marketing inspection large model API. By testing the model using power marketing inspection work order data, intelligent research and analysis of abnormal work orders are performed, thereby helping business personnel to more efficiently identify potential abnormal or irregular work orders, and significantly improving single review efficiency.

[0091] The abnormal work order identification and research result analysis outputs the classification and abnormal check result of the work order text of the power industry user, and the probability is 1, which means that the work order meets the industry regulations, and the lower the probability, the more the model judges that the work order does not meet the industry regulations.

[0092] The second aspect of the present application relates to an inspection and audit large model based on multi-modal supervised learning and an intelligent work order audit system. The system is implemented using the method described in the first aspect of the present application. The system includes a multi-modal instruction data set construction module, an inspection and audit large model development module, a model pre-training and fine-tuning module, and a lightweight deployment and application module. The multi-modal instruction data set construction module collects power marketing inspection data from the marketing business support system, cleans and labels the collected power marketing inspection data, and forms an inspection field instruction data set, which is input into the model pre-training and fine-tuning module. The inspection and audit large model development module uses multi-modal encoding to fully extract the sequence features of the inspection field instruction data set, and inputs the extracted sequence features into a base model with a multi-layer Transformer model architecture for deep feature extraction. The model is constructed by introducing domain expert knowledge through a retrieval enhancement generation method. The model pre-training and fine-tuning module pre-trains and fine-tunes the constructed marketing inspection large model using the inspection field instruction data set, and performs unified modeling, understanding and generation tasks on text, time series and other data types in the power marketing inspection scenario. The PPO algorithm based on heterogeneous strategy optimization is used to adjust the instruction fine-tuning model parameters. The lightweight deployment and application module approximates part of the layer weight matrix of the model to the product of low-rank matrices based on the improved adaptive low-rank adaptation technology, and provides API interface services for abnormal detection of input marketing inspection industry work order data.

[0093] In a third aspect, the present application provides a terminal comprising a processor and a storage medium; the storage medium is configured to store instructions; the processor is configured to operate according to the instructions to perform the steps of the method in the first aspect of the present application.

[0094] In a fourth aspect, the present application provides a computer readable storage medium having stored thereon a computer program which, when executed by a processor, carries out the steps of the method in the first aspect of the present application.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate but not to limit the technical solutions of the present application. Although the present application has been described in detail with reference to the above embodiments, it should be understood by those of ordinary skill in the art that the technical solutions of the present application still include modifications or equivalent replacements to the specific embodiments of the present application. Any modifications or equivalent replacements not departing from the spirit and scope of the present application should be covered within the protection scope of the claims of the present application.

Claims

1. A large-scale audit and review model based on multimodal supervised learning and an intelligent work order review method, characterized in that: The method includes the following steps: The marketing business support system collects electricity marketing inspection data, cleans and labels the collected electricity marketing inspection data to form an inspection domain instruction dataset, and inputs it into the model pre-training and fine-tuning module; Multimodal coding is used to extract sequence features from the instruction dataset in the audit domain. These extracted sequence features are then input into a base model based on a multi-layer Transformer model for deep feature extraction. Domain expert knowledge is incorporated through retrieval-enhanced generation methods to construct a large-scale marketing audit model. Specifically, this includes: By utilizing a multi-layer Transformer architecture to perform deep nonlinear operations on multimodal features, a deep feature representation of the entire marketing audit data domain is achieved. A multimodal decoder is constructed to perform text reconstruction and time series reconstruction on marketing audit industry text and time series data to obtain multimodal output. Search enhancement generation technology is used to externally expand the knowledge base of the marketing audit big model, and automatically expand the prompt words that match the questions without updating the parameters of the big model. The marketing audit model was pre-trained and fine-tuned using an audit domain instruction dataset. This enabled unified modeling, understanding, and task generation for different types of data in the power marketing audit scenario. The PPO algorithm based on heterogeneous strategy optimization was used to adjust the parameters of the instruction fine-tuning model. Based on the improved adaptive low-rank adaptation technique, the weight matrix of some layers of the model is approximated as the product of low-rank matrices, and an API interface service is provided to perform anomaly detection on the input marketing audit work order data.

2. The large-scale audit and review model based on multimodal supervised learning and the intelligent work order review method according to claim 1, characterized in that: The process involves collecting electricity marketing audit data from the marketing business support system, cleaning and labeling the collected data to form an audit domain instruction dataset, and inputting it into the model pre-training and fine-tuning module, including: Acquire industry-wide marketing audit data involved in the marketing business support system. This industry-wide marketing audit data includes multi-source, multi-modal information such as audit work orders, rules and regulations, policy reports, and the number of work orders per day. The collected data from the entire domain was preprocessed, and the data annotation was used to obtain a fine-tuning dataset for power marketing inspection instructions.

3. The large-scale audit and review model and intelligent work order review method based on multimodal supervised learning according to claim 2, characterized in that: The method of fully extracting sequence features from the inspection domain instruction dataset using multimodal coding includes: The key features of multimodal data, including marketing audit industry work order data, regulations, policy reports, and daily work order numbers, are extracted, aligned, and reconstructed using a multimodal encoding-decoding structure.

4. The large-scale audit and review model based on multimodal supervised learning and the intelligent work order review method according to claim 3, characterized in that: The process of extracting, aligning, and reconstructing key features from multimodal data, including marketing audit industry work order data, regulations, policy reports, and daily work order quantities, using a multimodal encoding-decoding structure, includes: For unstructured text data including inspection work orders and regulations, a large model word segmenter is used to segment the original text, construct a numerical vector representation of each word according to the vocabulary, and extract text features through a text encoder. The structured time-series data, including daily and monthly work order volumes, is segmented into multiple patches. The time-series features are transformed into text features through a multi-head attention mechanism, and a linear mapping layer is used to output a time-series representation with the same dimension as the text vector. A multimodal encoder is used to map different modal data to a unified feature space, and a fully connected layer is used to map different modal data in the unified feature space to the text embedding space of a large model.

5. The large-scale audit and review model based on multimodal supervised learning and the intelligent work order review method according to claim 4, characterized in that: The aforementioned pre-training and fine-tuning of the constructed marketing audit model using an audit domain instruction dataset, and the unified modeling, understanding, and generation tasks for text and time-series data types in the power marketing audit scenario, include: The base model is used to perform deep feature representation of unified multimodal features, and the original input is reconstructed and output through a trained text decoder and a temporal decoder.

6. The large-scale audit and review model based on multimodal supervised learning and the intelligent work order review method according to claim 5, characterized in that: The improved adaptive low-rank adaptation technique approximates the weight matrices of some layers of the model as products of low-rank matrices, including: An adaptive low-rank adaptation technique based on singular value decomposition freezes the remaining parameters of the large model and only updates the bias parameters of the Transformer network or the parameters of the linear layers.

7. The large-scale audit and review model based on multimodal supervised learning and the intelligent work order review method according to claim 6, characterized in that: The improved adaptive low-rank adaptation technique approximates the weight matrix of some layers of the model as a product of low-rank matrices, and provides an API interface service for anomaly detection of input marketing audit industry work order data, including: The marketing audit big data model API is called to generate audit work order data. The model is tested using the power marketing audit work order data to perform intelligent judgment and analysis of abnormal work orders.

8. A large-scale audit and review model and an intelligent work order review system based on multimodal supervised learning, utilizing the method described in any one of claims 1-7, characterized in that: The system includes a multimodal instruction dataset construction module, an audit and review large model development module, a model pre-training and fine-tuning module, and a lightweight deployment and application module; The multimodal instruction dataset construction module is formed by the marketing business support system collecting electricity marketing inspection data, cleaning and labeling the collected electricity marketing inspection data to form an instruction dataset in the inspection field, which is then input into the model pre-training and fine-tuning module. The large-scale audit and review model development module utilizes multimodal coding to extract sequence features from the instruction dataset in the audit domain. The extracted sequence features are then input into a base model based on a multi-layer Transformer model for deep feature extraction. Domain expert knowledge is introduced through retrieval enhancement generation methods to construct a large-scale marketing audit model. The model pre-training and fine-tuning module uses the instruction dataset in the inspection field to pre-train and fine-tune the constructed marketing inspection model. It performs unified modeling, understanding and generation of different types of data in the power marketing inspection scenario, and uses the PPO algorithm based on heterogeneous strategy optimization to adjust the parameters of the instruction fine-tuning model. The lightweight deployment and application module, based on an improved adaptive low-rank adaptation technique, approximates the weight matrix of some layers of the model as a product of low-rank matrices, and provides API interface services to perform anomaly detection on the input marketing audit industry work order data.

9. A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the steps of the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Joint entity relation extraction method based on semi-supervised learning and large language model

    CN120562418A

  • Large language model knowledge preference alignment method and system based on self-supervised learning

    CN120562561A