Method for constructing power marketing multi-modal large model demonstration application platform integrated with multiple scene functions and power marketing multi-modal large model demonstration application platform integrated with multiple scene functions

CN122838533APending Publication Date: 2026-09-29GUANGXI POWER GRID CO LTD NANNING POWER SUPPLY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610858021.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

传统系统依赖人工设定的问答库和流程,处理能力有限;而通用多模态模型虽具备较强的跨模态理解基础,但缺乏电力营销的专业知识支撑,往往通过简单的微调或检索增强生成(RAG)方式适应业务

Benefits of technology

[0008]借由上述技术方案,本申请实施例提供的一种集成多场景功能的电力营销多模态大模型示范应用平台构建方法及集成多场景功能的电力营销多模态大模型示范应用平台,通过构建电力营销专业知识图谱并为多模态大模型提供外部知识注入与微调训练,使通用大模型获得电力营销领域的专业响应能力;通过意图识别与场景协同工作管线的路由分发机制,实现对不同任务类别的差异化处理;通过采集用户采纳行为并作为强化反馈信号更新模型权重,使模型在实际使用中持续优化,逐步贴合业务人员的操作习惯与判断偏好。以上手段协同作用,构建了一个集成多场景功能、具备专业深度和在线学习能力的电力营销多模态大模型示范应用平台。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838533A_ABST
    Figure CN122838533A_ABST
Patent Text Reader

Abstract

The application discloses a power marketing multi-modal large model demonstration application platform construction method integrating multiple scene functions and the power marketing multi-modal large model demonstration application platform integrating multiple scene functions. The method comprises the following steps: constructing a power marketing professional knowledge graph; injecting the power marketing professional knowledge graph as external incremental knowledge into a basic multi-modal large model, and generating a power marketing professional model through fine-tuning training; receiving a multi-modal interaction request sent by a business terminal, analyzing the task category of the multi-modal interaction request by using an intention recognition mechanism, and distributing the multi-modal interaction request to a corresponding scene collaborative work pipeline; in the corresponding scene collaborative work pipeline, calling the power marketing professional model to perform feature extraction and cross-modal alignment on the multi-modal interaction request, generating a professional execution result of a specific business scene, feeding back the professional execution result to the business terminal, collecting an adoption behavior action returned by the business terminal, and updating the neuron weight of the power marketing professional model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power information technology, and in particular to a method for constructing a multi-modal large-scale power marketing demonstration application platform integrating multiple scenario functions, and a multi-modal large-scale power marketing demonstration application platform integrating multiple scenario functions. Background Technology

[0002] With the deepening of power system reform and the rapid development of the energy internet, power marketing is transforming from traditional single-service processing to diversified, personalized, and interactive comprehensive energy services. Scenarios such as user inquiries, fault reporting, electricity bill calculation, contract management, and policy interpretation frequently involve multimodal information such as text, voice, images, and tables, and the business rules are complex, regulations are rapidly updated, and professional terminology is abundant. To improve service efficiency and user experience, power companies urgently need to build intelligent platforms capable of understanding multimodal inputs and integrating professional knowledge to achieve rapid response and accurate decision-making in different marketing scenarios.

[0003] Currently, the power marketing field generally uses rule-based traditional customer service systems or general pre-trained language models. Traditional systems rely on manually set question-and-answer databases and processes, resulting in limited processing capabilities. While general multimodal models possess a strong foundation in cross-modal understanding, they lack the support of professional knowledge in power marketing and often adapt to business needs through simple fine-tuning or retrieval-enhanced generation (RAG). At the system architecture level, existing solutions mostly adopt a centralized model of "one model handling all requests," or deploy multiple dedicated models independently for different business scenarios. The former struggles to meet the differentiated needs of multiple scenarios, while the latter leads to resource waste and information silos.

[0004] However, the aforementioned technical approaches still have significant shortcomings in practical applications: the depth of professional knowledge injected into the general model is insufficient, resulting in low accuracy in responding to electricity marketing-specific terminology, regulations, and operational procedures, easily leading to "illusions" or erroneous conclusions; a single model architecture cannot efficiently match diverse business scenarios, leading to high task processing latency and low resource utilization; simultaneously, existing systems lack effective utilization of user feedback, making it difficult to continuously optimize model performance. Therefore, how to construct a multimodal large-scale model platform that deeply integrates electricity marketing knowledge graphs, supports multi-scenario collaborative scheduling, and possesses online learning capabilities has become an urgent technical problem to be solved in this field. Summary of the Invention

[0005] In view of this, embodiments of this application provide a method for constructing a multi-modal large-scale power marketing demonstration application platform integrating multiple scenario functions, and a multi-modal large-scale power marketing demonstration application platform integrating multiple scenario functions.

[0006] According to one aspect of this application, a method for constructing a multimodal large-scale demonstration application platform for power marketing integrating multiple scenario functions is provided, the method comprising: Collect historical business data in the field of electricity marketing, preprocess the historical business data, and construct a professional knowledge map of electricity marketing that includes professional terms, business regulations and operating procedures based on the preprocessed historical business data; Obtain a basic multimodal large model, inject the power marketing professional knowledge graph as external incremental knowledge into the basic multimodal large model, and generate a power marketing professional model through fine-tuning training; Receive multimodal interaction requests sent by business terminals, use an intent recognition mechanism to parse the task category of the multimodal interaction requests, and route and distribute the multimodal interaction requests to the corresponding scenario collaborative work pipeline; In the corresponding scenario collaborative work pipeline, the power marketing professional model is invoked to perform feature extraction and cross-modal alignment on the multimodal interaction request, generating professional execution results for specific business scenarios; The professional execution results are fed back to the business terminal, and the adoption behavior actions returned by the business terminal are collected. The adoption behavior actions are used as feedback reinforcement signals to update the neuron weights of the power marketing professional model.

[0007] According to another aspect of this application, a multi-modal large-scale power marketing demonstration application platform integrating multiple scenario functions is provided, which is constructed by the above-described method for constructing a multi-modal large-scale power marketing demonstration application platform integrating multiple scenario functions.

[0008] By employing the above technical solutions, this application provides a method for constructing a multi-scenario large-scale power marketing model demonstration application platform and a multi-scenario large-scale power marketing model demonstration application platform. This is achieved by constructing a professional knowledge graph of power marketing and providing external knowledge injection and fine-tuning training for the multi-modal large-scale model, enabling the general large-scale model to acquire professional responsiveness in the power marketing field. Through the routing and distribution mechanism of intent recognition and scenario collaborative work pipelines, differentiated processing of different task categories is realized. By collecting user adoption behavior and using it as a reinforcing feedback signal to update model weights, the model is continuously optimized in actual use, gradually aligning with the operating habits and judgment preferences of business personnel. The synergistic effect of these methods constructs a multi-scenario large-scale power marketing model demonstration application platform with integrated multi-scenario functions, professional depth, and online learning capabilities.

[0009] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description

[0010] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This document illustrates a flowchart of a method for constructing a multi-modal large-scale demonstration application platform for power marketing that integrates multiple scenario functions, as provided in an embodiment of this application. Figure 2 This paper illustrates a flowchart of another method for constructing a multi-modal large-scale demonstration application platform for power marketing that integrates multiple scenario functions, as provided in an embodiment of this application. Figure 3 This illustration shows a flowchart of another method for constructing a multi-modal large-scale demonstration application platform for power marketing that integrates multiple scenario functions, provided in an embodiment of this application. Detailed Implementation

[0011] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0012] This embodiment provides a method for constructing a multi-modal large-scale demonstration application platform for power marketing that integrates multiple scenario functions, such as... Figure 1 As shown, the method includes: Step 101: Collect historical business data in the field of electricity marketing, preprocess the historical business data, and construct a professional knowledge map of electricity marketing that includes professional terms, business regulations and operating procedures based on the preprocessed historical business data.

[0013] Historical business data refers to data such as text work orders, voice call records, scanned copies of contracts, policy documents, and user operation logs generated during the business processing in the field of electricity marketing; preprocessing includes format unification, missing value filling, noise filtering, and text segmentation of unstructured data; the electricity marketing professional knowledge graph refers to an entity relationship network stored in a graph database, whose nodes contain professional terms, business regulations and operating procedures, and edges represent subordinate, causal, or sequential relationships between entities.

[0014] In this embodiment, historical business data is first collected from the power company's data platform. Data sources include customer service call recordings, business receipts from service halls, online service hall interaction logs, and internal training manuals. After collection, the data undergoes preprocessing: for text data, regular expressions are used to remove irrelevant symbols and garbled characters, and terminology is segmented using a word segmentation tool combined with a power industry dictionary; for voice data, it is first transcribed into text using a speech recognition model and then cleaned; for image-based contracts or invoices, optical character recognition is used to extract key fields. The preprocessed data is then sent to the knowledge graph construction module. Specifically, in one implementation, entity types are first defined (e.g., "electricity price policy," "default electricity usage behavior," "business expansion application process"), then specific entities are extracted using a named entity recognition model, and finally, relationship extraction models are used to establish relationships between entities. In another implementation, based on co-occurrence analysis and association rule mining, high-frequency terms and co-occurrence patterns are automatically discovered from the preprocessed text, and then manually verified to form a knowledge graph structure. When logical conflicts between entities are detected during the knowledge graph construction process (such as two contradictory interpretations of the same legal provision), the conflicting nodes are automatically marked and a manual review process is triggered. The knowledge graph is updated only after the review is completed. Furthermore, after the knowledge graph is built, a graph embedding algorithm is used to map entities and relationships into low-dimensional vector representations, facilitating subsequent model invocation. Thus, by constructing a knowledge graph, fragmented knowledge in the power marketing field can be structured and organized, providing support for injecting knowledge into subsequent models.

[0015] Step 102: Obtain the basic multimodal large model, inject the power marketing professional knowledge graph as external incremental knowledge into the basic multimodal large model, and generate the power marketing professional model through fine-tuning training.

[0016] Among them, the basic multimodal large model refers to a pre-trained large model that can simultaneously process input signals such as text, images, and speech, and contains a cross-modal attention mechanism; external incremental knowledge injection refers to introducing structured information from the knowledge graph as additional input or constraints into the model training process without changing the original parameter architecture of the model; fine-tuning training refers to using labeled data from the power marketing field to make targeted adjustments to some parameters of the model based on the general capabilities of the basic model.

[0017] In this embodiment, a basic multimodal model, either open-source or self-built, is first obtained. In one implementation, a pre-trained model supporting text, image, and speech trimodality is selected, with its input layers containing text encoders, image encoders, and speech encoders, respectively. A power marketing knowledge graph is injected into the basic model as external incremental knowledge. The injection can be implemented in several ways: First, entity vectors from the graph are concatenated with the hidden states of each layer of the basic model, adding them as additional key-value pairs to the cross-attention module; second, a graph neural network is trained on the graph, and the output graph-aware vector is added to the embedding vector at the model input, achieving implicit injection; third, triples (head entity, relation, tail entity) from the graph are serialized into natural language prompts and concatenated before user input, serving as explicit context injection. After injection, fine-tuning training is performed: multimodal labeled data in power marketing scenarios (such as fault reporting samples containing electricity marketing intent and user voice) are collected, and low-rank adaptation or adapter modules are used to efficiently fine-tune the parameters of the basic model. Through this step of injection and fine-tuning, the general multimodal model gains the professional knowledge response capability in the field of power marketing, and the generated results are more in line with the actual business.

[0018] Step 103: Receive the multimodal interaction request sent by the business terminal, use the intent recognition mechanism to parse the task category of the multimodal interaction request, and route and distribute the multimodal interaction request to the corresponding scene collaborative work pipeline.

[0019] Among them, multimodal interaction requests refer to mixed-type inputs initiated by business terminals, which may simultaneously include information such as text descriptions, voice commands, on-site photos, or table screenshots; the intent recognition mechanism refers to the algorithm module that classifies the request content and outputs task category labels (such as regional digital marketing scenarios, customer service risk assessment scenarios, and grassroots on-site operation inspection scenarios); the scenario collaborative work pipeline refers to the predefined processing flow links for different task categories, with each pipeline containing independent input validation, model call parameters, and output formatting rules.

[0020] In this embodiment, a multimodal interaction request is received from a business terminal (such as a customer service computer, mobile inspection device, or self-service terminal). First, an intent recognition mechanism is invoked to parse the task category of the request. In one implementation, a fast discrimination method based on a lightweight text classifier is used: extracting text or speech-to-text characters from the request, calculating the similarity with predefined category keywords, and outputting the category with the highest confidence. In another implementation, a multimodal joint classification network is used: encoding image, speech, and text features separately and then fusing them, outputting the category distribution through a soft maximization layer. After determining the task category, the request is routed and distributed to the corresponding scene collaboration pipeline.

[0021] In one optional embodiment, an intent recognition mechanism is used to parse the task category of the multimodal interaction request, and the multimodal interaction request is routed and distributed to the corresponding scene collaboration pipeline, such as... Figure 2 As shown, it includes: Step 201: Use a sequence encoder to extract the context semantic vector and modal feature vector from the multimodal interaction request; Step 202: Concatenate the context semantic vector and the modal feature vector, input the concatenation into a fully connected classification network, and calculate the matching probability distribution of the multimodal interaction request belonging to the regional digital marketing scenario, the customer service risk assessment scenario, and the grassroots on-site operation inspection scenario. Step 203: Select the scenario with the largest value in the matching probability distribution as the target business scenario, activate the scenario collaboration pipeline bound to the target business scenario, and load the data payload of the multimodal interaction request into the activated scenario collaboration pipeline.

[0022] Among them, a sequence encoder refers to a neural network module that can extract features in the order of input and preserve temporal dependencies, and can be a recurrent neural network, a transformer encoder, or a variant thereof; a contextual semantic vector refers to a fixed-length representation extracted by the sequence encoder from multimodal interaction requests that reflects the overall contextual meaning; a modal feature vector refers to an independent feature representation obtained after encoding a single modality such as text, speech, or image; a fully connected classification network refers to a classifier composed of multiple linear layers and nonlinear activation functions, whose output is the matching strength of each scenario category. Regional digital marketing scenarios refer to business types such as the promotion of electricity products, the recommendation of electricity packages, and the marketing of energy efficiency services targeting specific administrative regions or user groups; customer service risk assessment scenarios refer to business types that identify potential service risks or credit risks based on user complaint records, abnormal payment behavior, or contract performance status; and grassroots field operation scenarios refer to the types of operations such as taking photos, recording audio, and filling out forms submitted by inspectors, meter readers, or repair personnel at user sites via mobile terminals.

[0023] In this embodiment, a sequence encoder is first used to extract the context semantic vector and modal feature vector from the multimodal interaction request. In one implementation, a pre-trained transformer encoder is used as the sequence encoder to concatenate the text, image description, and speech-transcribed characters in the multimodal request into a labeled sequence in the input order. After multi-layer self-attention calculation, the last hidden state is taken as the context semantic vector. Simultaneously, independent text encoders, image encoders, and speech encoders are used to extract the specific features of each modality, and these feature vectors are compressed into modal feature vectors through average pooling. After extraction, the context semantic vector and the modal feature vector are concatenated to obtain a joint feature vector. This joint feature vector is input into a fully connected classification network to calculate the matching probability distribution of the current multimodal interaction request belonging to the regional digital marketing scenario, customer service risk assessment scenario, and grassroots field operation scenario, respectively. In one implementation, the fully connected classification network uses two-layer linear transformation. The first layer maps the joint feature vector to 256 dimensions, and after an activation function, the second layer maps it to 3 dimensions. Then, a soft maximization function is used to output the normalized probability. In another implementation, a multi-task classification network is used, with three independent binary classifiers on top of a shared hidden layer. These classifiers determine whether a scenario belongs to a digital marketing scenario, a risk assessment scenario, or an on-site operation scenario, respectively. The confidence scores of the three binary classification results are then combined to form a probability distribution. After obtaining the probability distribution, the scenario with the highest confidence score is selected as the target business scenario. Once the target business scenario is determined, the scenario collaboration pipeline bound to that scenario is activated. Specifically, the activation operation may include: loading the dedicated model parameters required for the pipeline, allocating an independent computing resource queue, and initializing the output format template. Subsequently, the original data payload of the multimodal interaction request (including user-uploaded images, audio files, form fields, etc.) is loaded into the activated pipeline, ready for subsequent processing. Through the sequence encoding and fully connected classification network design of this embodiment, multimodal interaction requests can be accurately identified as specific power marketing business scenarios, avoiding misjudgments of complex requests by a single classifier.

[0024] In one alternative embodiment, after parsing the task category of the multimodal interaction request using an intent recognition mechanism, such as... Figure 3 As shown, the method further includes: Step 301: Based on the task category of the multimodal interaction request obtained from the parsing, extract the corresponding role setting clause, task constraint clause and security boundary clause from the preset instruction template library; Step 302: Concatenate the role setting clause, the task constraint clause, and the security boundary clause into a system-level guidance prompt; Step 303: Force the system-level guidance prompt to be placed in the header input sequence of the multimodal interaction request in order to limit the output style and business boundaries of the power marketing professional model.

[0025] The instruction template library pre-stores a knowledge base of various structured text fragments, each fragment corresponding to a specific task category's role definition, behavioral constraints, or compliance boundaries. Role setting clauses are natural language descriptions used to assign a specific business identity to the model, such as "You are an electricity marketing customer service expert" or "You serve as a grassroots on-site operation inspection assistant." Task constraint clauses are instructions that clearly limit the scope of the model's output content, such as "Only answer questions related to electricity billing, not involving other areas" or "The output results must include the location of the hidden danger, the type of the hidden danger, and disposal suggestions." Safety boundary clauses are prohibitive rules that prevent the model from generating illegal, misleading, or unauthorized content, such as "No compensation may be promised beyond the power company's authority" or "It is forbidden to draw conclusions directly about unverified electricity theft."

[0026] In this embodiment, after parsing the task category of the multimodal interaction request using the intent recognition mechanism and before sending the request into the power marketing professional model, the following guidance prompt injection operation is performed. First, based on the parsed task category (such as regional digital marketing, customer service risk assessment, grassroots on-site operation inspection and its sub-scenarios), the corresponding role setting clause, task constraint clause, and security boundary clause are extracted from a preset instruction template library. In one implementation, the instruction template library is stored in key-value pair format, with each task category corresponding to a set of independent triplet templates. For example, the role setting clause corresponding to the regional digital marketing scenario is "You are a power marketing contract specialist," the task constraint clause is "Contract text is generated only based on the user-provided identification information and on-site survey parameters, and fictitious clauses are not allowed," and the security boundary clause is "It is forbidden to include any false electricity price promises, and the contract's effectiveness is subject to the final approval of the power supply bureau." After extraction, the role setting clause, task constraint clause, and security boundary clause are concatenated into a system-level guidance prompt. Then, the system-level guidance prompt is forcibly placed in the header input sequence of the multimodal interaction request. Specifically, for text-based inputs, the concatenated guidance prompt is added directly before the user's original text request, with explicit tokens (such as "[SYS]" and "[USER]") used to distinguish them. For multimodal inputs (such as image + text), the guidance prompt is inserted as text before all modal encoding: the model first encodes the guidance prompt, then sequentially encodes the user-provided text description, image features, or speech features. Finally, the input sequence with the system-level guidance prompt is fed into the power marketing professional model, which outputs professional execution results that conform to the business style and boundaries under the constraints of the prompt. By extracting task-adaptive clauses from the instruction template library and forcibly inserting them into the input header, the power marketing professional model can automatically switch to matching output styles and behavioral boundaries in different business scenarios, avoiding role confusion, unauthorized commitments, or illegal outputs that may occur with general models.

[0027] Step 104: In the corresponding scenario collaborative work pipeline, call the power marketing professional model to perform feature extraction and cross-modal alignment on the multimodal interaction request, and generate professional execution results for specific business scenarios.

[0028] Feature extraction refers to extracting task-related abstract representation vectors from multimodal inputs; cross-modal alignment refers to mapping features from different modalities to the same semantic space, establishing a correspondence between text descriptions and image regions and speech segments.

[0029] In this embodiment, within the scenario collaboration pipeline, a power marketing professional model matched to the scenario collaboration pipeline is invoked to process multimodal interaction requests. First, feature extraction is performed. For example, for text modality, a transformer encoder within the model extracts context-aware word-level features; for image modality, a visual transformer or convolutional network extracts tile-level features; for speech modality, Mel-spectrum processing is used before being fed into an audio encoder. Then, cross-modal alignment is performed. In one implementation, a contrastive learning alignment strategy is used to calculate the cosine similarity matrix between text features and image features. By maximizing the similarity of positive sample pairs (such as the text description of "damaged meter casing" and the damaged area in the photo), features from different modalities are brought closer together in the latent space. In another implementation, a cross-attention alignment strategy is used, with text features as queries and image features as keys and values, and attention weighting is used to obtain an aligned multimodal fusion representation. After alignment, the model generates professional execution results for specific business scenarios based on the fusion representation. Through feature extraction and cross-modal alignment in this step, the model can comprehensively utilize multi-source information to generate scenario-specific professional conclusions.

[0030] In an optional embodiment, when the target business scenario is the regional digital marketing scenario, the power marketing professional model performs feature extraction and cross-modal alignment on the multimodal interaction request to generate professional execution results for the specific business scenario, including: extracting the document scan image from the multimodal interaction request; locating the character bounding boxes in the document scan image using an optical character recognition algorithm and extracting the character sequence; inputting the character sequence into the natural language sequence annotation layer in the power marketing professional model, extracting key-value pair data representing customer identity and property address number through joint encoding of spatial coordinates and semantic features, and assembling them into a structured business expansion work order attribute set; obtaining on-site survey parameters; retrieving the corresponding power supply scheme template and standard legal clauses in the power marketing professional knowledge graph; using the text generation layer in the power marketing professional model, filling the business expansion work order attribute set and the on-site survey parameters into the blank slots of the power supply scheme template and the standard legal clauses; performing compliance lexical scanning on the filled text content to generate the regional digital marketing contract text to be signed.

[0031] Among them, scanned document images refer to photographs or scans of paper or electronic documents such as ID cards, business licenses, and property ownership certificates submitted by users when handling electricity business; optical character recognition algorithms refer to the computational process of detecting and extracting text information from images, including two sub-steps: character bounding box localization and character sequence recognition; natural language sequence labeling layers refer to neural network modules that assign a label to each position in the input sequence, often used for extracting named entities; joint encoding of spatial coordinates and semantic features refers to mapping the positional information of characters in an image with their textual semantic information into a unified representation; and the business expansion order attribute set refers to the set of attributes that includes customer identification. The system includes structured data records with fields such as property address number, electricity category, and installed capacity; on-site survey parameters refer to actual data measured by power company personnel, such as the location of the power supply access point, line direction, and transformer capacity; power supply scheme templates refer to predefined, standardized power supply access scheme text frameworks containing fillable slots; standard legal clauses refer to the legal clause templates that must be included in the power supply contract; text generation layer refers to the module that generates natural language text based on a pre-trained language model or template filling mechanism; compliance lexical scanning refers to checking the generated contract text for keywords and sentence structure to ensure that it does not violate power regulatory provisions or internal risk control requirements.

[0032] In this embodiment, when the identified target business scenario is a regional digital marketing scenario, the power marketing professional model performs the following operations on the multimodal interaction request. First, it extracts the document scan image from the multimodal interaction request. For example, it identifies and separates the image type field from the multimodal data payload of the request, and filters document images based on file names or metadata tags (such as "front of ID card" or "copy of business license"). After extracting the document scan image, it uses an optical character recognition algorithm to locate the character bounding boxes in the image and extract the character sequence. In one implementation, it uses a detection network based on connected text regions to generate character line-level bounding boxes, and then feeds them into a convolutional recurrent neural network for sequence recognition, outputting the character sequence corresponding to each line. After obtaining the character sequence, it inputs the character sequence into the natural language sequence annotation layer in the power marketing professional model, and extracts key-value pair data representing customer identity and property address number through joint encoding of spatial coordinates and semantic features. In one implementation, the sequence labeling layer employs a bidirectional long short-term memory network plus a conditional random field architecture. Each character is assigned a label such as "customer name" or "address." Simultaneously, the coordinates of the bounding box center points corresponding to each character are normalized and concatenated with the semantic embedding vector, enabling the model to utilize spatial layout (e.g., the "name" label is usually located below the photo) to aid in discrimination. In another implementation, a transformer-based multimodal document understanding model is used. Image block features and character sequence features are fused through cross-attention, directly outputting the key-value pair extraction results. After extraction, the key-value pair data is assembled into a structured set of business expansion work order attributes, such as {"customer name":"Zhang San","ID number":"410XXXXXXXXXXXXXXX","property address":"No. 10, XX Road, XX City, XX Province"}. Simultaneously, on-site survey parameters are acquired. These parameters can be filled in in real-time by frontline personnel via mobile terminals and attached to the interaction request in the form of a structured form. After acquiring the on-site survey parameters, the corresponding power supply scheme template and standard legal clauses are retrieved from the power marketing professional knowledge graph. The search criteria include electricity usage category (residential / industrial / commercial / agricultural), installed capacity, and power supply area. For example, a graph query language is used to construct a matching subgraph, returning the power supply scheme template node most similar to the input conditions and its associated legal clause node. Then, using the text generation layer in the power marketing professional model, the business expansion order attribute set and on-site survey parameters are filled into the blank slots of the power supply scheme template and standard legal clauses. In one implementation, the text generation layer uses a rule-based natural language generation engine, filling predefined slots in the template such as "[Customer Name]", "[Property Address]", and "[Power Supply Capacity]" through string replacement. After filling, a compliance lexical scan is performed on the filled text content.The scanning rules include: checking for prohibited words (such as "unconditional power supply" or "permanent free" – illegal promises), checking if all required slots are filled, and checking if numerical formats conform to specifications (such as voltage levels must be "220V" or "380V"). When the lexical scan detects violations or omissions, the problem location can be automatically marked and a correction process can be triggered. For example, for issues that can be automatically corrected (such as incorrect date formats), the code is directly rewritten according to specifications; for issues that cannot be automatically corrected (such as missing power supply capacity), a task is generated and pushed to business reviewers for manual completion. Finally, a regional digital marketing contract text to be signed is generated, ready to be fed back to the business terminal. Through the combined encoding of optical character recognition and sequence labeling with spatial coordinates in this embodiment, key customer information in document images can be accurately extracted and structured, avoiding manual input errors; through template filling and compliance lexical scanning, the generation of contract text has standardized and compliant guarantees, thus efficiently outputting formal contract documents that can be directly used for regional digital marketing business.

[0033] In an optional embodiment, when the target business scenario is the customer service risk assessment scenario, the power marketing professional model performs feature extraction and cross-modal alignment on the multimodal interaction request to generate professional execution results for the specific business scenario, including: extracting high-dimensional acoustic feature vectors from the customer service call audio in the multimodal interaction request using an acoustic coding network; extracting high-dimensional semantic feature vectors from the transcribed text corresponding to the customer service call audio using a text coding network; and calculating a service risk assessment index using a cross-modal bilinear projection resonance model, the calculation formula of which is:

[0034] in, This indicates the service risk assessment index; This represents the high-dimensional acoustic feature vector; This represents the high-dimensional semantic feature vector; Represents the Hadamard product of vectors; and Let L1 and L2 norms of the vector be represented respectively. This represents a pre-trained cross-modal coupling weight matrix used to map acoustic space features to text semantic space; Represents an exponential function with the natural constant as its base; This represents the resonance surge factor.

[0035] Among them, the acoustic coding network refers to a deep neural network used to extract paralinguistic information such as speaker's emotion, speech rate, pitch, and tremor from the original speech signal; the high-dimensional acoustic feature vector refers to the numerical representation output by the acoustic coding network that contains multiple dimensions (such as Mel-frequency cepstral coefficients, fundamental frequency, and energy change rate); the text coding network refers to a language model that converts the text transcribed from customer service calls into a high-dimensional semantic representation, which can adopt a transformer or recurrent neural network architecture; the cross-modal bilinear projection resonance model refers to a computational structure that simultaneously utilizes acoustic and text features and measures the degree of consistency or conflict between the two through bilinear interaction terms; the service risk assessment index refers to a continuous value that reflects the potential service quality risks (such as escalation of customer dissatisfaction, complaint tendency, and breach of contract implications) in the current customer service interaction process; and the cross-modal coupling weight matrix refers to a pre-trained learnable parameter matrix used to project acoustic features onto the text semantic space to achieve metric alignment between the two modalities.

[0036] In this embodiment, when the target business scenario is a customer service risk assessment scenario, the power marketing professional model performs feature extraction and cross-modal alignment on multimodal interaction requests to generate service risk assessment results. First, an acoustic coding network is used to extract high-dimensional acoustic feature vectors from the customer service call audio in the multimodal interaction requests. In one implementation, the acoustic coding network uses a pre-trained Wiener filter front-end combined with a temporal convolutional network to segment the audio into fixed-length frames. For each frame, Mel-spectral coefficients, linear prediction coefficients, and short-time energy are extracted, and after deep convolution processing, a 128-dimensional high-dimensional acoustic feature vector is output. In another implementation, a self-attention-based audio transformer is used to directly process the original waveform, output frame-level features, and then obtain a fixed-length acoustic vector through global average pooling. Simultaneously, a text coding network is used to extract high-dimensional semantic feature vectors from the transcribed text corresponding to the customer service call audio. In one implementation, the text coding network uses a lightweight pre-trained language model. The transcribed text is segmented and input into the model, and the average of all labeled vectors in the last layer is taken as the high-dimensional semantic feature vector. In another implementation, an encoder based on a long short-term memory network is used to process the text sequentially, concatenating the final hidden state with the attention-pooled text representation to form a richer semantic vector. After extracting the acoustic and text feature vectors, a cross-modal bilinear projection resonance model is used to calculate the service risk assessment index. This model is designed so that when acoustic features (such as a customer's rapid tone and increased volume) and text features (such as a customer saying words like "complaint," "refund," or "find a leader") resonate semantically, a higher risk index is output; conversely, if the two contradict each other (such as a complaint in the text but a calm tone), a lower risk index is output. In one specific implementation, the Hadamard product of two vectors is first calculated to obtain a joint excitation vector of the same dimension. The L1 norm of this vector is then used as the baseline resonance intensity. Simultaneously, the acoustic features are projected onto the text space through a cross-modal coupling weight matrix. The cosine similarity between the projected acoustic features and the original text features is calculated, and the influence of the similarity is amplified using an exponential function. The resonance surge factor is used to control the amplification degree (when the acoustic and text are highly consistent, this factor makes the exponential term significantly greater than 1, thereby increasing the risk index). If null values ​​or all-zero vectors are found in the acoustic or text feature vectors during the calculation process (such as missing audio or empty transcribed text), the feature vector corresponding to the missing modality can be automatically replaced with the average feature learned from historical samples of the same scenario. A "single-modal assessment" label is added to the risk assessment index output to remind downstream reviewers to refer to it with caution. The calculated service risk assessment index will be used as part of the professional execution results to trigger different levels of risk handling actions, such as ordinary attention, early warning intervention, or emergency transfer to manual intervention.Through the acoustic and textual cross-modal resonance model design in this embodiment, emotional and semantic cues in customer service calls can be jointly evaluated, avoiding hidden risks that may be overlooked by single-modal analysis, such as strong complaint intentions under a calm tone or simple inquiries under angry emotions.

[0037] In one optional embodiment, when the target business scenario is the intelligent identification sub-scenario of power hazard under the grassroots on-site operation inspection scenario, the power marketing professional model performs feature extraction and cross-modal alignment on the multimodal interaction request to generate a professional execution result for the specific business scenario, including: extracting the on-site inspection image and environmental description text from the multimodal interaction request; extracting pixel-level feature maps of the on-site inspection image using the visual encoder of the power marketing professional model; using a cross-modal cross-attention mechanism, with the word embedding vector of the environmental description text as the query condition, aggregating highly relevant target region features with a relevance greater than a preset relevance in the pixel-level feature map; inputting the aggregated target region features into a discriminator network, and outputting the classification label of the hazard category and the coordinate regression box surrounding the hazard category as the professional execution result.

[0038] Among them, on-site inspection images refer to original photos or video frames taken by grassroots workers at the user's site, showing the electrical equipment, wiring layout, meter box environment, etc.; environmental description text refers to supplementary explanatory text submitted by workers along with the images, which may include equipment location, preliminary judgment of abnormal phenomena, or verbal feedback from users; visual encoder refers to a deep convolutional network or visual transformer used to extract structured features from images, outputting multi-channel pixel-level feature maps; cross-modal cross-attention mechanism refers to a computational module that uses the representation of one modality as the query and the representation of another modality as the key and value, and aggregates information related to the query through attention weights; relevance refers to the similarity between the query vector and the feature vectors at various spatial locations in the feature map, which can be calculated using dot product or scaled dot product; discriminator network refers to a neural network that combines classification and localization, usually containing two output branches: the classification branch outputs the probability of the hazard category, and the regression branch outputs the bounding box coordinate parameters; the coordinate regression box of the hazard category refers to the location information of the rectangular area where the hazard is located in the original image, including the center point coordinates, width, and height.

[0039] In this embodiment, when the target business scenario is the intelligent identification of potential electricity hazards under the grassroots on-site operation inspection scenario, the power marketing professional model processes the multimodal interaction request as follows. First, the on-site inspection images and environmental description text are extracted from the multimodal interaction request. In one implementation, image fields and text fields are separated from the data payload of the multimodal request. For example, the image field may contain one or more on-site photos, and the text field may be a natural language sentence input by the operator through a mobile terminal. It should be noted that if the interaction request only contains images and does not provide environmental description text, an empty string can be automatically used as the default text, and the threshold of the cross-modal attention mechanism can be lowered in subsequent processing to run in pure visual mode. After extraction, the pixel-level feature map of the on-site inspection image is extracted using the visual encoder of the power marketing professional model. Then, using the cross-modal cross-attention mechanism, with the word embedding vector of the environmental description text as the query condition, highly relevant target region features with a relevance greater than a preset relevance are aggregated in the pixel-level feature map. In one implementation, the environmental description text is first segmented into words. A vector representation of each word is obtained through a pre-trained word embedding model. These word embedding vectors are averaged or subjected to self-attention pooling to obtain a global query vector. Then, the dot product similarity between this query vector and the feature vectors of each spatial location in the pixel-level feature map is calculated to obtain an attention heatmap. After soft-maximization normalization of the heatmap, the location features with a relevance greater than a preset threshold are weighted and summed to obtain the aggregated target region features. After obtaining the aggregated target region features, these features are input into a discriminator network, which outputs the classification label of the hazard category and the coordinate regression box surrounding the hazard category as the professional execution result. In one implementation, a single-stage detection network (such as an improved YOLO or SSD) is used to predict the category confidence and bounding box offset for each location on the full-map feature map. The hazard categories output by the discriminator network may include typical hazard types in the power marketing field, such as "damaged meter box," "unauthorized wiring," "aging lines," "missing seals," and "overheating marks." The coordinate regression bounding box represents the location of the hazard in the original image using relative coordinates (x_center, y_center, width, height). If the highest category confidence score output by the discriminator network is lower than a preset confidence threshold (e.g., 0.5), the result is not output directly. Instead, the image and text are packaged and pushed to the manual review queue, with a message displayed on the terminal: "Suspected hazard, but insufficient confidence, on-site verification required." Finally, the hazard category label and its coordinate regression bounding box are returned to the business terminal in a visual overlay (drawing colored rectangles and text labels on the original image) or in a structured data format.Through the cross-modal cross-attention mechanism in this embodiment, the environmental description text can guide the visual model to focus on the image area related to the text semantics, reducing background interference in the on-site inspection images, improving the pertinence of hazard detection, thereby realizing intelligent identification and location of electrical hazards, and assisting grassroots workers in quickly discovering and handling risks.

[0040] In one optional embodiment, when the target business scenario is the intelligent identification task sub-scenario of meter information under the grassroots on-site operation inspection scenario, the power marketing professional model performs feature extraction and cross-modal alignment on the multimodal interaction request to generate a professional execution result for the specific business scenario, including: extracting the on-site inspection image from the multimodal interaction request, locating the meter screen area and asset nameplate area in the on-site inspection image; using a character sequence recognition network to extract the meter reading value in the meter screen area and the asset number sequence in the asset nameplate area; retrieving the ledger benchmark value corresponding to the asset number sequence from the marketing business database; comparing the difference between the meter reading value and the ledger benchmark value, and if the difference exceeds the preset reasonable electricity consumption range, adding an abnormal warning mark to the professional execution result.

[0041] Among them, the character sequence recognition network refers to a deep learning model used to detect and recognize character sequences from image regions, which can be implemented based on connection-time classification or attention decoders; the marketing business database refers to the database in which power companies store business data such as user files, metering asset information, and historical electricity consumption; the ledger benchmark value refers to the reasonable reading of the last on-site meter reading or remote collection recorded in the database, combined with historical usage to estimate the current expected reading range.

[0042] In this embodiment, when the target business scenario is the intelligent identification task sub-scenario of meter information under the grassroots on-site operation inspection scenario, the power marketing professional model processes the multimodal interaction request as follows: First, extract the on-site inspection image from the multimodal interaction request, for example, directly obtain the meter photo taken by the operator from the data payload of the multimodal request. After obtaining the image, locate the meter screen area and asset nameplate area in the on-site inspection image. Next, use a character sequence recognition network to extract the meter reading value in the meter screen area and the asset number sequence in the asset nameplate area. Retrieve the ledger benchmark value corresponding to the asset number sequence from the marketing business database. For example, use the asset number sequence as the primary key to query the metering asset ledger table in the database to obtain the last recorded reading, recording date, user electricity category, and average daily electricity consumption of the meter. Calculate the current expected reading range based on the number of days from the last recording to the current date and the average daily electricity consumption. For example, if the previous reading was 1000, the average daily electricity consumption was 10 kWh, and the interval is 30 days, the expected reading would be 1300 kWh. The preset reasonable electricity consumption range can be set to 1300 ± 10% (i.e., 1170~1430). It should be noted that if the asset number does not exist in the database or the meter is newly installed with no historical readings, the ledger baseline value will be marked as "No Baseline," and the comparison will be skipped. The professional execution result will only output the identified reading and number, with the message "New meter installed, manual filing required." If the ledger baseline value can be determined, the difference between the meter reading and the ledger baseline value will be compared. If the difference exceeds the preset reasonable electricity consumption range, an abnormal warning will be added to the professional execution result. For example, the absolute value of the meter reading minus the ledger baseline value will be calculated. If this absolute value is greater than half the width of the reasonable range, it will be considered abnormal. In one implementation, the anomaly warning indicator includes three levels: minor anomaly (difference exceeding the range but within 20%), indicating "significant fluctuation in electricity consumption, recheck recommended"; moderate anomaly (20%~50%), indicating "suspected metering anomaly, please check meter wiring"; and severe anomaly (exceeding 50%), indicating "serious deviation, may involve electricity theft or meter malfunction, report immediately." When a meter reading is detected to be lower than the previous reading (i.e., running backwards), a special "meter running backwards" warning indicator can also be automatically generated. After comparison, the identified meter reading, asset number sequence, ledger baseline value, difference, and anomaly warning indicator (if any) are assembled into a structured professional execution result.

[0043] In one optional embodiment, during the process of generating the professional execution result by the power marketing professional model, the method further includes: The power marketing professional model generates business diagnostic text based on the input image through autoregression, and captures the two-dimensional cross-attention matrix activated at the bottom layer when generating each word in the process of generating the business diagnostic text based on the input image through autoregression. The global visual factual certainty is calculated using the visual anchoring focal convergence evaluation formula, which is as follows: ,in, This indicates the degree of certainty regarding the global visual facts; This represents the total number of valid tokens in the generated business diagnostic text. This indicates that the electricity marketing professional model generates the first... Original prediction confidence probability for each word element; Indicates the generation of the first A two-dimensional cross-attention matrix that associates input image regions with each word; This indicates the extraction of peak response weights from the two-dimensional cross-attention matrix; This represents the spatial information entropy calculation operation of the two-dimensional cross-attention matrix. The larger the entropy value, the more diffuse the attention. A preset smoothing constant is used to prevent overflow during division by zero; This indicates a cumulative multiplication operation.

[0044] Among them, business diagnostic text refers to the descriptive text automatically generated by the power marketing professional model based on input images (such as on-site inspection photos, meter screen images, and equipment anomaly images), which may include hazard identification, fault cause inference, and handling suggestions; autoregressive generation refers to the sequence generation method in which the model predicts the next word word one word word at a time, and the generation of each word word depends on all previously generated words words; the two-dimensional cross-attention matrix refers to the weight matrix generated by the interaction between the model decoder and the visual encoder when generating each word word, with its rows corresponding to the text word word positions and the columns corresponding to the spatial block positions of the input image, and each element in the matrix representing the intensity of attention to a certain image block when generating the current word word; global visual fact confidence refers to the comprehensive evaluation of the model's faithful dependence on visual information throughout the generation process, reflecting whether the generated content is based on image facts rather than the model's own memory or illusion; peak response weight refers to the maximum value in the attention matrix, representing the highest intensity of the model's attention to a certain image block; spatial information entropy refers to the uniformity of the distribution of attention weights in the image space, with a large entropy value indicating that attention is scattered across multiple blocks and lacks clear focus.

[0045] In this embodiment, during the generation of professional execution results by the power marketing professional model, specifically for scenarios requiring the output of business diagnostic text (such as a description of hazard diagnosis after on-site inspections at the grassroots level), the following operations are performed. First, the power marketing professional model generates business diagnostic text in an autoregressive manner based on the input image. In one implementation, an encoder-decoder architecture can be adopted, where the encoder processes the input image to extract visual features, and the decoder generates text word by word, with each decoding step conditional on the previously generated word sequence and image features. When generating the first word, the model uses a special starting marker as input; the generation of each subsequent word depends on all previously generated words. During the process of the power marketing professional model autoregressively generating business diagnostic text based on the input image, the two-dimensional cross-attention matrix activated at the bottom layer when generating each word is captured. Specifically, when the decoder performs cross-attention calculation at each layer, the weight tensor is extracted from the attention layer, and the cross-attention weight matrix of the last layer decoder is captured. The shape of this matrix is ​​(current generated word position, number of image blocks). Then, the global visual factual certainty is calculated using the visual anchoring convergence evaluation formula. The basic principle of this formula is: for each generated word element, the higher its original predicted confidence probability, and the more significant and spatially concentrated the peak of the cross-attention matrix, the higher the visual factual certainty of that word element; the global index is obtained by geometrically averaging the confidence probabilities of all word elements. Specifically, the original predicted confidence probability of the k-th word element is first extracted. (The probability value of the model's softmax output layer for the selected word); find the maximum value from the two-dimensional cross-attention matrix of that word. ; Calculate the spatial information entropy of the matrix That is, after normalizing the weights of all spatial locations in the matrix, the sum of the negative weighted logarithms is calculated; and Multiply the results and then divide by (spatial information entropy plus a small constant) to obtain the visual fact anchoring coefficient of the word; multiply the coefficients of all N valid words (excluding start and filler tags) together and then take the Nth root to obtain the global visual fact confidence level. .

[0046] Furthermore, the calculated global visual fact confidence level is also output as an additional field in the professional execution result. For example, when the confidence level is higher than a high threshold (e.g., 0.8), it is marked as "strong visual evidence," indicating that the diagnostic text is reliably based on image facts; when the confidence level is lower than a low threshold (e.g., 0.4), it is marked as "weak visual evidence," and manual review or re-acquiring a clearer image is recommended. This allows the model to quantitatively assess its reliance on image information while outputting business diagnostic text, effectively distinguishing between reasoning truly based on image facts and arbitrary generation without images.

[0047] Step 105: Feed back the professional execution results to the business terminal, collect the adoption behavior actions returned by the business terminal, and use the adoption behavior actions as feedback reinforcement signals to update the neuron weights of the power marketing professional model.

[0048] Among them, adoption behavior refers to the subsequent actions of business end users on the model output results, including direct adoption, partial modification and use, rejection or supplementary questions; feedback reinforcement signal refers to converting the user's adoption behavior into reward or penalty values, which are used to adjust the model parameters; neuron weight refers to the numerical parameters of the internal connection units of the power marketing professional model, and its update direction is guided by reinforcement signal.

[0049] In this embodiment, the professional execution results are fed back to the business terminal, and the adoption behavior actions returned by the business terminal are collected. In one implementation, explicit feedback buttons are set on the terminal interface: "Useful," "Useless," and "Needs Modification." Clicking these buttons records the user's behavior. In another implementation, implicit behavior collection is used: the degree of adoption is judged by monitoring the user's subsequent actions. For example, if the user copies the result content and pastes it into the work order system, it is considered indirect adoption; if the user immediately re-enters a different question, it is considered implicit rejection. The collected adoption behavior actions are used as feedback reinforcement signals to update the neuron weights of the power marketing professional model. In a specific implementation of the update method, policy gradient reinforcement learning can be used, treating the model as a policy network, with adoption / rejection as reward signals. Weights are updated through gradient ascent to increase the probability of high-reward outputs. Through this step of feedback collection and reinforcement updates, the model has the ability to continuously evolve, becoming increasingly aligned with actual business preferences over time.

[0050] In one optional embodiment, updating the neuron weights of the power marketing professional model using feedback reinforcement signals includes: freezing the original weight matrix of the pre-trained backbone layer in the power marketing professional model; injecting an update increment matrix consisting of the multiplication of two low-dimensional feature matrices into the bypass of the pre-trained backbone layer; calculating the gradient descent bias using the feedback reinforcement signal; iteratively correcting the parameters of the low-dimensional feature matrix only during backpropagation; and adding the corrected update increment matrix to the original weight matrix to complete the business-adaptive fine-tuning of the network model.

[0051] The pre-trained backbone layer refers to the core network layer in the power marketing professional model that has been pre-trained on large-scale general data. Its weight matrix carries the basic cross-modal understanding capability. For example, the backbone layer can include the following types: Transformer encoder layer in a text encoding network: This layer contains a multi-head self-attention module and a feedforward neural network module to extract high-dimensional semantic feature vectors from the text. Visual transformer layer or convolutional network backbone in a visual encoder: When a visual transformer is used, it contains a multi-head self-attention module and a feedforward neural network module; when a convolutional network is used, its deep convolutional blocks constitute the backbone feature extraction part. Temporal convolutional network layer or audio transformer layer in an acoustic encoding network: This layer also contains a self-attention or convolutional temporal module and its corresponding feedforward network. Attention layer containing the cross-modal cross-attention mechanism: This layer is used to align features from different modalities, and its core cross-attention weight matrix is ​​part of the backbone parameters. Decoder layer in text generation layer: When the power marketing professional model uses an encoder-decoder architecture to generate business diagnostic text or contracts, the self-attention module, cross-attention module, and feedforward neural network module in the decoder layer all belong to the backbone layer.

[0052] In this embodiment, the neuron weights of the electricity marketing professional model are updated using the collected feedback reinforcement signals (such as reward values ​​converted from user adoption, rejection, or modification behaviors), specifically employing a low-rank bypass incremental update strategy. First, the original weight matrix of the pre-trained backbone layer in the electricity marketing professional model is frozen. After freezing, an update increment matrix consisting of the product of two low-dimensional feature matrices is injected into the bypass of the pre-trained backbone layer. Specifically, for each weight matrix (of shape d×k, where d and k are the input and output dimensions, respectively) that needs to be adapted in the backbone layer, two low-dimensional matrices A (of shape d×r) and B (of shape r×k) are constructed, where r is much smaller than d and k (e.g., r=4, 8, or 16). The update increment matrix is ​​the product of A and B, with the same shape as the original weight matrix. Then, the gradient descent bias is calculated using the feedback reinforcement signals. In one implementation, the feedback reinforcement signal is treated as a reward value, and the loss function is calculated using a policy gradient method. For example, when the model's output is accepted by the user, the loss is the negative log-likelihood multiplied by the positive reward; when rejected, the loss is the negative log-likelihood multiplied by the negative reward. After calculating the gradient descent bias, the parameters of the low-dimensional feature matrices (i.e., matrices A and B) are iteratively corrected only during backpropagation; the parameters of the original weight matrix do not participate in gradient updates. Finally, the corrected update increment matrix is ​​added to the original weight matrix to complete the business-adaptive fine-tuning of the network model. During the inference phase, the weights actually used are the sum of the original weight matrix and the updated increment matrix. This ensures that when the power marketing professional model uses the feedback reinforcement signal for business-adaptive fine-tuning, it does not destroy the general cross-modal understanding ability learned in the pre-training phase, while reducing the number of parameters and computational overhead that need to be trained.

[0053] By applying the technical solution of this embodiment, a professional knowledge graph of power marketing is constructed, and external knowledge is injected and fine-tuned for training the multimodal large model, enabling the general large model to acquire professional responsiveness in the field of power marketing. Through the routing and distribution mechanism of intent recognition and scenario-based collaborative pipelines, differentiated processing of different task categories is achieved. By collecting user adoption behavior and using it as a reinforcing feedback signal to update model weights, the model is continuously optimized in actual use, gradually aligning with the operational habits and judgment preferences of business personnel. The synergistic effect of these methods constructs a demonstration application platform for a multimodal large model of power marketing that integrates multi-scenario functions, possesses professional depth, and online learning capabilities.

[0054] Furthermore, this application embodiment also provides a multi-scenario power marketing multimodal large-scale model demonstration application platform that integrates multiple scenarios, constructed through the above-described method for constructing a multi-scenario power marketing multimodal large-scale model demonstration application platform.

[0055] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0056] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for constructing a multi-modal large-scale demonstration application platform for power marketing that integrates multiple scenario functions, characterized in that: The method includes: Collect historical business data in the field of electricity marketing, preprocess the historical business data, and construct a professional knowledge map of electricity marketing that includes professional terms, business regulations and operating procedures based on the preprocessed historical business data; Obtain a basic multimodal large model, inject the power marketing professional knowledge graph as external incremental knowledge into the basic multimodal large model, and generate a power marketing professional model through fine-tuning training; Receive multimodal interaction requests sent by business terminals, use an intent recognition mechanism to parse the task category of the multimodal interaction requests, and route and distribute the multimodal interaction requests to the corresponding scenario collaborative work pipeline; In the corresponding scenario collaborative work pipeline, the power marketing professional model is invoked to perform feature extraction and cross-modal alignment on the multimodal interaction request, generating professional execution results for specific business scenarios; The professional execution results are fed back to the business terminal, and the adoption behavior actions returned by the business terminal are collected. The adoption behavior actions are used as feedback reinforcement signals to update the neuron weights of the power marketing professional model.

2. The method according to claim 1, characterized in that, The task category of the multimodal interaction request is parsed using an intent recognition mechanism, and the multimodal interaction request is routed and distributed to the corresponding scene collaboration pipeline, including: The context semantic vector and modal feature vector in the multimodal interaction request are extracted using a sequence encoder; The context semantic vector and the modal feature vector are concatenated and input into a fully connected classification network to calculate the matching probability distribution of the multimodal interaction request belonging to the regional digital marketing scenario, the customer service risk assessment scenario, and the grassroots on-site operation inspection scenario. The scenario with the largest value in the matching probability distribution is selected as the target business scenario. The scenario collaboration pipeline bound to the target business scenario is activated, and the data payload of the multimodal interaction request is loaded into the activated scenario collaboration pipeline.

3. The method according to claim 2, characterized in that, When the target business scenario is the regional digital marketing scenario, the power marketing professional model performs feature extraction and cross-modal alignment on the multimodal interaction request to generate professional execution results for the specific business scenario, including: Extract the scanned image of the document from the multimodal interaction request; use an optical character recognition algorithm to locate the character bounding box in the scanned image of the document and extract the character sequence; input the character sequence into the natural language sequence annotation layer in the power marketing professional model, and extract key-value pair data representing customer identity and property address number through joint encoding of spatial coordinates and semantic features, and assemble them into a structured business expansion work order attribute set; Obtain on-site survey parameters; retrieve the corresponding power supply scheme template and standard legal clauses from the power marketing professional knowledge graph; use the text generation layer in the power marketing professional model to fill the business expansion work order attribute set and the on-site survey parameters into the blank slots of the power supply scheme template and the standard legal clauses; perform compliance lexical scanning on the filled text content to generate the regional digital marketing contract text to be signed.

4. The method according to claim 2, characterized in that, When the target business scenario is the customer service risk assessment scenario, the power marketing professional model performs feature extraction and cross-modal alignment on the multimodal interaction request to generate professional execution results for the specific business scenario, including: High-dimensional acoustic feature vectors of customer service call audio in the multimodal interaction request are extracted using an acoustic coding network; A high-dimensional semantic feature vector of the transcribed text corresponding to the customer service call audio is extracted using a text encoding network. The service risk assessment index is calculated using a cross-modal bilinear projective resonance model, and its calculation formula is as follows: in, This indicates the service risk assessment index; This represents the high-dimensional acoustic feature vector; This represents the high-dimensional semantic feature vector; Represents the Hadamard product of vectors; and Let L1 and L2 norms of the vector be represented respectively. This represents a pre-trained cross-modal coupling weight matrix used to map acoustic space features to text semantic space; Represents an exponential function with the natural constant as its base; This represents the resonance surge factor.

5. The method according to claim 2, characterized in that, When the target business scenario is the intelligent identification sub-scenario of potential electricity hazards under the grassroots on-site operation inspection scenario, the power marketing professional model performs feature extraction and cross-modal alignment on the multimodal interaction request to generate professional execution results for the specific business scenario, including: Extract the on-site inspection image and environmental description text from the multimodal interaction request; extract pixel-level feature maps of the on-site inspection image using the visual encoder of the power marketing professional model; use a cross-modal cross-attention mechanism, with the word embedding vector of the environmental description text as the query condition, aggregate highly relevant target region features with a relevance greater than a preset relevance in the pixel-level feature map; input the aggregated target region features into a discriminator network, and output the classification label of the hazard category and the coordinate regression box surrounding the hazard category as the professional execution result.

6. The method according to claim 2, characterized in that, When the target business scenario is the intelligent identification task sub-scenario of meter information under the grassroots on-site operation inspection scenario, the power marketing professional model performs feature extraction and cross-modal alignment on the multimodal interaction request to generate professional execution results for the specific business scenario, including: Extract the on-site inspection image from the multimodal interaction request, and locate the meter screen area and asset nameplate area in the on-site inspection image; use a character sequence recognition network to extract the meter reading value in the meter screen area and the asset number sequence in the asset nameplate area; retrieve the ledger benchmark value corresponding to the asset number sequence from the marketing business database; compare the difference between the meter reading value and the ledger benchmark value, and if the difference exceeds the preset reasonable electricity consumption range, add an abnormal warning mark to the professional execution result.

7. The method according to any one of claims 1 to 6, characterized in that, In the process of generating the professional execution results by the power marketing professional model, the method further includes: The power marketing professional model generates business diagnostic text based on the input image through autoregression, and captures the two-dimensional cross-attention matrix activated at the bottom layer when generating each word in the process of generating the business diagnostic text based on the input image through autoregression. The global visual factual certainty is calculated using the visual anchoring focal convergence evaluation formula, which is as follows: ,in, This indicates the degree of certainty regarding the global visual facts; This represents the total number of valid tokens in the generated business diagnostic text. This indicates that the electricity marketing professional model generates the first... Original prediction confidence probability for each word element; Indicates the generation of the first A two-dimensional cross-attention matrix that associates input image regions with each word; This indicates the extraction of peak response weights from the two-dimensional cross-attention matrix; This represents the spatial information entropy calculation operation of the two-dimensional cross-attention matrix. The larger the entropy value, the more diffuse the attention. A preset smoothing constant is used to prevent overflow during division by zero; This indicates a cumulative multiplication operation.

8. The method according to any one of claims 1 to 6, characterized in that, The neuron weights of the electricity marketing expertise model are updated using feedback reinforcement signals, including: Freeze the original weight matrix of the pre-trained backbone layer in the power marketing professional model; inject an update increment matrix consisting of the multiplication of two low-dimensional feature matrices into the bypass of the pre-trained backbone layer; calculate the gradient descent bias using the feedback reinforcement signal; iteratively correct the parameters of the low-dimensional feature matrix only during backpropagation; add the corrected update increment matrix to the original weight matrix to complete the business adaptability fine-tuning of the network model.

9. The method according to claims 1 to 6, characterized in that, After parsing the task category of the multimodal interaction request using an intent recognition mechanism, the method further includes: Based on the task category of the multimodal interaction request obtained from the parsing, the corresponding role setting clause, task constraint clause, and security boundary clause are extracted from the preset instruction template library; the role setting clause, the task constraint clause, and the security boundary clause are concatenated into a system-level guidance prompt; the system-level guidance prompt is forcibly placed in the header input sequence of the multimodal interaction request to limit the output style and business boundaries of the power marketing professional model.

10. A multi-modal large-scale demonstration application platform for power marketing integrating multiple scenario functions, characterized in that, This platform is constructed using the method described in any one of claims 1 to 9 for building a multi-modal large-scale demonstration application platform for integrated multi-scenario functions in power marketing.