Intelligent operation method and system, electronic equipment and storage medium

By converting user voice information into structured data and establishing a large language model knowledge base, building a two-layer LSTM network and detection model, we have solved the problem of low efficiency of traditional customer service, realized automatic generation of work orders and efficient retrieval of the knowledge base, improved the real-time performance of inspections and work efficiency, and reduced human supervision costs.

CN120655244APending Publication Date: 2025-09-16JIANGXI SHUITOUJIANG INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510981729.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Traditional customer service is inefficient, lacks intelligent information extraction capabilities, has delayed knowledge base updates, high prediction errors for complex time series data, insufficient multi-source heterogeneous sensor data fusion capabilities, poor real-time performance, and difficulty tracing violations. Furthermore, traditional customer service scenarios lack intelligent information extraction capabilities, resulting in low work efficiency and the inability to achieve semantic association and dynamic knowledge reasoning.

Method used

By acquiring user voice information and converting it into structured data, establishing a large language model knowledge base, building a two-layer LSTM network for early warning, training the detection model and associating the detection results with the knowledge base, using the CleanS2S and FunASR models to process voice information, using NebulaGraph to build a knowledge graph for retrieval, and combining the Milvus database and GAN-generated data for multi-source data fusion.

Benefits of technology

It realizes the automatic generation of work orders, shortens response time, improves work efficiency, improves the matching accuracy of knowledge retrieval and the real-time performance of inspections, reduces human supervision costs, and reduces manual inspection costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655244A_ABST
    Figure CN120655244A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent operation method and system, electronic equipment and a storage medium. The method comprises the steps that user voice information is acquired, and a voice text is processed to output structured data; identifying a user request type and an asset object in the structured data, and performing context completion on the structured data and reporting the structured data; establishing a knowledge base, converting the knowledge document into structured knowledge data, and importing the structured knowledge data; performing Boolean matching management detection retrieval on the structured knowledge data according to a query demand of a user; early warning is carried out based on a double-layer LSTM network and fusion input features; and training the detection model, detecting the field image based on the trained detection model to obtain a detection result, and associating the detection result with the knowledge base. According to the invention, the response time of condition disposal can be shortened, the manpower supervision cost can be reduced, and the working efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent operation method, system, electronic device and storage medium. Background Art

[0002] The shortcomings of traditional customer service in water conservancy projects, asset operation and maintenance, and customer service scenarios include: manual recording of work orders takes an average of 15-20 minutes and has a high error rate; there is a lack of intelligent information extraction capabilities, requiring manual secondary confirmation of asset codes; and the lack of customer emotion recognition leads to blind spots in the formulation of follow-up strategies.

[0003] Limitations of current knowledge bases include: long average response times for document retrieval; low accuracy in parsing unstructured documents such as PDFs; knowledge update cycles of up to 2-3 weeks, lagging behind actual operational needs; and document retrieval relying on keyword matching, which fails to achieve semantic association and dynamic knowledge reasoning.

[0004] In summary, among existing technologies, traditional statistical operation models have high prediction errors for complex time series data, manual inspections have a high rate of missed detections, traditional models such as ARIMA are insufficient in the ability to fuse multi-source heterogeneous sensor data, resulting in delayed warnings, low manual inspection efficiency, poor real-time performance, and difficulty in tracing violations. In addition, they lack intelligent information extraction capabilities in traditional customer service scenarios, resulting in low work efficiency and the inability to achieve semantic association and dynamic knowledge reasoning. Single-modal technology is difficult to support end-to-end closed-loop management of complex operation scenarios. Summary of the Invention

[0005] Based on this, the purpose of the present invention is to provide a smart operation method, system, electronic device and storage medium to solve the deficiencies in the above-mentioned prior art.

[0006] In a first aspect, the present invention provides a smart operation method, the method comprising:

[0007] Acquiring user voice information, converting the voice information into voice text, and processing the voice text to output structured data;

[0008] Identify the user request type and asset object in the structured data, complete the context of the structured data, and report it;

[0009] Building a knowledge base based on a large language model, converting knowledge documents into structured knowledge data, and importing the structured knowledge data into the knowledge base;

[0010] Performing Boolean matching detection and retrieval on the structured knowledge data according to user query requirements;

[0011] Constructing a two-layer LSTM network and fusing input features to generate an early warning based on the two-layer LSTM network and the fused input features;

[0012] A detection model is trained, and on-site image information is collected. The on-site image is detected based on the trained detection model to obtain a detection result, and the detection result is associated with the knowledge base.

[0013] Compared with the existing technology, the beneficial effects of the present invention are: by contextually completing and reporting structured data, it is possible to automatically generate and report work orders, shorten response time, and improve work efficiency. By establishing a knowledge base through a large language model and detecting and retrieving through Boolean matching, it can support downstream knowledge retrieval and reasoning, improve matching accuracy, and improve work order processing efficiency through multi-source data fusion. Early warning through a double-layer LSTM network not only improves the real-time performance of inspections, but also effectively shortens processing response time. It detects on-site information through a detection model and associates the detection results with a knowledge base to reduce human supervision costs, shortening the time from violation handling to response, and reducing manual inspection costs.

[0014] Furthermore, the step of converting the voice information into voice text and processing the voice text to output structured data includes:

[0015] Using the CleanS2S model to filter background noise in the voice information, and using FunASR to convert the voice information into speech text;

[0016] The speech text is jointly parsed based on a large language model to output the structured data.

[0017] Furthermore, the step of identifying the user request type and asset object in the structured data, and performing context completion on the structured data and reporting the structured data includes:

[0018] Performing intent recognition on the structured data based on a large language model to classify user request types and asset objects;

[0019] populating an event field in the structured data based on the conversation history;

[0020] The emotional features in the voice information are extracted and marked using the SenseVoice model.

[0021] Furthermore, the step of importing the structured knowledge data into the knowledge base includes:

[0022] Obtain multi-format knowledge documents based on operation management systems, solutions to common problems, station operation and maintenance, and water conservancy project inspection and oxidation manuals;

[0023] Converting the reading mode of the multi-format knowledge document to obtain a Markdown format document, processing the Markdown format document using the open source framework Easy Dataset, and generating a fine-tuning dataset with a thought chain based on the processed Markdown format document;

[0024] Sequentially performing chapter segmentation, dynamic recursive splitting, and metadata extraction on the processed Markdown format document to obtain segmented text blocks and structured metadata respectively;

[0025] The structured metadata is imported into the knowledge base, and the segmented text blocks are stored in the Milvus database.

[0026] Furthermore, the step of performing Boolean matching detection and retrieval on the structured metadata according to the user query requirement includes:

[0027] Define entity relationships based on NebulaGraph and associate the structured documents with real-time data to build a knowledge graph;

[0028] Based on the knowledge graph, NebulaGraph is used to search for the entity relationships and disposal rules for graph retrieval, and Boolean matching is performed on the structured knowledge data for keyword retrieval to obtain retrieval results;

[0029] The search results are scored for relevance according to the recall results and sorted by weight.

[0030] Furthermore, the step of training the detection model includes:

[0031] Build a basic training set based on COCO-Safety or Safety-Helmet-Wearing-Dataset;

[0032] Capture live images in real time, annotate and dynamically enhance them in turn, and generate synthetic data based on GAN

[0033] Embed the CBAM attention mechanism in the YOLOv8 backbone network Backbone, and perform the first phase training on the YOLOv8 backbone network using the training set;

[0034] The YOLOv8 backbone network is trained in the second stage based on the synthetic data to obtain a detection model.

[0035] Furthermore, the step of associating the detection result with the knowledge base includes:

[0036] Associating the detection results with the safety construction specifications in the knowledge base to generate treatment recommendations;

[0037] The illegal data in the detection result is stored in the atlas database.

[0038] In a second aspect, the present invention further provides a smart operation system, comprising:

[0039] An acquisition module, configured to acquire user voice information, convert the voice information into voice text, and process the voice text to output structured data;

[0040] an identification module, configured to identify the user request type and asset object in the structured data, and perform contextual completion on the structured data and report the result;

[0041] A building module, configured to build a knowledge base based on a large language model, convert knowledge documents into structured knowledge data, and import the structured knowledge data into the knowledge base;

[0042] A retrieval module, configured to perform Boolean matching and detection retrieval on the structured knowledge data according to user query requirements;

[0043] A construction module is used to construct a two-layer LSTM network and fuse input features to generate an early warning based on the two-layer LSTM network and the fused input features;

[0044] The detection module is used to train the detection model and collect scene image information, detect the scene image based on the trained detection model to obtain the detection result, and associate the detection result with the knowledge base.

[0045] In a third aspect, the present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned intelligent operation method when executing the computer program.

[0046] In a fourth aspect, the present invention also provides a storage medium on which a computer program is stored, which implements the above-mentioned intelligent operation method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 Flowchart of the smart operation method in the first embodiment of the present invention;

[0048] Figure 2 A schematic diagram of the construction of a knowledge base and data storage in the first embodiment of the present invention;

[0049] Figure 3is a structural block diagram of the smart operation system in the second embodiment of the present invention;

[0050] Figure 4 FIG. 4 is a schematic structural diagram of an electronic device in a third embodiment of the present invention.

[0051] Description of main component symbols:

[0052] 10. Acquisition module; 20. Recognition module; 30. Establishment module; 40. Retrieval module; 50. Construction module; 60. Detection module;

[0053] 70. Bus; 71. Processor; 72. Memory; 73. Communication interface.

[0054] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION

[0055] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. The drawings illustrate several embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present invention.

[0056] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be an intermediate element. When an element is referred to as being "connected to" another element, it may be directly connected to the other element or there may be an intermediate element. The terms "vertical," "horizontal," "left," "right," and similar expressions used herein are for illustrative purposes only.

[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0058] Example 1

[0059] See also Figure 1 , which shows the smart operation method in the first embodiment of the present invention, and includes steps S1 to S6:

[0060] S1, obtaining user voice information, converting the voice information into voice text, and processing the voice text to output structured data;

[0061] Specifically, the step S1 includes steps S11 to S12:

[0062] S11, using the CleanS2S model to filter background noise in the voice information, and using FunASR to convert the voice information into speech text;

[0063] S12, jointly parsing the speech text based on the large language model to output the structured data;

[0064] It is understandable that the speech-to-text capability converts the content of communication between customer service and customers into text and then uses the deepseek large language model to perform multiple rounds of task joint analysis, and performs intent recognition, asset recognition (station assets, water conservancy project assets, software assets, hardware assets), and customer information recognition on these texts. At the same time, natural language understanding (NLU) and natural language processing (NLP) are combined with the context to summarize the chat content for event registration, pre-selection of assets, pre-recording of customer information, and pre-recording of event descriptions. After the customer or customer service hangs up the phone, the SenseVoice model is used to recognize the emotions of the chat voice so that users can provide high-quality service when they return.

[0065] S2, identifying the user request type and asset object in the structured data, performing context completion on the structured data, and reporting the structured data;

[0066] Specifically, step S2 includes steps S21 to S23:

[0067] S21, performing intent recognition on the structured data based on a large language model to classify the user request type and the asset object;

[0068] S22, filling the event field in the structured data based on the conversation history;

[0069] S23, extracting emotional features from the voice information using a SenseVoice model and marking them;

[0070] It can be understood that intent recognition: classifies user request types, such as fault reporting and consultation questions; entity extraction: identifies asset objects, such as the name of the measuring station asset, the location of the water conservancy project, and the customer contact information; context completion: automatically fills in the event report fields based on the conversation history, such as event classification, description, customer contact information, asset object, etc.

[0071] It's worth noting that after the call, the SenseVoice model extracts emotional features (pitch, speaking rate, and pause frequency) from the voice and generates an emotion tag (happy / angry) to mark high-priority return calls. This improves the efficiency of automatically generating incident tickets, enhances the completeness of information entered, enhances emotion recognition, and shortens response times for high-risk customers.

[0072] S3, establishing a knowledge base based on the large language model, converting the knowledge documents into structured knowledge data, and importing the structured knowledge data into the knowledge base;

[0073] Specifically, step S3 includes steps S31 to S34:

[0074] S31, obtains multi-format knowledge documents based on operation management systems, solutions to common problems, station operation and maintenance, and water conservancy project inspection and oxidation manuals;

[0075] S32, converting the reading mode of the multi-format knowledge document to obtain a Markdown format document, processing the Markdown format document using the open source framework Easy Dataset, and generating a fine-tuning dataset with a thought chain based on the processed Markdown format document;

[0076] S33, sequentially performing chapter segmentation, dynamic recursive splitting, and metadata extraction on the processed Markdown format document to obtain segmented text blocks and structured metadata respectively;

[0077] S34, importing the structured metadata into the knowledge base, and storing the segmented text blocks into the Milvus database;

[0078] It is understandable that the knowledge base is built based on the large language model technology. The knowledge base imports the operation manuals of various application systems, the management system of the operation department, the solution to common problems, the operation and maintenance of the measuring station, the inspection and maintenance standard manual of the water conservancy project, etc., so that customers can solve problems conveniently and quickly. When new employees of the operation department join the company, they can quickly find the operation and maintenance standards and reduce the training cost for new employees. The operation knowledge base adopts the following technologies for construction. The open source project EASYDATASET selects the 70b deepseek deployed by ollam as the model for processing documents in PDF, word, excel and other formats and generating fine-tuning data sets. The data set and the documents divided by the enhanced recursive block method are vectorized by the vector model bge-m3 and stored in the milvus vector database. For details, please refer to Figure 2 ;

[0079] It should be noted that the PDF / Word / Excel / Text format is converted into Markdown format suitable for model reading. Then the open source framework Easy Dataset is used to parse the processed Markdown format documents, and the DeepSeek-R1-70B model is used to generate a fine-tuning dataset with thought chain.

[0080] Chapter-priority segmentation: Automatically identify document chapter structure based on Markdown titles and segment text blocks according to semantic integrity;

[0081] Dynamic recursive splitting: Recursively splits very long paragraphs (>1000 characters) into chunks, ensuring paragraph integrity and a character count between 200 and 1000. There will be a 50-character overlap between adjacent chunks, ensuring that important information is preserved even if it falls on a chunk boundary.

[0082] Metadata extraction: Synchronously extract document outlines and text block summaries (such as "Station equipment maintenance cycle: once a quarter") to generate structured metadata;

[0083] Vectorized storage: Text blocks are vectorized by the BGE-M3 model and stored in the Milvus database, supporting high-dimensional similarity retrieval (cosine similarity > 0.85).

[0084] It should be explained that the semantic integrity of document segmentation is improved by 90%, the processing speed reaches 200 pages / minute, the metadata extraction accuracy is ≥95%, and downstream knowledge retrieval and reasoning are supported.

[0085] S4, performing Boolean matching detection and retrieval on the structured knowledge data according to the user's query requirements;

[0086] Specifically, the step S4 includes steps S41 to S43:

[0087] S41, defining entity relationships based on NebulaGraph, and associating the structured documents with real-time data to construct a knowledge graph;

[0088] S42, searching the entity relationships and processing rules based on the knowledge graph through NebulaGraph to perform graph retrieval, and performing Boolean matching on the structured knowledge data to perform keyword retrieval to obtain retrieval results;

[0089] S43, scoring the relevance of the search results according to the recall results, and sorting them by weight;

[0090] It is understandable that the knowledge graph is constructed by defining entity relationships based on NebulaGraph (such as "water pump equipment-operation and maintenance standards-maintenance cycle"), and associating structured documents with real-time IoT data (water level, displacement). Among them, the rule triggering mechanism: when sensor data is abnormal, such as displacement >15mm, the disposal rules in the knowledge graph are automatically matched to determine whether to "start the slope reinforcement plan"; multi-source data fusion: combining historical work orders and equipment manuals to generate dynamic disposal suggestions, such as "the current rainfall is 50mm / hour, it is recommended to open the flood discharge gate." The response time of knowledge graph reasoning is shortened, the matching accuracy of disposal suggestions is improved, and the fusion of multi-source data improves the efficiency of work order processing. The response time of knowledge graph reasoning is shortened, the matching accuracy of disposal suggestions is improved, and the fusion of multi-source data improves the efficiency of work order processing.

[0091] It should be explained that the hybrid retrieval mechanism, when a user queries, simultaneously executes the following: recalling the top 10 vector retrievals of similar text blocks from Milvus; performing graph retrieval by searching associated entities and disposal rules through NebulaGraph; performing Boolean matching on document metadata (abstracts, chapter titles) for keyword retrieval; and scoring the relevance of the recall results based on DeepSeek-R1-70B, sorting them by weight (vector similarity × 0.6 + graph relevance × 0.3 + keyword matching × 0.1).

[0092] In addition, high-frequency query results are cached in Redis to reduce database load, and the vector library and graph are incrementally updated daily at dawn to ensure knowledge currency. RAG recall accuracy is improved, response time is shortened, and multimodal retrieval increases the resolution rate of complex problems. A cross-modal correlation solution for "water pump leaks and rainfall surges" is also available. This allows for precise segmentation: an enhanced recursive segmentation algorithm avoids chapter truncation and adapts to LLM input restrictions; intelligent storage: the vector library and knowledge graph complement each other, taking into account semantic similarity and logical relevance; and efficient recall: multi-path retrieval and reordering address the redundancy and missed detection issues of traditional RAG recall.

[0093] S5, constructing a two-layer LSTM network and fusing input features to generate an early warning based on the two-layer LSTM network and the fused input features;

[0094] It is understandable that the LSTM (long short-term memory) algorithm model is used to improve the intelligence level of data processing, solve the limitations of traditional methods in time series prediction, anomaly detection, and multi-source data fusion, and use water level, temperature, reservoir rainfall, and geological sensor data as time series input to predict deformation trends in the next 1-7 days. If the predicted value exceeds the safety threshold (such as displacement >20mm), a graded warning will be triggered to avoid the risk of dam collapse.

[0095] It is important to explain that the input features include: water level (detected by ultrasonic sensors), temperature (detected by infrared sensors), rainfall (detected by weather station API), and geological displacement (detected by GNSS monitoring), with a sampling frequency of 1 time per minute. The model architecture uses a two-layer LSTM network with a hidden layer dimension of 128. Noisy data is filtered using a forget gate / input gate mechanism. Dynamic thresholds are used to calculate safety thresholds based on rolling historical data (for example, the displacement threshold is dynamically adjusted to 3σ of the historical mean).

[0096] This enables tiered early warnings, predicting trends for the next 1-7 days and outputting warning levels: Yellow (displacement 10-15mm): notifying operations and maintenance personnel for on-site verification; Red (displacement >20mm): triggering SMS notifications and initiating emergency response procedures. The mean absolute error of predictions has significantly decreased, the rate of missed warnings has been reduced, and the accuracy of dam failure risk identification has been improved.

[0097] S6, training a detection model, collecting on-site image information, detecting the on-site image based on the trained detection model to obtain a detection result, and associating the detection result with the knowledge base;

[0098] Specifically, step S6 includes steps S61 to S66:

[0099] S61, build a basic training set based on COCO-Safety or Safety-Helmet-Wearing-Dataset;

[0100] S62, capturing live images in real time, annotating and dynamically enhancing the live images in sequence, and generating synthetic data based on GAN

[0101] S63, embedding the CBAM attention mechanism in the YOLOv8 backbone network Backbone, and performing the first phase training on the YOLOv8 backbone network using the training set;

[0102] S64, performing a second phase of training on the YOLOv8 backbone network based on the synthetic data to obtain a detection model;

[0103] S65, associating the detection result with the safety construction specifications in the knowledge base to generate a disposal suggestion;

[0104] S66, storing the illegal data in the detection result into a graph database;

[0105] It is understandable that the YOLOv8 model is deployed to detect helmet wearing, construction compliance, equipment illegal stacking, and protective clothing detection through cameras; output violation event labels (time, location, and violation type);

[0106] It should be noted that the basic training set is built based on free commercial datasets such as COCO-Safety and Safety-Helmet-Wearing-Dataset, covering common targets such as safety helmets, protective clothing, and construction scenes. The camera is connected to the video cloud platform to capture images in real time, and a semi-automatic annotation tool (CVAT) is used to annotate specific scene targets such as "illegal stacking of equipment" and "missing foundation pit fences." The hard example mining strategy is supported to optimize the model's generalization capabilities. To address the lighting changes and occlusion issues in water conservancy project scenes, a dynamic enhancement strategy is adopted: adding rain and fog, and low-light simulation at night to simulate a noisy environment; random cropping and rotational spatial enhancement; and generating synthetic data based on GAN for rare violations, such as high-altitude work without a safety rope.

[0107] Furthermore, it's worth noting that the CBAM attention mechanism is embedded in the YOLOv8 backbone network to improve detection of small objects, such as protective gloves. To address class imbalance, where "compliant behavior" samples far outnumber "violation" samples, a joint optimization method called Focal Loss and IoU-Aware Loss is designed to focus on difficult examples and enhance localization accuracy. A two-stage training approach is employed as a transfer learning strategy: first, a base model is trained on open-source datasets; second, annotated field data collected from the video cloud is used as incremental data for domain adaptation training. The learning rate is reduced to 1e-5, freezing shallow network parameters.

[0108] During the inspection, safety equipment inspections are carried out: safety helmets, protective clothing, reflective clothing, and insulating gloves; construction compliance inspections are carried out: equipment stacking distances less than 1 meter from safety passages are considered violations, enclosure integrity, and the wearing of safety ropes for high-altitude operations; environmental risk inspections: no warning signs are set up in waterlogged areas, and flammable materials are stored illegally.

[0109] Using TensorRT quantization, the model was deployed to the NVIDIA Jetson AGX Orin edge server at the construction site, enabling real-time detection of video streams with a required frame rate of 30 FPS or higher. Through TensorRT quantization and model pruning, the YOLOv8 model was compressed to less than 40MB, supporting deployment on low-computing edge devices. A tiered alerting strategy was designed, with local processing of level 1 alerts and cloud-based decision-making for level 2 alerts, reducing bandwidth reliance.

[0110] Upload the screenshot, timestamp, and location coordinates (gis geographic location) of the violation to the smart operation platform, trigger an alarm work order and push it to the responsible person's mobile phone APP (Gan Shuitong).

[0111] A multi-level alarm mechanism is set up on the smart operation platform. With the support of the basic capabilities of the video cloud platform, when it is found that the safety helmet is not worn, an audible and visual alarm will be issued on site, and the violator will be notified by SMS, and a first-level alarm event will be pushed to the operation platform; when it is found that the other party of the equipment is blocking the escape route, the construction permit will be automatically frozen, and it can only be lifted after review by the safety officer, and a second-level alarm event will be generated and pushed to the operation platform.

[0112] The detection results are linked to the "Safety Construction Specifications" in the knowledge base, and disposal suggestions are automatically generated, such as "Not wearing a safety helmet → stop work immediately, fined 200 yuan." Historical violation data is stored in the atlas database to support risk profiling analysis, such as "XX construction team's monthly violation rate > 10% → mandatory safety training." Manual review is automatically triggered for low-confidence samples in the model (confidence <0.6), and after annotation, they are added to the training set for iterative optimization. The detection threshold is dynamically adjusted according to meteorological data (rainfall, strong winds), such as enhancing the detection sensitivity of waterlogged areas on rainy days. This can effectively reduce the incidence of safety accidents, reduce supervision labor costs, shorten the time from violation identification to disposal response, and reduce the cost of manual inspections.

[0113] In summary, the intelligent operation method in the above embodiment of the present invention can realize automatic generation and reporting of work orders by contextually completing and reporting structured data, shortening response time, and improving work efficiency. It establishes a knowledge base through a large language model, and supports downstream knowledge retrieval and reasoning through Boolean matching tube detection and retrieval, improves matching accuracy, and improves work order processing efficiency through multi-source data fusion. Early warning through a two-layer LSTM network not only improves the real-time inspection, but also can effectively shorten the processing response time. It detects on-site information through a detection model, and associates the detection results with the knowledge base to reduce human supervision costs, shortens the time from violation handling to response, and reduces manual inspection costs.

[0114] Example 2

[0115] The second embodiment of the present invention also provides a smart operation system, please refer to Figure 3 , which shows a smart operation system in a second embodiment of the present invention, the system includes:

[0116] An acquisition module 10 is configured to acquire user voice information, convert the voice information into voice text, and process the voice text to output structured data;

[0117] an identification module 20 for identifying the user request type and asset object in the structured data, and performing context completion on the structured data and reporting the same;

[0118] A building module 30 is used to build a knowledge base based on the large language model, convert knowledge documents into structured knowledge data, and import the structured knowledge data into the knowledge base;

[0119] A retrieval module 40 is configured to perform a Boolean matching test on the structured knowledge data according to a user query requirement;

[0120] A construction module 50 is configured to construct a two-layer LSTM network and fuse input features to generate an early warning based on the two-layer LSTM network and the fused input features;

[0121] The detection module 60 is used to train the detection model and collect scene image information, detect the scene image based on the trained detection model to obtain a detection result, and associate the detection result with the knowledge base.

[0122] In some optional embodiments, the acquisition module 10 includes:

[0123] A first conversion unit, configured to filter background noise in the voice information using a CleanS2S model, and convert the voice information into speech text using FunASR;

[0124] A parsing unit is used to jointly parse the speech text based on a large language model to output the structured data.

[0125] In some optional embodiments, the identification module 20 includes:

[0126] an identification unit, configured to perform intent recognition on the structured data based on a large language model to classify user request types and asset objects;

[0127] A filling unit, configured to fill an event field in the structured data based on the conversation history;

[0128] The extraction unit is used to extract the emotional features in the voice information through the SenseVoice model and mark them.

[0129] In some optional embodiments, the establishing module 30 includes:

[0130] An acquisition unit is used to obtain multi-format knowledge documents based on operation management systems, solutions to common problems, station operation and maintenance, and water conservancy project inspection and oxidation manuals;

[0131] a second conversion unit, configured to convert the reading mode of the multi-format knowledge document to obtain a Markdown format document, process the Markdown format document using the open source framework Easy Dataset, and generate a fine-tuning dataset with a thought chain based on the processed Markdown format document;

[0132] A segmentation unit, configured to sequentially perform chapter segmentation, dynamic recursive splitting, and metadata extraction on the processed Markdown format document to obtain segmented text blocks and structured metadata respectively;

[0133] The import unit is used to import the structured metadata into the knowledge base and store the segmented text blocks into the Milvus database.

[0134] In some optional embodiments, the retrieval module 40 includes:

[0135] A definition unit, configured to define entity relationships based on NebulaGraph and associate the structured documents with real-time data to construct a knowledge graph;

[0136] A search unit, configured to search the entity relationships and processing rules based on the knowledge graph through NebulaGraph to perform graph search, and perform Boolean matching on the structured knowledge data to perform keyword search to obtain search results;

[0137] The sorting unit is used to score the relevance of the retrieval results according to the recall results and sort them according to the weights.

[0138] In some optional embodiments, the detection module 60 includes:

[0139] Construction unit, used to build a basic training set based on COCO-Safety or Safety-Helmet-Wearing-Dataset;

[0140] The capture unit is used to capture the scene image in real time, annotate and dynamically enhance the scene image in turn, and generate synthetic data based on GAN

[0141] The first training unit is used to embed the CBAM attention mechanism in the YOLOv8 backbone network Backbone and perform the first phase training on the YOLOv8 backbone network using the training set;

[0142] A second training unit is used to perform a second phase of training on the YOLOv8 backbone network based on the synthetic data to obtain a detection model;

[0143] a generating unit, configured to associate the detection result with the safety construction specifications in the knowledge base to generate a disposal suggestion;

[0144] The storage unit is used to store the illegal data in the detection result into the atlas database.

[0145] The functions or operation steps implemented when the above modules and units are executed are substantially the same as those in the above method embodiments and will not be repeated here.

[0146] The smart operation system provided in the embodiment of the present invention has the same implementation principle and technical effects as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the system embodiment, reference can be made to the corresponding content in the aforementioned method embodiment.

[0147] Example 3

[0148] The present invention also provides an electronic device, see Figure 4 , shown is an electronic device in a third embodiment of the present invention.

[0149] The electronic device may include a processor 71 and a memory 72 storing computer program instructions.

[0150] Specifically, the processor 71 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the present application.

[0151] Among them, the memory 72 may include a large-capacity memory for data or instructions. By way of example and not limitation, the memory 72 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 72 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 72 may be inside or outside the data processing device. In a specific embodiment, the memory 72 is a non-volatile memory. In a specific embodiment, the memory 72 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. Under appropriate circumstances, the RAM can be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM can be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0152] The memory 72 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 71 .

[0153] The processor 71 implements the smart operation method of the above-mentioned embodiment 1 by reading and executing computer program instructions stored in the memory 72.

[0154] In some embodiments, the electronic device may further include a communication interface 73 and a bus 70. Figure 3 As shown, the processor 71, the memory 72, and the communication interface 73 are connected via a bus 70 and communicate with each other.

[0155] The communication interface 73 is used to implement communication between the various modules, devices, units and / or equipment in this application. The communication interface 73 can also implement data communication with other components such as: external devices, image / data acquisition equipment, databases, external storage, and image / data processing workstations.

[0156] The bus 70 includes hardware, software, or both, and couples the components of the device to each other. The bus 70 includes, but is not limited to, at least one of the following: a data bus, an address bus, a control bus, an expansion bus, and a local bus. By way of example and not limitation, bus 70 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of the above. Bus 70 may include one or more buses, where appropriate. Although this application describes and illustrates a particular bus, this application contemplates any suitable bus or interconnect.

[0157] The electronic device can obtain the smart operation system and execute the smart operation method of the first embodiment.

[0158] In addition, in conjunction with the smart operation method in the first embodiment, the present application may provide a storage medium for implementation. The storage medium stores computer program instructions; when the computer program instructions are executed by a processor, the smart operation method in the first embodiment is implemented.

[0159] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0160] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A smart operation method, characterized in that: The method comprises: Acquiring user voice information, converting the voice information into voice text, and processing the voice text to output structured data; Identify the user request type and asset object in the structured data, complete the context of the structured data, and report it; Building a knowledge base based on a large language model, converting knowledge documents into structured knowledge data, and importing the structured knowledge data into the knowledge base; Performing Boolean matching detection and retrieval on the structured knowledge data according to user query requirements; Constructing a two-layer LSTM network and fusing input features to generate an early warning based on the two-layer LSTM network and the fused input features; A detection model is trained, and on-site image information is collected. The on-site image is detected based on the trained detection model to obtain a detection result, and the detection result is associated with the knowledge base.

2. The smart operation method according to claim 1, characterized in that: The steps of converting the voice information into voice text and processing the voice text to output structured data include: Using the CleanS2S model to filter background noise in the voice information, and using FunASR to convert the voice information into speech text; The speech text is jointly parsed based on a large language model to output the structured data.

3. The smart operation method according to claim 1, characterized in that: The steps of identifying the user request type and asset object in the structured data, performing context completion on the structured data, and reporting the structured data include: Performing intent recognition on the structured data based on a large language model to classify user request types and asset objects; populating an event field in the structured data based on the conversation history; The emotional features in the voice information are extracted and marked using the SenseVoice model.

4. The smart operation method according to claim 1, characterized in that: The step of importing the structured knowledge data into the knowledge base includes: Obtain multi-format knowledge documents based on operation management systems, solutions to common problems, station operation and maintenance, and water conservancy project inspection and oxidation manuals; Converting the reading mode of the multi-format knowledge document to obtain a Markdown format document, processing the Markdown format document using the open source framework EasyDataset, and generating a fine-tuning dataset with a thought chain based on the processed Markdown format document; Sequentially performing chapter segmentation, dynamic recursive splitting, and metadata extraction on the processed Markdown format document to obtain segmented text blocks and structured metadata respectively; The structured metadata is imported into the knowledge base, and the segmented text blocks are stored in the Milvus database.

5. The smart operation method according to claim 1, characterized in that: The step of performing Boolean matching detection and retrieval on the structured metadata according to the user query requirement includes: Define entity relationships based on NebulaGraph and associate the structured documents with real-time data to build a knowledge graph; Based on the knowledge graph, NebulaGraph is used to search for the entity relationships and disposal rules for graph retrieval, and Boolean matching is performed on the structured knowledge data for keyword retrieval to obtain retrieval results; The search results are scored for relevance according to the recall results and sorted by weight.

6. The smart operation method according to claim 1, characterized in that: The steps of training the detection model include: Build a basic training set based on COCO-Safety or Safety-Helmet-Wearing-Dataset; Capture live images in real time, annotate and dynamically enhance them in turn, and generate synthetic data based on GAN Embed the CBAM attention mechanism in the YOLOv8 backbone network Backbone, and perform the first phase training on the YOLOv8 backbone network using the training set; The YOLOv8 backbone network is trained in the second stage based on the synthetic data to obtain a detection model.

7. The smart operation method according to claim 1, characterized in that: The step of associating the detection result with the knowledge base includes: Associating the detection results with the safety construction specifications in the knowledge base to generate treatment recommendations; The illegal data in the detection result is stored in the atlas database.

8. A smart operation system, characterized in that: The system comprises: An acquisition module, configured to acquire user voice information, convert the voice information into voice text, and process the voice text to output structured data; an identification module, configured to identify the user request type and asset object in the structured data, and perform contextual completion on the structured data and report the result; A building module, configured to build a knowledge base based on a large language model, convert knowledge documents into structured knowledge data, and import the structured knowledge data into the knowledge base; A retrieval module, configured to perform Boolean matching and detection retrieval on the structured knowledge data according to user query requirements; A construction module is used to construct a two-layer LSTM network and fuse input features to generate an early warning based on the two-layer LSTM network and the fused input features; The detection module is used to train the detection model and collect scene image information, detect the scene image based on the trained detection model to obtain the detection result, and associate the detection result with the knowledge base.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the smart operation method according to any one of claims 1 to 7 is implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the smart operation method according to any one of claims 1 to 7 is implemented.