Power guarantee scheme text recognition and analysis method based on OCR and NLP

By combining OCR and NLP technologies in the text processing of power-saving solution, dynamically analyzing the event timeline and logical dependencies of power-saving tasks, the problem of inaccurate priority scheduling in the existing technology is solved, and more efficient and intelligent power-saving task scheduling is achieved.

CN120014659APending Publication Date: 2025-05-16STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411908065.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, in the priority scheduling of power-saving solution text, it is impossible to dynamically optimize the event timeline and logical dependencies according to the event, resulting in inaccurate priority resolution, affecting scheduling efficiency, and may cause delays or conflicts.

Method used

Using an OCR and NLP-based method, the image data of the power-keeping scheme text is converted into structured text data through optical character recognition, and natural language processing analysis is carried out to identify event time tags and logical dependencies. Use dynamic semantic embedding models to evaluate the fluctuation characteristics of context consistency, determine the intensity changes in the logical association between power-keeping tasks, and generate a preliminary priority, and finally confirm the final priority through priority scheduling rules.

Benefits of technology

Dynamic semantic analysis and optimization of power maintenance tasks are realized, scheduling accuracy and efficiency are improved, workload and error rate of manual analysis are reduced, and the intelligent scheduling level of power maintenance tasks is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014659A_ABST
    Figure CN120014659A_ABST
Patent Text Reader

Abstract

The invention discloses a power guarantee scheme text recognition and analysis method based on OCR and NLP, particularly relates to the technical field of text processing, and is used for solving the problem that existing power guarantee scheme text priority analysis depends on a fixed rule and cannot be dynamically adjusted. The method comprises the following steps of: converting image data of a power guarantee scheme text into structured text data through an optical character recognition technology, recognizing an event time label and a logic dependency relationship related to a power guarantee task by utilizing a natural language processing technology, and generating a priority scheduling rule; extracting context representation in combination with a dynamic semantic embedding model, evaluating fluctuation characteristics of context consistency, quantifying logical association strength change between power guarantee tasks, and generating a primary power guarantee task priority; re-analyzing the structured text data, and synthesizing a priority scheduling rule and the primary power guarantee task priority to confirm the final power guarantee task priority; and the accuracy and efficiency of power guarantee task scheduling are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text processing, and more specifically, to a text recognition and analysis method for a power conservation solution based on OCR and NLP. Background Art

[0002] A power supply plan is a planning document formulated to ensure power supply in specific scenarios, which includes the time schedule and logical relationships of power supply tasks. Optical character recognition (OCR) technology can be used to convert image data in the power supply plan text into editable text data, and natural language processing (NLP) technology can further perform semantic analysis and information extraction on the text data. In large-scale power supply tasks, the power supply plan text usually records the event time information and logical dependencies of the power supply tasks. These text contents need to be identified and analyzed to clarify the sequence and logical relationships of power supply tasks, thereby assisting in the priority scheduling and execution of power supply tasks. In actual applications, the power supply plan text is often presented in an unstructured form, including the power supply task content, time nodes and their interrelated information described in natural language.

[0003] In the prior art, the priority adjustment of power supply tasks in the power supply plan text depends on predefined rules and cannot be dynamically optimized according to the event timeline and logical dependencies of the power supply tasks, which will lead to inaccurate priority parsing of power supply tasks, affect scheduling efficiency, and may cause delays or conflicts in the execution of power supply tasks. Summary of the invention

[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a text recognition and analysis method for a power conservation solution based on OCR and NLP to solve the problems raised in the above-mentioned background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A text recognition and analysis method for power conservation scheme based on OCR and NLP includes the following steps:

[0007] Acquire image data of the power supply protection plan text, and convert the image data into structured text data through optical character recognition processing;

[0008] Perform natural language processing analysis on structured text data to identify event time labels and logical dependencies related to power supply protection tasks, and map event time labels and logical dependencies into priority scheduling rules;

[0009] Extract contextual representation of structured text data through dynamic semantic embedding model and evaluate the fluctuation characteristics of contextual consistency;

[0010] Determine the intensity change of the logical association between power-saving tasks based on the fluctuation characteristics of context consistency, and generate preliminary power-saving task priorities according to the intensity change of the logical association between power-saving tasks;

[0011] The structured text data is parsed again, and the final power supply task priority is confirmed based on the priority scheduling rules and the preliminary power supply task priority.

[0012] In a preferred embodiment, obtaining image data of the power conservation plan text and converting the image data into structured text data through optical character recognition processing specifically includes:

[0013] Acquiring image data of a power supply protection plan text based on a scanning device, where the power supply protection plan text includes a paper text and an electronic document;

[0014] Preprocess the image data, including grayscale, noise removal and edge enhancement;

[0015] Use optical character recognition algorithm to extract characters from pre-processed image data and convert character information into digital text;

[0016] The digitized text is structured, and the logical block information of paragraphs, titles and power supply task descriptions is extracted to generate structured text data.

[0017] In a preferred embodiment, natural language processing analysis is performed on structured text data to identify event time labels and logical dependencies related to power supply tasks, and the event time labels and logical dependencies are mapped into priority scheduling rules, specifically including:

[0018] Perform syntactic analysis on structured text data through grammatical parsing to extract the subject, predicate and object relationships in the sentence;

[0019] Use named entity recognition algorithms to extract event time information from structured text data and annotate the event time information as event time tags;

[0020] Based on dependency syntactic analysis, the logical dependencies between power-saving tasks are identified. The logical dependencies include the order, conditional relationship, and parallel relationship of power-saving tasks.

[0021] The extracted event time labels and logical dependencies are stored in a graph structure to generate a dependency graph between power-saving tasks.

[0022] According to the event time labels and logical dependencies in the dependency graph between power supply tasks, priority scheduling rules are generated in combination with the priority estimation model.

[0023] In a preferred embodiment, the context representation of structured text data is extracted by a dynamic semantic embedding model to evaluate the fluctuation characteristics of context consistency, specifically including:

[0024] Divide the structured text data into context windows to obtain context segments with fixed window lengths, where each context segment covers consecutive sentences.

[0025] Use a dynamic semantic embedding model to extract semantic features for each context fragment. The dynamic semantic embedding model is built based on a pre-trained language model and generates a context representation for each context fragment through a multi-layer attention mechanism.

[0026] The consistency of context representation is quantified by calculating the semantic vector change rate between context segments; the semantic vector change rate is calculated by the difference between the semantic vectors of adjacent context segments;

[0027] Multifractal detrended fluctuation analysis is used to model the fluctuation characteristics of the semantic vector change rate, and a context consistency fluctuation characteristic index is generated to evaluate the fluctuation characteristics of context consistency.

[0028] In a preferred embodiment, multifractal detrended fluctuation analysis is used to model the fluctuation characteristics of the semantic vector change rate, and a context consistency fluctuation characteristic index is generated to evaluate the fluctuation characteristics of context consistency. Specifically, the semantic vector change rate sequence is set to , the volatility characteristic index is generated through multi-fractal detrended volatility analysis, and the calculation formula is: ;in, represents the context consistency fluctuation characteristic index, Indicates The change rate of the semantic vector of the context fragment, represents the mean value of the semantic vector change rate, represents the total number of context fragments, Indicates the number of the context fragment.

[0029] In a preferred embodiment, the intensity change of the logical association between power-saving tasks is determined based on the fluctuation characteristics of context consistency, and the preliminary power-saving task priority is generated according to the intensity change of the logical association between the power-saving tasks, which specifically includes:

[0030] The larger the context consistency fluctuation characteristic index is, the more significant the change in the strength of the logical association between power-saving tasks is;

[0031] All power-saving tasks are sorted according to the size of the context consistency fluctuation characteristic index, and a preliminary power-saving task priority list is generated according to the sorting results. The power-saving tasks are arranged in order from the smallest to the largest context consistency fluctuation characteristic index.

[0032] In a preferred embodiment, the structured text data is parsed again, and the final power supply task priority is determined according to the priority scheduling rule and the preliminary power supply task priority, which specifically includes:

[0033] Obtain the priority scheduling rules and preliminary priority of power supply protection tasks;

[0034] Parse the structured text data again to identify the event time labels and logical dependencies related to each power supply task;

[0035] According to the order of power supply tasks in the priority scheduling rules and the arrangement of the preliminary power supply task priorities, combined with the logical dependencies in the structured text data, the priorities of all power supply tasks are adjusted and confirmed to generate the final power supply task priority.

[0036] The technical effects and advantages of the text recognition and analysis method of the power conservation scheme based on OCR and NLP of the present invention are as follows:

[0037] 1. By converting the image data of the power supply plan text into structured text data and combining it with natural language processing technology, the event time labels and logical dependencies in the text are accurately identified and mapped; the present invention extracts context representation through a dynamic semantic embedding model, evaluates the fluctuation characteristics of context consistency, and quantifies the changes in the logical association strength between power supply tasks, thereby generating preliminary power supply task priorities and confirming the final priorities; compared with the traditional static priority scheduling method, this method can dynamically combine the timeline and logical dependencies of the power supply task for analysis and optimization, effectively improving the accuracy and efficiency of power supply task scheduling.

[0038] 2. By processing unstructured power supply protection plan text, complex semantic information is converted into structured data that is easy to analyze, solving the problem of traditional priority adjustment relying on regularized models; through multi-level data analysis and the introduction of dynamic models, the generation process of power supply protection task priorities is optimized; when processing complex logical associations between power supply protection tasks, the method of the present invention can adapt to the diverse power supply protection task scenario requirements, reduce the workload and error rate of manual analysis, significantly improve the intelligence level of power supply protection task scheduling, and provide important technical support for achieving efficient and stable power supply guarantee. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a schematic diagram of a text recognition and analysis method for a power conservation solution based on OCR and NLP in the present invention. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0041] The present invention provides a power supply plan text analysis method combining OCR (optical character recognition) and NLP (natural language processing) technologies, which is used for automatically processing the power supply plan text and optimizing the priority of power supply tasks. First, the image data in the power supply plan text is converted into structured text data by OCR technology, and then the time information and logical relationship of each power supply task are extracted by NLP technology to generate preliminary power supply task scheduling rules. On this basis, the change characteristics between text contexts are analyzed by a dynamic semantic embedding model, and the change in the logical association strength between power supply tasks is further evaluated to preliminarily determine the priority of the power supply task. Finally, the power supply task is analyzed again in combination with the preliminary priority and scheduling rules to confirm the final priority of the power supply task. The present invention can solve the problems of many manual operations and inflexible priority rules in the existing power supply plan processing, and realize intelligent scheduling and optimization of power supply tasks.

[0042] Example: Figure 1 The present invention provides a text recognition and analysis method for a power conservation solution based on OCR and NLP, which includes the following steps:

[0043] The image data of the power conservation plan text is obtained, and the image data is converted into structured text data through optical character recognition processing.

[0044] Perform natural language processing analysis on structured text data to identify event time labels and logical dependencies related to power supply tasks, and map event time labels and logical dependencies into priority scheduling rules.

[0045] The contextual representation of structured text data is extracted through a dynamic semantic embedding model, and the fluctuation characteristics of contextual consistency are evaluated.

[0046] Based on the fluctuation characteristics of context consistency, the intensity change of logical associations between power-saving tasks is determined, and the preliminary power-saving task priority is generated according to the intensity change of logical associations between power-saving tasks.

[0047] The structured text data is parsed again, and the final power supply task priority is confirmed based on the priority scheduling rules and the preliminary power supply task priority.

[0048] Obtain image data of the power supply protection plan text, and convert the image data into structured text data through optical character recognition processing, including:

[0049] The image data of the power supply plan text is obtained based on the scanning device. The power supply plan text includes paper text and electronic document:

[0050] For paper texts, high-resolution scanners are used for image capture to ensure that the text and graphic information are fully presented.

[0051] For electronic documents, convert the electronic files into image form by taking screenshots or printing and then scanning to unify the data input format.

[0052] Preprocess the image data, including grayscale, noise removal, and edge enhancement:

[0053] Grayscale processing: Convert a color image to a grayscale image to reduce the interference of color information and retain only pixel brightness information.

[0054] Noise Removal: Use the median filter algorithm to remove possible noise in the image, such as spots produced during scanning or defects in electronic documents.

[0055] Edge enhancement: Use a convolution kernel-based image enhancement algorithm to highlight image edge features, making the outline of characters clearer and facilitating subsequent recognition.

[0056] Use optical character recognition algorithm to extract characters from preprocessed image data and convert character information into digital text:

[0057] The optical character recognition algorithm is based on a deep learning text recognition model to segment, classify and extract each character in the image. During the character extraction process, the recognition algorithm combines the specific power protection solution domain vocabulary and contextual error correction mechanism to optimize the recognition of industry terms and special characters. The extraction results are stored in text form to ensure that each character information is complete and accurate.

[0058] Perform structural processing on the digitized text, extract the logical block information of paragraphs, titles and power protection task descriptions, and generate structured text data:

[0059] Logical segmentation: Based on text grammatical features and paragraph separators, the digitized text is logically segmented into titles, paragraphs, and power-saving task descriptions.

[0060] Information extraction: Use regular expressions and natural language processing techniques to extract the core information of the power supply task, such as time tags, power supply task names, and power supply task logical dependencies.

[0061] Structured organization: The divided text content is stored according to the established hierarchical structure to form structured text data including titles, power supply task descriptions and logical relationships.

[0062] Perform natural language processing analysis on structured text data to identify event time labels and logical dependencies related to power supply tasks, and map event time labels and logical dependencies into priority scheduling rules, including:

[0063] Through grammatical parsing, the structured text data is analyzed syntactically to extract the subject, predicate and object relationships in the sentence:

[0064] The structured text data is divided into the smallest word units; the word segmentation algorithm is based on the hidden Markov model, which segments the text content through the dictionary and context probability, and annotates the part-of-speech information (such as nouns, verbs, etc.); for example, for the sentence "Power supply task A needs to be carried out after power supply task B is completed", the word segmentation results include "power supply task A / noun, need / verb, in / preposition, power supply task B / noun, complete / verb, after / adverb, proceed / verb".

[0065] The subject, predicate and object relationships in a sentence are identified through a predefined syntactic rule library. The syntactic rules include:

[0066] Subject-predicate relationship: used to identify the subject and core verb in a sentence, such as "power supply task A" and "completed" in "power supply task A completed".

[0067] Object relation: used to identify the target object of the verb, such as "power supply task B" in "complete power supply task B".

[0068] Use the dependency syntactic analysis algorithm to generate a syntactic tree to clarify the dependency relationship between the components in the sentence. For example, the syntactic tree of the sentence "Power supply task A needs to be performed after power supply task B is completed" can be represented as the subject "power supply task A", the predicate "need", the object "perform", and the additional condition "after power supply task B is completed"; the syntactic relationship can be expressed by the following formula: ;in, represents a set of syntactic dependencies, and Indicates two words in a syntactic relationship, Indicates the dependency relationship between two words (such as subject-predicate, verb-object), is a set of syntactic relation types.

[0069] Use named entity recognition algorithms to extract event time information from structured text data and annotate the event time information as event time tags:

[0070] Predefine "event time" category labels for named entity recognition algorithms, covering absolute time (such as "January 1, 2024") and relative time (such as "next Monday").

[0071] Use a time normalization algorithm to convert relative times (such as "tomorrow") to standard times (such as "January 2, 2024") to ensure uniform time information.

[0072] Based on dependency syntactic analysis, the logical dependencies between power-saving tasks are identified. The logical dependencies include the sequence, conditional relationship, and parallel relationship of power-saving tasks:

[0073] Based on the syntax tree, identify the logical dependencies between power-saving tasks, including:

[0074] Sequential relationship: Power supply task A can only be started after power supply task B is completed;

[0075] Conditional relationship: Completion of power supply task C is the prerequisite for starting power supply task D;

[0076] Parallel relationship: Power supply task E and power supply task F can be carried out at the same time.

[0077] Logical dependencies can be expressed as: ;in, A collection representing the logical dependencies between power-saving tasks; and Used to represent two power-saving tasks respectively. ; Represents a set of all power supply tasks extracted from structured text data, for example, including power supply task A, power supply task B, power supply task C, etc.; Indicates the types of logical dependencies between power-saving tasks, including sequential relationships (the order in which power-saving tasks are completed, such as "power-saving task A is completed after power-saving task B"), conditional relationships (necessary conditions for the completion of power-saving tasks, such as "the completion of power-saving task C is the prerequisite for power-saving task D"), and parallel relationships (power-saving tasks are carried out simultaneously, such as "power-saving task E and power-saving task F can be executed at the same time"); Represents a collection of logical dependency types.

[0078] The extracted event time labels and logical dependencies are stored in a graph structure to generate a dependency graph between power-saving tasks:

[0079] The logical dependency relationship between power supply protection tasks is represented as a graph structure, where nodes are power supply protection tasks and edges are logical dependencies.

[0080] The dependency graph between power-saving tasks can be expressed as: ;in, Represents the dependency graph between power-keeping tasks; Represents the power supply task set, including all power supply tasks; Represents a set of dependencies between power-saving tasks.

[0081] Based on the event time labels and logical dependencies in the dependency graph between power-saving tasks, combined with the priority estimation model, priority scheduling rules are generated:

[0082] The priority of each power supply task is calculated through a weighted model, taking into account factors such as the urgency of the power supply task, resource requirements, and time constraints.

[0083] The priority of the power supply task is expressed as: ;in, Indicates the priority of the power supply task; Indicates the urgency of the power supply task, which is calculated based on the time limit and criticality of the power supply task; It indicates the resource requirements of the power supply protection task, which is derived from the estimation of the human, material or equipment resources required to perform the power supply protection task; It indicates the time constraint of the power supply task, which is calculated by the start time, end time or duration of the power supply task; , and They represent the urgency of the power supply task, the resource requirements, and the weight coefficients of the time constraint, respectively, and , and Both are greater than 0.

[0084] According to the priority of the power-saving task in the priority calculation result, the power-saving tasks are sorted to form a priority scheduling rule. For example, the power-saving task with the highest priority is executed first, followed by the power-saving task with the second highest priority.

[0085] The context representation of structured text data is extracted through a dynamic semantic embedding model to evaluate the fluctuation characteristics of context consistency, including:

[0086] The structured text data is divided into context windows to obtain context segments with fixed window lengths. Each context segment covers consecutive sentences:

[0087] The structured text data is divided into fixed window lengths to generate multiple context segments. Each context segment covers consecutive sentences, and the window length is set according to specific application requirements, such as containing a fixed number of sentences or paragraphs.

[0088] When dividing, the structured text data is first segmented into sentences to ensure that the start and end positions of each segment completely cover the boundaries of the sentence to avoid truncation of semantic information; the division process follows a sliding window strategy, that is, adjacent context segments can partially overlap to preserve the continuity between contexts.

[0089] The dynamic semantic embedding model is used to extract semantic features for each context fragment. The dynamic semantic embedding model is built based on the pre-trained language model and generates the context representation of each context fragment through a multi-layer attention mechanism:

[0090] The dynamic semantic embedding model is built based on a pre-trained language model, such as BERT and other deep learning models, and generates contextual representations through a multi-layer attention mechanism.

[0091] The pre-trained language model has been trained on a large-scale corpus and can capture rich semantic features; the multi-layer attention mechanism can extract semantic relations in fragments hierarchically, including dependencies between words and global correlations in the context.

[0092] For each context segment, the sentences in the segment are converted into word vector sequences according to the word order. Then, the word vectors are embedded using the dynamic semantic embedding model to generate the semantic feature representation of the context segment.

[0093] The context embedding process is expressed as: ;in, A semantic feature vector representing the context fragment; A sequence of word vectors representing the context fragments of the input; Represents a dynamic semantic embedding model and generates embedded representations through an attention mechanism.

[0094] The consistency of context representation is quantified by calculating the semantic vector change rate between context segments; the semantic vector change rate is calculated by the difference between the semantic vectors of adjacent context segments:

[0095] The semantic vector change rate between context segments is used to describe the degree of semantic change between adjacent segments and is obtained by calculating the difference between the semantic feature vectors of adjacent context segments.

[0096] Assume that the semantic feature vectors of two adjacent context fragments are and , then the calculation formula of semantic vector change rate is: ;in, Indicates the change rate of the semantic vector; and Represent the semantic feature vectors of two adjacent context fragments respectively; Represents the norm of the vector difference, usually the Euclidean norm (L2 norm).

[0097] The smaller the semantic vector change rate, the higher the consistency between adjacent context fragments; the larger the semantic vector change rate, the lower the consistency.

[0098] Multifractal detrended fluctuation analysis is used to model the fluctuation characteristics of the semantic vector change rate, and a context consistency fluctuation characteristic index is generated to evaluate the fluctuation characteristics of context consistency:

[0099] Based on the semantic vector change rate, the fluctuation characteristics of context consistency are analyzed. Multifractal detrended fluctuation analysis is a nonlinear dynamic modeling method used to characterize the complex distribution characteristics of semantic changes in context.

[0100] The process of fluctuation characteristic modeling: taking the time series of semantic vector change rate as input; modeling the fractal characteristics of local fluctuations of the change rate through multi-fractal detrended fluctuation analysis; generating a context-consistent fluctuation characteristic index to quantify the distribution characteristics of the change rate.

[0101] Assume that the sequence of semantic vector change rate is , the volatility characteristic index is generated through multi-fractal detrended volatility analysis, and the calculation formula is: ;in, represents the context consistency fluctuation characteristic index, Indicates The change rate of the semantic vector of the context fragment, represents the mean value of the semantic vector change rate, represents the total number of context fragments, Indicates the number of the context fragment.

[0102] The larger the context consistency fluctuation characteristic index is, the greater the fluctuation characteristic of the context consistency is, indicating that the semantic change of the context is more drastic, reflecting the low consistency between different context fragments in the text. The intensification of this fluctuation characteristic may lead to insufficient coherence and consistency of the context semantics, and then introduce uncertainty and erroneous parsing in subsequent power supply tasks (such as logical dependency analysis or priority scheduling rule generation). For example, logical associations may not be accurately extracted due to semantic fragmentation, which in turn affects the dynamic adjustment of priorities and the accuracy of power supply task scheduling.

[0103] Based on the fluctuation characteristics of context consistency, the intensity change of the logical association between power-saving tasks is determined, and the preliminary power-saving task priority is generated according to the intensity change of the logical association between power-saving tasks, including:

[0104] The greater the fluctuation characteristics of contextual consistency, the more significant the change in the strength of logical associations between power-saving tasks. This indicates that the semantic consistency between different power-saving tasks in the text is low, and there may be obvious association fluctuations, reflecting the instability or complexity of the logical relationship of power-saving tasks. In this case, the logical association needs to be further quantified to accurately describe the dynamic changes between power-saving tasks.

[0105] The context consistency fluctuation characteristic index is used to represent the degree of semantic fluctuation of each power-saving task. The larger the context consistency fluctuation characteristic index is, the lower the semantic consistency of the power-saving task is and the more significant the change in the strength of logical association is.

[0106] All power-saving tasks are sorted according to the size of the context consistency fluctuation characteristic index. The power-saving tasks with higher context consistency fluctuation characteristic index have lower priority, and the power-saving tasks with lower context consistency fluctuation characteristic index have higher priority.

[0107] A preliminary priority list of power supply tasks is generated according to the sorting results, and the power supply tasks are arranged in order from the smallest to the largest context consistency fluctuation characteristic index.

[0108] The structured text data is parsed again, and the final power supply task priority is confirmed according to the priority scheduling rules and the initial power supply task priority, including:

[0109] Get the priority scheduling rules and preliminary priority of power supply tasks:

[0110] The priority scheduling rule is obtained by reading the output data of the previous step and formatted into a list containing power supply task identifiers and their priority values, for example: power supply task 1, priority value: 10; power supply task 2, priority value: 8.

[0111] By retrieving the stored preliminary priority results, a power supply task priority list is generated, for example: power supply task 1, preliminary priority: 1; power supply task 2, preliminary priority: 2.

[0112] Parse the structured text data again to identify the event time labels and logical dependencies related to each power protection task:

[0113] The structured text data is parsed segment by segment to extract the contextual information of each power supply task, including the power supply task description and the event time labels and logical dependencies associated with the power supply task.

[0114] The parsing process includes: locating the time attributes of the power supply tasks based on the grammatical features of the event time label (such as date or time format); identifying the logical dependencies between the power supply tasks through logical connectives (such as "if", "after", and "at the same time").

[0115] In structured text data, named entity recognition algorithms are used to extract time information related to each power supply task and annotate it as an event time label. Event time labels are used to determine the time constraints of power supply tasks, for example: the time label of power supply task 1 is December 1, 2024; the time label of power supply task 2 is December 5, 2024.

[0116] The logical dependencies between power-saving tasks are extracted based on the dependency syntax analysis algorithm, including: sequential relationship: power-saving task 1 must be completed before power-saving task 2; conditional relationship: the completion of power-saving task 3 is the prerequisite for the start of power-saving task 4; parallel relationship: power-saving task 5 and power-saving task 6 can be executed at the same time.

[0117] By marking the logical dependency of each power supply task, a logical dependency table is formed, for example: power supply task 1 depends on power supply task 2 (sequential relationship); power supply task 3 depends on power supply task 4 (conditional relationship).

[0118] According to the order of power supply tasks in the priority scheduling rules and the arrangement of the preliminary power supply task priorities, combined with the logical dependencies in the structured text data, the priorities of all power supply tasks are adjusted and confirmed to generate the final power supply task priorities:

[0119] The power supply task ranking in the priority scheduling rule is compared with the preliminary power supply task priority to identify priority conflicts and power supply tasks that need to be adjusted; for example, if power supply task 1 is ranked after power supply task 2 in the priority scheduling rule, but power supply task 1 has a higher priority in the preliminary priority, the conflict is marked.

[0120] According to the logical dependencies in the structured text data, the power supply tasks with priority conflicts are adjusted: if there is a sequential relationship between the power supply tasks, the sequential relationship is given priority; if there is a conditional relationship between the power supply tasks, the priority requirements of the prerequisite power supply tasks are met; for the power supply tasks with parallel relationships, they are optimized and sorted based on the preliminary priority values.

[0121] The final priority value of each power supply task is recalculated by integrating priority scheduling rules, preliminary power supply task priorities and logical dependencies.

[0122] The final priority value sorting logic is as follows:

[0123] Give priority to satisfying logical dependencies; when satisfying logical dependencies, give priority to power supply tasks with high preliminary priority values; sort power supply tasks with the same conditions in accordance with priority scheduling rules.

[0124] The finally generated power supply task priority list includes the final priority value and ranking of each power supply task, for example, power supply task 1, final priority: 1; power supply task 2, final priority: 2; power supply task 3, final priority: 3.

[0125] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.

[0126] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or may be transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium may be a solid-state hard disk.

[0127] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0128] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0129] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0130] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0131] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0132] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.

[0133] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0134] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A text recognition and analysis method for power conservation scheme based on OCR and NLP, characterized in that: The steps include: Acquire image data of the power supply protection plan text, and convert the image data into structured text data through optical character recognition processing; Perform natural language processing analysis on structured text data to identify event time labels and logical dependencies related to power supply protection tasks, and map event time labels and logical dependencies into priority scheduling rules; Extract contextual representation of structured text data through dynamic semantic embedding model and evaluate the fluctuation characteristics of contextual consistency; Determine the intensity change of the logical association between power-saving tasks based on the fluctuation characteristics of context consistency, and generate preliminary power-saving task priorities according to the intensity change of the logical association between power-saving tasks; The structured text data is parsed again, and the final power supply task priority is confirmed based on the priority scheduling rules and the preliminary power supply task priority.

2. According to the method for text recognition and analysis of power conservation scheme based on OCR and NLP in claim 1, it is characterized in that: Obtain image data of the power supply protection plan text, and convert the image data into structured text data through optical character recognition processing, including: Acquiring image data of a power supply protection plan text based on a scanning device, where the power supply protection plan text includes a paper text and an electronic document; Preprocess the image data, including grayscale, noise removal and edge enhancement; Use optical character recognition algorithm to extract characters from pre-processed image data and convert character information into digital text; The digitized text is structured, and the logical block information of paragraphs, titles and power supply task descriptions is extracted to generate structured text data.

3. The method for text recognition and analysis of power conservation scheme based on OCR and NLP according to claim 2 is characterized in that: Perform natural language processing analysis on structured text data to identify event time labels and logical dependencies related to power supply tasks, and map event time labels and logical dependencies into priority scheduling rules, including: Perform syntactic analysis on structured text data through grammatical parsing to extract the subject, predicate and object relationships in the sentence; Use named entity recognition algorithms to extract event time information from structured text data and annotate the event time information as event time tags; Based on dependency syntactic analysis, the logical dependencies between power-saving tasks are identified. The logical dependencies include the order, conditional relationship, and parallel relationship of power-saving tasks. The extracted event time labels and logical dependencies are stored in a graph structure to generate a dependency graph between power-saving tasks. According to the event time labels and logical dependencies in the dependency graph between power supply tasks, priority scheduling rules are generated in combination with the priority estimation model.

4. The method for text recognition and analysis of power conservation scheme based on OCR and NLP according to claim 1 is characterized in that: The context representation of structured text data is extracted through a dynamic semantic embedding model to evaluate the fluctuation characteristics of context consistency, including: Divide the structured text data into context windows to obtain context segments with fixed window lengths, where each context segment covers consecutive sentences. Use a dynamic semantic embedding model to extract semantic features for each context fragment. The dynamic semantic embedding model is built based on a pre-trained language model and generates a context representation for each context fragment through a multi-layer attention mechanism. The consistency of context representation is quantified by calculating the semantic vector change rate between context segments; the semantic vector change rate is calculated by the difference between the semantic vectors of adjacent context segments; Multifractal detrended fluctuation analysis is used to model the fluctuation characteristics of the semantic vector change rate, and a context consistency fluctuation characteristic index is generated to evaluate the fluctuation characteristics of context consistency.

5. The method for text recognition and analysis of power conservation scheme based on OCR and NLP according to claim 4 is characterized in that: Multifractal detrended fluctuation analysis is used to model the fluctuation characteristics of semantic vector change rate, and a context consistency fluctuation characteristic index is generated to evaluate the fluctuation characteristics of context consistency. Specifically, let the semantic vector change rate sequence be , the volatility characteristic index is generated through multi-fractal detrended volatility analysis, and the calculation formula is: ;in, represents the context consistency fluctuation characteristic index, Indicates The change rate of the semantic vector of the context fragment, represents the mean value of the semantic vector change rate, represents the total number of context fragments, Indicates the number of the context fragment.

6. The method for text recognition and analysis of power conservation scheme based on OCR and NLP according to claim 5 is characterized in that: Based on the fluctuation characteristics of context consistency, the intensity change of the logical association between power-saving tasks is determined, and the preliminary power-saving task priority is generated according to the intensity change of the logical association between power-saving tasks, including: The larger the context consistency fluctuation characteristic index is, the more significant the change in the strength of the logical association between power-saving tasks is; All power-saving tasks are sorted according to the size of the context consistency fluctuation characteristic index, and a preliminary power-saving task priority list is generated according to the sorting results. The power-saving tasks are arranged in order from the smallest to the largest context consistency fluctuation characteristic index.

7. The method for text recognition and analysis of power conservation scheme based on OCR and NLP according to claim 1, characterized in that: The structured text data is parsed again, and the final power supply task priority is confirmed according to the priority scheduling rules and the initial power supply task priority, including: Obtain the priority scheduling rules and preliminary priority of power supply protection tasks; Parse the structured text data again to identify the event time labels and logical dependencies related to each power supply task; According to the order of power supply tasks in the priority scheduling rules and the arrangement of the preliminary power supply task priorities, combined with the logical dependencies in the structured text data, the priorities of all power supply tasks are adjusted and confirmed to generate the final power supply task priority.

Citation Information

Cited By

  • Data security classification and grading method and system based on large language model

    CN120523956A