Electric power operation site intelligent control image-text fusion labeling method and related device
By acquiring images at power operation sites, performing feature extraction and preprocessing, and using image recognition and natural language processing models to fuse feature vectors, the problem of weak correlation between images and text is solved, efficient and accurate image-text fusion annotation is achieved, and the accuracy and efficiency of annotation are improved.
Patent Information
- Application Number
- CN202510895740.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies have weak correlation between image recognition and text generation at power operation sites, and lack effective image-text fusion annotation methods, resulting in poor annotation accuracy and affecting the comprehensive understanding and analysis of on-site conditions.
By acquiring images of power operation sites, performing feature extraction and preprocessing, and using image recognition models and natural language processing models to extract image and text feature vectors respectively, they are fused through the attention mechanism and annotated in combination with preset annotation rules to achieve accurate association between images and text.
It improves the accuracy and efficiency of annotation, realizes the deep integration of image and text information, significantly improves the accuracy and consistency of annotation results, and can more accurately reflect the actual situation at the power operation site.
Smart Images

Figure HDA0005475948530000011 
Figure HDA0005475948530000012
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of intelligent management and control of power operation sites, and relates to an intelligent management and control image-text fusion labeling method for power operation sites and related devices. BACKGROUND
[0002] In the power industry, with the development of intelligentization and digitization, the management and control requirements for power operation sites are becoming higher and higher, and it is necessary to efficiently, accurately record, analyze and process the site conditions to ensure the safety and efficiency of power operation. In the power operation site, the traditional management and control method mainly relies on manual recording and simple video monitoring. Manual recording is prone to information omission and inaccuracy, and simple video monitoring can only provide picture information, making it difficult to analyze and label the key content in the picture. With the development of computer vision and natural language processing technology, some power operation site management and control methods based on image recognition and information extraction have gradually emerged, but these methods mostly only focus on single modal information of images or texts, and fail to fully utilize the complementarity of multi-modal information.
[0003] The existing technology has the following deficiencies in the management and control of power operation sites: first, the correlation between image recognition and text generation is not strong, resulting in inconsistent image-text information; second, there is a lack of an effective image-text fusion labeling method, so it is difficult to accurately correspond the key information in the image with the corresponding text description, affecting the comprehensive understanding and analysis of the power operation site, and leading to a large deviation in the accuracy of labeling. SUMMARY
[0004] The purpose of the present application is to overcome the above-mentioned deficiencies of the prior art, and to provide an intelligent management and control image-text fusion labeling method for power operation sites and related devices, which can improve the accuracy of labeling.
[0005] To achieve the above-mentioned purpose, the present application discloses an intelligent management and control image-text fusion labeling method for power operation sites, comprising:
[0006] obtaining a power operation site image;
[0007] performing feature extraction on the power operation site image to obtain an image feature vector and a text feature vector;
[0008] fusing the image feature vector and the text feature vector to obtain fused image-text information;
[0009] labeling the fused image-text information to complete the intelligent management and control image-text fusion labeling for power operation sites.
[0010] Further improvement of the intelligent management and control image-text fusion labeling method for power operation sites is as follows:
[0011] Further, the feature extraction on the power operation site image further includes:
[0012] The power operation site image is subjected to Gaussian filter denoising and histogram equalization enhancement.
[0013] Further, the feature extraction on the power operation site image includes:
[0014] The power operation site image is input into the trained image recognition model to obtain an image feature vector.
[0015] The image feature vector is input into a natural language processing model to obtain a text feature vector.
[0016] Further, the fusion of the image feature vector and the text feature vector to obtain the fused image-text information includes:
[0017] The image feature vector and the text feature vector are fused by using an attention mechanism to obtain the fused image-text information.
[0018] Further, the annotation of the fused image-text information includes:
[0019] The fused image-text information is annotated according to a preset annotation rule.
[0020] The application discloses an intelligent power operation site management and control image-text fusion annotation system, which comprises:
[0021] An acquisition module is configured to acquire a power operation site image.
[0022] An extraction module is configured to extract features of the power operation site image to obtain an image feature vector and a text feature vector.
[0023] A fusion module is configured to fuse the image feature vector and the text feature vector to obtain fused image-text information.
[0024] An annotation module is configured to annotate the fused image-text information to complete intelligent power operation site management and control image-text fusion annotation.
[0025] The application further improves the intelligent power operation site management and control image-text fusion annotation system.
[0026] Further, the feature extraction on the power operation site image further includes:
[0027] The power operation site image is subjected to Gaussian filter denoising and histogram equalization enhancement.
[0028] Further, the extraction module comprises:
[0029] A first extraction unit is configured to input the power operation site image into the trained image recognition model to obtain an image feature vector;
[0030] A second extraction unit is configured to input the image feature vector into a natural language processing model to obtain a text feature vector.
[0031] Further, the process of fusing the image feature vector and the text feature vector to obtain the fused image-text information comprises:
[0032] The image feature vector and the text feature vector are fused by using an attention mechanism to obtain the fused image-text information.
[0033] Further, the process of labeling the fused image-text information comprises:
[0034] The fused image-text information is labeled according to a preset labeling rule.
[0035] The application discloses a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the power operation site intelligent management and control image-text fusion labeling method when executing the computer program.
[0036] The application discloses a computer readable storage medium, which stores a computer program, wherein the computer program implements the steps of the power operation site intelligent management and control image-text fusion labeling method when executed by a processor.
[0037] The application has the following beneficial effects:
[0038] The power operation site intelligent management and control image-text fusion labeling method and related device disclosed by the application fuse the image feature vector and the text feature vector to obtain the fused image-text information, and then label the fused image-text information, thereby breaking through the limitation of traditional single modal information processing and improving the labeling accuracy.
[0039] Further, the image feature vector and the text feature vector are deeply fused by using an attention mechanism, and the image and text information are accurately associated.
[0040] Further, the image feature vector is obtained by using an image recognition model, and the text feature vector is obtained by using a natural language processing model, thereby reducing the artificial workload, avoiding errors and delays that may occur in manual operation, enabling more tasks to be completed in a shorter time, and significantly improving the labeling efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings, which constitute part of the present invention, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0042] Figure 1 is a flow chart of the method of the present invention;
[0043] Figure 2 This is a system structure diagram of the present invention. DETAILED DESCRIPTION
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0045] In the description of the present invention, it is to be understood that the terms “include” and “comprise” indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0046] It should also be understood that the terms used in the present specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0047] It should be further understood that the term "and / or" as used in the present specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally indicates that the associated objects are in an "or" relationship.
[0048] It should be understood that although the terms "first," "second," and "third" may be used to describe preset ranges in embodiments of the present invention, these preset ranges should not be limited to these terms. These terms are merely used to distinguish one preset range from another. For example, without departing from the scope of embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0049] Depending on the context, the word "if" as used herein can be interpreted to mean "when" or "while" or "in response to determining" or "in response to detecting." Similarly, the phrase "if it is determined" or "if [a stated condition or event] is detected" can be interpreted to mean "when it is determined" or "in response to determining" or "when [a stated condition or event] is detected" or "in response to detecting [a stated condition or event]."
[0050] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0051] Various structural schematic diagrams according to the disclosed embodiments of the present application are shown in the drawings. These diagrams are not drawn to scale, in which some details are exaggerated for the purpose of clarity and some details can be omitted. The shapes of various regions, layers and their relative sizes and positional relationships shown in the drawings are only exemplary, and in actuality, there can be deviations due to manufacturing tolerances or technical limitations, and a person skilled in the art can additionally design regions / layers with different shapes, sizes and relative positions according to actual needs.
[0052] It is known that image recognition technology: can perform target detection, classification and other operations on the image of the power operation site, and identify key elements such as equipment and personnel.
[0053] Natural language processing technology: used for processing text information, such as generating descriptions of image content, extracting key information, etc.
[0054] Multimodal fusion technology: fuses information of different modalities such as images and texts to obtain more comprehensive and accurate information.
[0055] Embodiment one
[0056] Reference Figure 1 The intelligent management and control image-text fusion labeling method for the power operation site described in the present application comprises the following steps:
[0057] 1) Obtain the image of the power operation site;
[0058] 2) preprocessing the power operation site image to obtain a preprocessed image;
[0059] Specifically, the power operation site image is subjected to Gaussian filter denoising and histogram equalization enhancement to improve the quality of the image and obtain a preprocessed image; wherein the power operation site image is sequentially subjected to Gaussian filter denoising and histogram equalization enhancement.
[0060] 3) inputting the preprocessed image into a trained image recognition model to identify key elements in the image, the key elements including equipment, personnel and safety signs, to obtain an image feature vector;
[0061] 4) inputting the image feature vector into a natural language processing model to obtain a text feature vector, the text feature vector being used to describe the content of the image in text, including the position and state of the key elements;
[0062] It should be noted that before inputting the image feature vector into the natural language processing model, a large number of power operation site images and their corresponding text description data need to be collected, and the natural language processing model is trained based on the data, so that it can accurately learn the mapping relationship between images and texts.
[0063] 5) fusing the image feature vector and the text feature vector through a graphic-text fusion algorithm to determine the association between the image and the text, to obtain fused graphic-text information, wherein the graphic-text fusion algorithm can be an attention mechanism;
[0064] 6) according to a preset annotation rule, annotating the fused graphic-text information, for example, for the identified equipment, annotating the name, model and running state of the equipment, and for the identified personnel, annotating the work content, safety equipment wearing condition and whether there is a violation behavior.
[0065] The annotation rule is: combining the actual needs and management specifications of the power operation site, a detailed and accurate annotation rule is formulated to ensure the consistency and standardization of the annotation results. The annotation rule includes the annotation standard of equipment type, state, position and other information, and the annotation specification of personnel behavior, safety equipment wearing condition, and whether there is a violation.
[0066] It should be noted that the present application integrates the processing functions of multiple modal information such as images and texts, from data acquisition, preprocessing, to the execution of the labeling process, to the update and maintenance of the labeled version, and finally to the generation and storage of the labeling result, forming a complete process. Through such a complete tool chain, the entire labeling work is more systematic and standardized, providing a solid framework support for subsequent labeling work. In addition, with the help of advanced algorithms and technologies, the present application quickly and accurately completes the association and labeling of image and text information. By using deep learning algorithm to extract image features and analyze text semantics, the intelligent matching algorithm is used to accurately associate the two, and the labeling result is automatically generated, which reduces the artificial workload, avoids errors and delays that may occur in manual operation, and enables the labeling work to complete more tasks in a shorter time, significantly improving the labeling efficiency. Finally, it should be noted that the present application realizes the accurate association of image and text information through the image-text fusion algorithm, and the accuracy of the labeling result is improved by 30% compared with the traditional method, which can more accurately reflect the actual situation of the power operation site and ensure that the labeling can truly and accurately reflect the actual situation of the power operation site, providing a reliable basis for subsequent work.
[0067] Embodiment two
[0068] Reference Figure 2 The power operation site intelligent management and control image-text fusion labeling system provided by the present application comprises:
[0069] The acquisition module is configured to acquire the power operation site image.
[0070] The extraction module is configured to extract features of the power operation site image to obtain an image feature vector and a text feature vector.
[0071] The fusion module is configured to fuse the image feature vector and the text feature vector to obtain fused image-text information.
[0072] The labeling module is configured to label the fused image-text information to complete the power operation site intelligent management and control image-text fusion labeling.
[0073] In this embodiment, before the feature extraction of the power operation site image, the following steps are further included:
[0074] The power operation site image is subjected to Gaussian filter denoising and histogram equalization enhancement.
[0075] In this embodiment, the extraction module comprises:
[0076] The first extraction unit is configured to input the power operation site image into the trained image recognition model to obtain an image feature vector.
[0077] A second extraction unit is configured to input the image feature vector into a natural language processing model to obtain a text feature vector.
[0078] In this embodiment, the process of fusing the image feature vector and the text feature vector to obtain the fused image-text information is as follows:
[0079] The image feature vector and the text feature vector are fused by using an attention mechanism to obtain the fused image-text information.
[0080] In this embodiment, the process of labeling the fused image-text information is as follows:
[0081] The fused image-text information is labeled according to a preset labeling rule.
[0082] The division of the modules in the embodiments of the present application is illustrative, and is merely a logical functional division. In actual implementation, another division manner can be used. In addition, each functional module in each embodiment of the present application can be integrated in one processor, or can be physically separated, or two or more modules can be integrated in one module. The integrated module can be implemented in the form of hardware or in the form of a software functional module.
[0083] Embodiment three
[0084] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the power operation site intelligent management and control image-text fusion labeling method are implemented, for example, including: obtaining a power operation site image; performing feature extraction on the power operation site image to obtain an image feature vector and a text feature vector; fusing the image feature vector and the text feature vector to obtain fused image-text information; and labeling the fused image-text information to complete power operation site intelligent management and control image-text fusion labeling. The memory can include a memory, such as a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk memory. The processor, network interface, and memory are connected to each other through an internal bus, which can be an industry standard architecture bus, a peripheral component interconnect standard bus, an extended industry standard architecture bus, etc. The bus can be divided into an address bus, a data bus, and a control bus. The memory is used to store programs, and specifically, the programs can include program codes, and the program codes include computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0085] Embodiment four
[0086] A computer readable storage medium stores a computer program, the computer program is executed by a processor to implement steps of the power operation site intelligent management and control graphic-text fusion labeling method, for example, comprising: acquiring a power operation site image; performing feature extraction on the power operation site image to obtain an image feature vector and a text feature vector; fusing the image feature vector and the text feature vector to obtain fused graphic-text information; and labeling the fused graphic-text information to complete the power operation site intelligent management and control graphic-text fusion labeling. Specifically, the computer readable storage medium includes but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory can include random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include read-only memory (ROM), hard disk, flash memory, optical disc, magnetic disc, etc.
[0087] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0088] The present application is described with reference to flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks. Figure 1 The functions specified in one or more flows and / or blocks.
[0089] These computer program instructions can also be stored in a computer readable storage medium that can direct the computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer readable storage medium produce a manufactured product including instruction devices that implement the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks. Figure 1 The functions specified in one or more flows and / or blocks.
[0090] These computer program instructions can also be loaded into a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 Figure 1
[0091] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the application being indicated by the following claims.
[0092] It is to be understood that the application is not limited to the precise details of construction and the method described above and illustrated in the drawings. Various modifications and changes can be made thereunto without departing from the scope of the application. The scope of the application is indicated by the appended claims, rather than by the foregoing description.
[0093] The above description is only preferred embodiments of the present application, not any limitation thereto, any simple modification, change and equivalent structure change of the above embodiments according to the technical essence of the present application are still within the protection scope of the technical scheme of the present application.
Claims
1. A method for integrating graphics and text into intelligent management and control of power operation sites, characterized in that: include: Acquire images of power operation sites; Extracting features from the power operation site image to obtain an image feature vector and a text feature vector; Fusing the image feature vector with the text feature vector to obtain fused image and text information; The fused graphic and text information is annotated to complete the graphic and text fusion annotation for the intelligent management and control of the power operation site.
2. The method for integrating graphics and text into intelligent management and control of power operation sites according to claim 1 is characterized in that: Before extracting features from the power operation site image, the following steps are further included: Gaussian filtering denoising and histogram equalization enhancement are performed on the power operation site image.
3. The method for integrating graphics and text into intelligent management and control of power operation sites according to claim 1 is characterized in that: The process of extracting features from the power operation site image to obtain image feature vectors and text feature vectors is as follows: Inputting the power operation site image into a trained image recognition model to obtain an image feature vector; The image feature vector is input into a natural language processing model to obtain a text feature vector.
4. The method for integrating graphics and text into intelligent management and control of power operation sites according to claim 1 is characterized in that: The process of fusing the image feature vector and the text feature vector to obtain the fused image-text information is as follows: The image feature vector and the text feature vector are fused by using the attention mechanism to obtain fused image and text information.
5. The method for integrating graphics and text into intelligent management and control of power operation sites according to claim 1 is characterized in that: The process of labeling the fused graphic and text information is as follows: The fused graphic and text information is annotated according to preset annotation rules.
6. A system for integrating graphics and text into intelligent management and control of power operation sites, characterized in that: include: An acquisition module, used to acquire images of the power operation site; An extraction module, configured to extract features from the power operation site image to obtain an image feature vector and a text feature vector; A fusion module, configured to fuse the image feature vector with the text feature vector to obtain fused image and text information; The annotation module is used to annotate the fused graphic and text information to complete the graphic and text fusion annotation for the intelligent management and control of the power operation site.
7. The electric power operation site intelligent management and control graphic and text fusion annotation system according to claim 6 is characterized in that: Before extracting features from the power operation site image, the following steps are further included: Gaussian filtering denoising and histogram equalization enhancement are performed on the power operation site image.
8. The electric power operation site intelligent management and control graphic and text fusion annotation system according to claim 6 is characterized in that: The extraction module includes: A first extraction unit is configured to input the power operation site image into a trained image recognition model to obtain an image feature vector; The second extraction unit is used to input the image feature vector into a natural language processing model to obtain a text feature vector.
9. The intelligent control and management graphic and text fusion annotation system for electric power operation sites according to claim 6 is characterized in that: The process of fusing the image feature vector and the text feature vector to obtain the fused image-text information is as follows: The image feature vector and the text feature vector are fused by using the attention mechanism to obtain fused image and text information.
10. The electric power operation site intelligent management and control graphic and text fusion annotation system according to claim 6, characterized in that: The process of labeling the fused graphic and text information is as follows: The fused graphic and text information is annotated according to preset annotation rules.
11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the graphic and text fusion annotation method for intelligent management and control of power operation sites as described in any one of claims 1 to 5 are implemented.
12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for fusion annotation of graphics and text for intelligent management and control of power operation sites as described in any one of claims 1 to 5 are implemented.