Target detection method based on image-text multi-modal optimization

By adopting a multimodal optimization of graphics and text in the distribution station, combining voltage and current sensors and camera data, the accuracy problem of traditional object detection models in identifying power equipment failure types is solved, and more efficient and accurate fault identification and processing is achieved.

CN119942562APending Publication Date: 2025-05-06NANJING SHENDA ENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510053321.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When traditional target detection models identify the fault types of power equipment in power station buildings, due to the single input data, accurate identification cannot be achieved.

Method used

The object detection method based on multimodal optimization of graphics and text is adopted, and the actual voltage and current of the power equipment and the power topology diagram of the distribution station building are collected through voltage and current sensors, fault location information and fault question text are generated, and the monitoring images taken by the camera are combined with the encoding process through the image encoder and text encoder to calculate the similarity of the coded vector to determine the fault type.

Benefits of technology

Fault identification based on multimodal data is realized, the accuracy and efficiency of fault identification are improved, manual intervention is reduced, and the level of automation of fault processing is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942562A_ABST
    Figure CN119942562A_ABST
Patent Text Reader

Abstract

The invention provides a target detection method based on image-text multi-modal optimization. The method comprises the following steps: when it is determined that power equipment in a power distribution station room has a fault according to actual voltage and current of each piece of power equipment and a power topological graph of the power distribution station room, generating fault position information according to the voltage and current of each piece of power equipment, and generating a fault question text according to the fault position information, acquiring a monitoring image shot by a camera according to the fault position information, and processing the fault question text and the monitoring image by using an input layer to generate an input vector; an image encoder is used for encoding an input vector to generate an image encoding vector, the similarity between the image encoding vector and a text encoding vector of each description text in a preset description text set is calculated, the description text with the maximum similarity is selected to determine a fault type, and fault recognition based on image-text multi-modal data is achieved. And the question text does not need to be manually input, so that the fault processing accuracy and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of electric power Internet of Things, and in particular to a target detection method based on multi-modal optimization of graphics and text. Background Art

[0002] With the development of science and technology, the scale and complexity of power systems are increasing. As an important part of the power system, the safe and stable operation of distribution substations is crucial.

[0003] Traditional human-based inspection and management methods can no longer fully meet the needs of modern power distribution stations. In recent years, object detection models, as an important application in the field of artificial intelligence, have gradually been introduced into the management and maintenance process of power distribution stations.

[0004] However, target detection models usually only take image data as input and classify it after extracting features from the image data. Due to the single input data, it is impossible to accurately identify the fault type of power equipment in the distribution station. Summary of the invention

[0005] The present application provides a target detection method based on image-text multimodal optimization, which is used to achieve the effect of determining the fault type of power equipment in a distribution station based on monitoring images.

[0006] The present application provides a target detection method based on multi-modal optimization of graphics and text. The target detection method is applied to a background server, which is communicated with a camera located in a power distribution station. Voltage and current sensors are arranged on various power equipment in the power distribution station. The method includes: Obtain the actual voltage and current of each power device collected by the voltage and current sensors and the power topology diagram of the distribution station; Determine whether there is a fault in the power equipment in the power distribution station based on the actual voltage and current of each power equipment and the power topology diagram of the power distribution station; If yes, generate fault location information according to the voltage and current of each power device, generate a fault question text according to the fault location information, and obtain a monitoring image taken by a camera according to the fault location information; Use the input layer to process the fault question text and the monitoring image to generate an input vector; use the image encoder to encode the input vector to generate an image encoding vector; Calculate the similarity between the image coding vector and the text coding vector of each description text in the preset description text set; select the description text with the greatest similarity as the target description text output; and determine the fault type according to the target description text.

[0007] In the above technical scheme, multiple cameras and multiple voltage and current sensors are arranged in the distribution station. The voltage and current sensors are used to collect the voltage and current of each power equipment. The fault location information is generated based on the voltage and current of the power equipment and the power topology map in the distribution station. In this way, fault question information is generated based on the fault location, and then the fault question information and the monitoring image are fused and input into the encoder to generate an image coding vector. The descriptive text most similar to the image coding vector is selected for output, and the fault type is determined based on the descriptive text to realize fault recognition based on multi-modal data of pictures and texts. The text data is determined based on the voltage and current, so the fault question text can be generated more accurately, thereby improving the accuracy of fault recognition. There is no need to manually input the question text, thereby improving the efficiency of fault handling.

[0008] In a possible implementation, determining whether there is a fault in the power equipment in the power distribution station according to the actual voltage and current of each power equipment and the power topology diagram of the power distribution station specifically includes: According to the power topology diagram of the distribution station and the input voltage of the distribution station, the theoretical voltage and current of each power equipment are calculated; For each electric device, the difference between the theoretical voltage and current and the actual voltage and current of the electric device is calculated. If the difference is greater than a preset threshold, it is determined that a fault exists.

[0009] In the above technical scheme, the theoretical voltage and current of each power equipment are calculated according to the power topology diagram of the distribution station and the input voltage of the distribution station, and the theoretical voltage and current of each power equipment are compared with the actual voltage and current. According to the comparison result, it is determined whether there is a fault. In this way, preliminary fault identification is achieved based on voltage and current, and a fault question text is generated based on the result of preliminary fault identification, so that the fault question text contains more fault information. Subsequently, fault type identification is performed based on the fault extraction text and the monitoring image, which can improve the accuracy of fault identification.

[0010] In a possible implementation manner, generating fault location information according to the voltage and current of each power device specifically includes: For each power device, if the difference of the power device is greater than a preset threshold, the power device is taken as a target power device; The power equipment adjacent to the target power equipment is obtained from the power topology map of the power distribution station, the target power equipment and the adjacent power equipment are regarded as suspicious power equipment, and the fault location information is generated according to the identification of the suspicious power equipment.

[0011] In the above technical scheme, if it is determined that the voltage difference and current difference of a certain power equipment are greater than the corresponding threshold values, a suspicious power equipment is determined based on the power equipment, and a fault question text is generated according to the information of the suspicious power equipment. In this way, the fault question text can contain more information, and then when the fault type is identified based on the fault question text and the monitoring image, the fault type can be identified more accurately.

[0012] In a possible implementation, generating a fault question text according to the fault location information, and acquiring a monitoring image taken by a camera according to the fault location information specifically includes: Acquire the voltage and current of the suspicious power equipment, and generate a fault question text for inquiring whether the suspicious power equipment has a fault according to the voltage and current of the suspicious power equipment; Obtain the camera identification corresponding to the suspicious power equipment, generate a shooting instruction according to the camera identification corresponding to the suspicious power equipment, use the shooting instruction to control the camera corresponding to the suspicious power equipment to shoot monitoring images, and receive the monitoring images sent back by the camera corresponding to the suspicious power equipment.

[0013] In the above technical solution, a fault question text is generated according to the voltage and current of the suspicious power equipment, so that the fault question text can contain more information, and then when the fault type is identified based on the fault question text and the monitoring image, the fault type can be identified more accurately. In addition, after the suspicious power equipment is determined, the camera corresponding to the suspicious power equipment is controlled to capture the monitoring image, and the camera does not need to continuously capture, thereby reducing the power consumption of the camera.

[0014] In a possible implementation, calculating the similarity between the image encoding vector and the text encoding vector of each description text in the preset description text set specifically includes: Extracting a classification coding vector corresponding to the classification fusion vector from the image coding vector; The similarity between the classification encoding vector corresponding to the classification fusion vector and the text encoding vector of each description text in the preset description text set is calculated.

[0015] In the above technical scheme, the input vector is obtained by concatenating the fusion vector of each image block and the classification fusion vector, and then the encoder encodes the input vector to output an encoding vector with the same dimension as the input vector, and extracts the classification-related components from the encoding vector, so that the target description text can be obtained based on the components.

[0016] In a possible implementation, before using the input layer to process the fault question text and the monitoring image to generate an input vector, the method further includes: Crawling training images and training description texts of each training image from the Internet, using the input layer to vectorize the training image to obtain the input vector of the training image, and using the image encoder to encode the input vector of the training image to obtain the encoded vector of the training image; Using a word vector conversion module to vectorize the training description text to obtain an input vector of the training description text, and using a text encoder to encode the input vector of the training description text to obtain an encoded vector of the training description text; For each coding vector of a training image, calculate the paired similarity between the coding vector of the training image and the coding vector of the corresponding training description text, and calculate the unpaired similarity between the coding vector of the training image and the coding vector of each other training description text; The loss value is calculated according to the paired similarity and the unpaired similarity, and the parameters of the image encoder, the parameters of the input layer, the parameters of the text encoder, and the parameters of the word vector conversion module are optimized according to the loss value.

[0017] In the above technical solution, training images and training description texts are used for model training, the parameters of the image encoder, the parameters of the input layer, the parameters of the text encoder, and the parameters of the word vector conversion module are optimized, and the most similar vectors are screened from the encoding vectors of multiple description texts based on the image encoding vector output by the image encoder to obtain the most matching description text and improve recognition accuracy.

[0018] In a possible implementation manner, after determining the fault type according to the target description text, the method further includes: Generate a maintenance work order based on the fault type and send the maintenance work order to the procurement system; enable the procurement system to generate a purchase list based on the maintenance work order.

[0019] In the above technical solution, after the background server determines the fault type of the faulty power equipment, a maintenance work order is generated according to the fault type, and the maintenance work order is sent to the procurement system, so that the procurement system generates a purchase list based on the maintenance work order, purchases the items required for maintenance in a timely manner, improves maintenance efficiency, and reduces the impact of power equipment failures in distribution stations on the powered areas.

[0020] The present application provides an object detection device based on image-text multimodal optimization, comprising: An acquisition module is used to acquire the actual voltage and current of each power device collected by the voltage and current sensors and the power topology diagram of the power distribution station; A processing module, used to determine whether there is a fault in the power equipment in the power distribution station according to the actual voltage and current of each power equipment and the power topology diagram of the power distribution station; The processing module is further used to generate fault location information according to the voltage and current of each power device, generate a fault question text according to the fault location information, and obtain a monitoring image taken by a camera according to the fault location information; The processing module is further used to use the input layer to process the fault question text and the monitoring image to generate an input vector; use the image encoder to encode the input vector to generate an image encoding vector; The processing module is also used to calculate the similarity between the image coding vector and the text coding vectors of each description text in the preset description text set; select the description text with the greatest similarity as the target description text output; and determine the fault type according to the target description text.

[0021] The present application provides a backend server, including: a memory, a processor; Memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes various possible implementations as described above.

[0022] The present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the various possible implementation methods described above.

[0023] The present application provides a computer program product, including a computer program, which implements various possible implementation modes as described above when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0025] Figure 1 A schematic diagram of an application scenario of a target detection method based on multimodal optimization of images and texts provided in this application; Figure 2 A flowchart of a target detection method based on multimodal optimization of images and texts provided in this application; Figure 3 A schematic diagram of the structure of the backend server provided for this application.

[0026] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0027] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0028] Traditional human-based inspection and management methods can no longer fully meet the needs of modern power distribution stations. In recent years, object detection models, as an important application in the field of artificial intelligence, have gradually been introduced into the management and maintenance process of power distribution stations.

[0029] However, target detection models usually only take image data as input and classify it after extracting features from the image data. Due to the single input data, it is impossible to accurately identify the fault type of power equipment in the distribution station.

[0030] The present application provides a target detection method based on multi-modal optimization of graphics and text. The voltage and current of each power equipment are collected through voltage and current sensors, and the fault location is determined based on the voltage and current. In this way, fault question information is generated based on the fault location, and then the fault question information and the monitoring image are fused and input into the encoder to generate an image coding vector. The descriptive text most similar to the image coding vector is selected for output, and the fault type is determined based on the descriptive text, so as to realize fault recognition based on multi-modal data of graphics and text. The text data is determined based on the voltage and current, so the fault question text can be generated more accurately, thereby improving the accuracy of fault recognition, and there is no need to manually input the question text, thereby improving the efficiency of fault handling.

[0031] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0032] Figure 1 A schematic diagram of an application scenario of a target detection method based on multimodal optimization of images and text provided in this application, such as Figure 1As shown, a plurality of power equipment are arranged in the power distribution station, and a voltage and current sensor is arranged on each power equipment, and the voltage and current sensor is used to detect the voltage and current of the corresponding power equipment. A camera is also arranged in the power distribution station, and the camera is used to capture monitoring images of the power equipment. Each camera is connected to the background server in communication, and the camera sends the captured monitoring images to the background server. Each voltage and current sensor collects the voltage and current of the corresponding power equipment, and sends the voltage and current of the power equipment to the background server. The background server performs fault identification based on the voltage and current of each power equipment collected by the voltage and current sensor and the monitoring images collected by the camera.

[0033] Figure 2 A flowchart of a target detection method based on multimodal optimization of images and texts provided in this application is shown in FIG. Figure 2 As shown, the method includes: S101. The backend server obtains the actual voltage and current of each power device collected by the voltage and current sensors and the power topology diagram of the power distribution station.

[0034] The voltage and current sensors are connected to the backend server in communication. The voltage and current sensors collect the actual voltage and current of the power equipment in real time and send the actual voltage and current of the power equipment to the backend server. The backend server extracts the power topology diagram of the power distribution station from the design data of the power distribution station.

[0035] S102: The backend server determines whether there is a fault in the power equipment in the power distribution station according to the actual voltage and current of each power equipment and the power topology diagram of the power distribution station. If yes, proceed to S103; if no, return to S101.

[0036] The theoretical voltage and current of each power device are determined based on the power topology diagram of the power distribution station, the actual voltage and current of each power device are compared with the theoretical voltage and current, and whether a fault occurs is determined based on the comparison result.

[0037] S103, generating fault location information according to the voltage and current of each power device, generating a fault question text according to the fault location information, and acquiring a monitoring image taken by a camera according to the fault location information.

[0038] If it is determined that a fault exists, the fault location information is determined based on the comparison results of the actual voltage and current of each power device with the theoretical voltage and current.

[0039] The fault location information is used to indicate suspicious power equipment that may have a fault. A fault question text is generated based on the fault location information so that the fault question text contains more fault information, so that when the fault type is identified based on the fault question text and the monitoring image, the fault identification accuracy can be improved.

[0040] The camera to be used for shooting is determined according to the fault location information, and the camera is controlled to shoot the monitoring image of the suspicious power equipment. The camera is connected to the background server for communication, and the camera sends the monitoring image of the suspicious power equipment to the background server.

[0041] S104. The backend server uses the input layer to process the fault question text and the monitoring image to generate an input vector, and uses the image encoder to encode the input vector to generate an image encoding vector.

[0042] Among them, the background server uses the input layer to vectorize the fault question text and the monitoring image respectively, and then fuses the two vectors to obtain the input vector.

[0043] The image encoder includes multiple encoding modules, each of which includes a multi-head attention mechanism module and a feedforward neural network module. The multi-head attention mechanism module and the feedforward neural network module can refer to the existing solution and will not be repeated here. The input vector is encoded by multiple encoding modules to generate an image encoding vector.

[0044] S105, the background server calculates the similarity between the image coding vector and the text coding vectors of each description text in the preset description text set, selects the description text with the greatest similarity as the target description text output, and determines the fault type according to the target description text.

[0045] Among them, for each description text in the preset description text set, the word vector conversion module is used to convert the description text into a text input vector, and the text encoder is used to transform the text input vector to generate a text encoding vector, and each description text in the description text set is traversed to obtain multiple text encoding vectors. For each text encoding two boxes, the similarity between the image encoding vector and the text encoding vector is calculated according to the cosine similarity calculation formula. The description text with the largest similarity is selected as the target description text output, and the fault type is determined according to the target description text.

[0046] For example, if the most similar description text is "This is an image of a switch cabinet with a short-circuit fault in the incoming cable", the fault type can be extracted from the target description text as a short-circuit fault in the incoming cable of the switch cabinet, so that the fault type of the power equipment can be determined based on the target description text.

[0047] In the above technical scheme, multiple cameras and multiple voltage and current sensors are arranged in the distribution station. The voltage and current sensors are used to collect the voltage and current of each power equipment. The fault location information is generated based on the voltage and current of the power equipment and the power topology map in the distribution station. In this way, fault question information is generated based on the fault location, and then the fault question information and the monitoring image are fused and input into the encoder to generate an image coding vector. The descriptive text most similar to the image coding vector is selected for output, and the fault type is determined based on the descriptive text to realize fault recognition based on multi-modal data of pictures and texts. The text data is determined based on the voltage and current, so the fault question text can be generated more accurately, thereby improving the accuracy of fault recognition. There is no need to manually input the question text, thereby improving the efficiency of fault handling.

[0048] Optionally, in the above scheme, by first acquiring the actual voltage and current data of each power device collected by the voltage and current sensor, combined with the power topology diagram of the distribution station, the operating status of the power device can be monitored comprehensively and in real time. When the voltage and current are abnormal, the fault location information is generated according to the abnormal data. This process ensures the accuracy of fault location and provides a reliable basis for subsequent fault identification. Based on the fault location information, the fault question text is automatically generated, and the monitoring image of the corresponding position is obtained. This process avoids manual intervention and realizes the automation of fault identification. The fault question text and the monitoring image are processed by the input layer to generate an input vector, and the input vector is encoded by the image encoder to generate an image encoding vector. This step makes full use of the advantages of multimodal data (text and image) and provides rich feature information for subsequent similarity calculation. Then, the similarity between the image encoding vector and the text encoding vector of each description text in the preset description text set is calculated, and the description text with the largest similarity is selected as the target description text output. This process quickly locks the description text that best matches the target fault through an efficient similarity matching algorithm. The fault type is then determined based on the target description text. This process not only improves the speed of fault identification, but also ensures the accuracy of the fault type, providing a clear direction for subsequent fault handling.

[0049] It is worth noting that since the above scheme automatically generates fault question text based on voltage and current data, no manual input is required, which significantly improves the efficiency of fault handling. At the same time, through the fusion and intelligent processing of multimodal data, the false alarm rate is reduced, and unnecessary operation and maintenance costs and time waste are reduced. The above scheme is not only suitable for power equipment fault detection in distribution stations, but can also be extended to other fields, such as intelligent transportation, industrial automation, etc., and has broad application prospects.

[0050] In a possible implementation, S102, the backend server determines whether there is a fault in the power equipment in the power distribution station according to the actual voltage and current of each power equipment and the power topology diagram of the power distribution station, specifically including: S201. The backend server calculates and obtains the theoretical voltage and current of each power device according to the power topology diagram of the power distribution station and the input voltage of the power distribution station.

[0051] Among them, as one of the implementation methods, a power model of the distribution station is built in the power modeling software based on the power topology diagram of the distribution station, the input voltage of the distribution station is added to the power model of the distribution station, and the theoretical voltage and current of each power equipment are calculated by the power modeling software.

[0052] As another implementation method, power parameters of each power device are extracted from the power topology diagram of the distribution station, the power parameters include resistance, capacitance or inductance, and the connection relationship of each power device is extracted from the power topology diagram of the distribution station. Based on the power parameters of each power device and the connection relationship of each power device, the theoretical voltage and current of each power device are calculated using the principle of electrical engineering.

[0053] S202. For each power device, the backend server calculates the difference between the theoretical voltage and current and the actual voltage and current of the power device. If the difference is greater than a preset threshold, it is determined that a fault exists.

[0054] The theoretical voltage and current include the theoretical voltage and the theoretical current, and the actual voltage and current include the actual voltage and the actual current. For each power device, the voltage difference between the theoretical voltage and the actual voltage of the power device is calculated, and the current difference between the theoretical current and the actual current of the power device is calculated. If the voltage difference is greater than the voltage threshold and the current difference is greater than the current threshold, it is determined that a fault occurs near the power device.

[0055] In the above technical scheme, the theoretical voltage and current of each power equipment are calculated according to the power topology diagram of the distribution station and the input voltage of the distribution station, and the theoretical voltage and current of each power equipment are compared with the actual voltage and current. According to the comparison result, it is determined whether there is a fault. In this way, preliminary fault identification is achieved based on voltage and current, and a fault question text is generated based on the result of preliminary fault identification, so that the fault question text contains more fault information. Subsequently, fault type identification is performed based on the fault extraction text and the monitoring image, which can improve the accuracy of fault identification.

[0056] Optionally, the theoretical voltage and current of each power device are calculated based on the power topology diagram and input voltage of the distribution station. This process ensures the accuracy of the theoretical value and provides a reliable benchmark for subsequent comparison. Then, by comparing the actual voltage and current of each power device with the theoretical voltage and current, when the difference exceeds the preset threshold, it is determined that the device has a fault, so that the potential fault point can be quickly identified, and the preliminary screening of the fault is achieved. In addition, based on the results of the preliminary fault identification, a question text containing rich fault information is automatically generated. Since the question text is generated based on the comparison results of the actual voltage and current with the theoretical value, it contains accurate fault location and possible fault type information, which provides strong support for subsequent fault type identification. In addition, by combining the fault question text with the monitoring image, multimodal data is used for fault type identification. Since the question text already contains preliminary fault information, the subsequent identification process can be more focused and accurate. At the same time, the fusion of multimodal data also further improves the accuracy and reliability of fault identification. Since the scheme can automatically and quickly identify the fault point and generate accurate fault information, the efficiency of fault handling is significantly improved. Operation and maintenance personnel can quickly locate problems based on the fault information provided by the system and take corresponding treatment measures, thereby reducing the impact of faults on the stable operation of the power system.

[0057] In a possible implementation, S103, the backend server generates fault location information according to the voltage and current of each power device, specifically including: S301. For each power device difference, if the power device difference is greater than a preset threshold, the backend server uses the power device as a target power device.

[0058] Among them, for each power device, the voltage difference between the theoretical voltage and the actual voltage of the power device is calculated, and the current difference between the theoretical current and the actual current of the power device is calculated. If the voltage difference is greater than the voltage threshold and the current difference is greater than the current threshold, the power device is taken as the target power device.

[0059] For example, there are four power equipments in a distribution station, which are marked as power equipment 1, power equipment 2, power equipment 3 and power equipment 4 respectively.

[0060] If the voltage difference of the power device 1 is greater than the preset voltage threshold, and the current difference of the power device 1 is greater than the preset current threshold, the power device 1 is determined to be the target power device.

[0061] If the voltage difference of the power device 2 is smaller than the preset voltage threshold, and the current difference of the power device 2 is smaller than the preset current threshold, it is determined that the power device 2 is not the target power device.

[0062] If the voltage difference of the power device 3 is greater than the preset voltage threshold, and the current difference of the power device 1 is greater than the preset current threshold, the power device 3 is determined to be the target power device.

[0063] If the voltage difference of the power device 4 is smaller than the preset voltage threshold, and the current difference of the power device 2 is smaller than the preset current threshold, it is determined that the power device 4 is not the target power device.

[0064] S302. The backend server obtains the power equipment adjacent to the target power equipment from the power topology map of the power distribution station, regards the target power equipment and the adjacent power equipment as suspicious power equipment, and generates fault location information according to the identification of the suspicious power equipment.

[0065] The identification of the target power equipment is compared with the identification of the power equipment in the power topology map of the power distribution station to determine the position of the target power equipment in the power topology map of the power distribution station.

[0066] According to the position of the target power equipment in the power topology map of the power distribution station, the power equipment adjacent to the target power equipment is obtained from the power topology map of the power distribution station, and the target power equipment and the adjacent power equipment are regarded as suspicious power equipment. Wherein, adjacent means that the power connection relationship is adjacent. For example: if power equipment 1 and power equipment 2 are electrically connected, then power equipment 1 is the power equipment adjacent to power equipment 2.

[0067] In the above technical scheme, if it is determined that the voltage difference and current difference of a certain power equipment are greater than the corresponding threshold values, a suspicious power equipment is determined based on the power equipment, and a fault question text is generated according to the information of the suspicious power equipment. In this way, the fault question text can contain more information, and then when the fault type is identified based on the fault question text and the monitoring image, the fault type can be identified more accurately.

[0068] Optionally, when both the voltage difference and current difference of a certain power device exceed preset thresholds, the device is immediately marked as a target power device, i.e., a potential fault point. This determination method based on voltage and current differences can quickly and accurately locate the source of the fault, providing a clear direction for subsequent processing. The power devices adjacent to the target power device are obtained from the power topology map of the distribution station, and these devices are considered as suspicious power devices together with the target power device. This step takes into account the interrelationships in the power system. Even if the fault may be caused by a single device, adjacent devices may be affected or reflect signs of fault. Therefore, including adjacent devices in the analysis scope helps to understand the fault situation more comprehensively.

[0069] Then, the fault location information is generated based on the identification of the suspected power equipment. This information not only includes the specific location of the target power equipment, but also covers the information of its adjacent equipment, forming a detailed description of the fault area. This detailed fault location information provides rich context for subsequent fault type identification, which helps to improve the accuracy of identification.

[0070] Based on the fault location information generated above, question texts containing more fault details can be constructed. These question texts not only reflect the fault status of the target power equipment, but also associate the potential impact of adjacent equipment, so that the subsequent multimodal fault type identification process based on text and images can obtain more comprehensive information input.

[0071] In the fault type identification stage, by combining detailed fault question text and monitoring images and taking advantage of multimodal data fusion, the fault type can be identified more accurately. Since the question text already contains rich fault location information, it helps to narrow the search space for fault identification and improve the accuracy and efficiency of identification.

[0072] In a possible implementation, S103, generating a fault question text according to the fault location information, and acquiring a monitoring image captured by a camera according to the fault location information, specifically includes: S401. Obtain voltage and current of a suspicious power device, and generate a fault question text for inquiring whether the suspicious power device has a fault according to the voltage and current of the suspicious power device.

[0073] The identification of the suspicious power equipment is extracted from the fault location information, and the voltage and current of each suspicious power equipment are obtained. For each suspicious power equipment, the possible fault type of the suspicious power equipment is determined according to the voltage of the suspicious power equipment. A fault question is generated according to the possible fault type of the suspicious power equipment.

[0074] For example, if it is determined that a short circuit fault occurs in the suspicious power equipment based on the voltage and current of the suspicious power equipment, a fault question text is generated to inquire whether a short circuit fault occurs in the suspicious power equipment.

[0075] S402, generating a shooting instruction according to the identifier of the camera corresponding to the suspicious power equipment, the shooting instruction controlling the camera corresponding to the suspicious power equipment to shoot a monitoring image, and receiving the monitoring image sent back by the camera corresponding to the suspicious power equipment.

[0076] Among them, after the suspicious power equipment is determined, the correspondence between the power equipment and the camera is determined according to the installation data of the power equipment and the camera, the camera corresponding to the suspicious power equipment is determined according to the correspondence, and a shooting instruction is generated according to the identification of the camera corresponding to the suspicious power equipment. The camera corresponding to the suspicious power equipment is controlled to capture the monitoring image of the suspicious power equipment, and the monitoring image of the suspicious power equipment is transmitted back to the background server.

[0077] In the above technical solution, a fault question text is generated according to the voltage and current of the suspicious power equipment, so that the fault question text can contain more information, and then when the fault type is identified based on the fault question text and the monitoring image, the fault type can be identified more accurately. In addition, after the suspicious power equipment is determined, the camera corresponding to the suspicious power equipment is controlled to capture the monitoring image, and the camera does not need to continuously capture, thereby reducing the power consumption of the camera.

[0078] Optionally, based on the voltage and current data of the suspicious power equipment, targeted fault question texts are generated. These texts not only reflect the current status of the suspicious power equipment (such as voltage anomalies, current fluctuations, etc.), but also imply possible fault types, providing valuable information input for subsequent multimodal fault identification. Since the fault question text is generated based on the actual voltage and current data, it contains more specific and accurate fault information, which helps to narrow the scope of fault identification and improve the accuracy of identification. After the suspicious power equipment is identified, a shooting instruction is generated according to its corresponding camera identification to control the camera to capture the monitoring image. This on-demand shooting method avoids continuous shooting of the camera, effectively reduces the power consumption of the camera, and extends its service life.

[0079] At the same time, since the shooting instructions are generated for specific suspicious power equipment, the acquired monitoring images are highly correlated with the fault location, providing intuitive and effective visual information for subsequent fault type identification. The combination of fault question text and monitoring images for fault type identification fully utilizes the advantages of multimodal data. The fault question text provides detailed fault information, while the monitoring image provides an intuitive on-site situation. The two complement each other and jointly improve the accuracy of fault type identification. By combining text information with image information, the system can understand the fault situation more comprehensively, thereby making a more accurate fault type judgment.

[0080] Since the fault question text and monitoring image acquisition are based on the fault location information, the system can quickly locate the suspicious power equipment and quickly obtain relevant data for fault identification. This targeted data processing method improves the response speed and efficiency of the system. In addition, by reducing unnecessary camera shooting and data transmission, the overall processing burden of the system is also reduced, further improving the system's operating efficiency.

[0081] It can be seen that the above technical solution not only improves the accuracy of fault identification, but also fully considers the issues of power consumption management and resource optimization. By taking monitoring images on demand, the camera is avoided from working continuously, effectively reducing the energy consumption of the system. This design concept of balancing power consumption management and resource optimization enables the system to achieve a greener and more sustainable operation while ensuring performance.

[0082] In a possible implementation, S104, the backend server uses the input layer to process the fault question text and the monitoring image to generate an input vector, specifically including: S501. The backend server uses a target detection model to perform target detection on the monitoring images taken by the camera corresponding to the suspicious power equipment to obtain the type and location of the power equipment.

[0083] Among them, the target detection model can choose Faster R-CNN, YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector) and other network models, which can quickly and accurately identify and locate target objects in surveillance images.

[0084] S502: The backend server extracts a region of interest from the monitoring image according to the location of the power equipment, and uses the input layer to process the fault question text and the region of interest to generate an input vector.

[0085] Among them, the region of interest algorithm is used to cut out the minimum circumscribed rectangular area where the suspicious power equipment is located. The input vector is generated by processing the fault question text and the region of interest.

[0086] In the above technical solution, the location of the suspicious power equipment is identified from the monitoring image of the suspicious power equipment by using the target detection model, and the region of interest extraction algorithm is used to extract the region of interest based on the location of the suspicious power equipment. By removing other pixel areas, the amount of data in the subsequent processing process is reduced, and the processing efficiency and recognition accuracy are improved.

[0087] Optionally, by processing the monitoring images using advanced target detection models, the type and location of the power equipment can be accurately identified. This step provides accurate target coordinates for subsequent extraction of the region of interest, ensuring that the extracted information is highly relevant to the suspicious power equipment. In addition, the use of the target detection model enables the system to automatically and quickly locate key information in the monitoring images, avoiding the tediousness and errors of manual intervention and improving processing efficiency.

[0088] Based on the power equipment location information provided by the target detection model, the monitoring image is processed using the region of interest extraction algorithm to accurately extract the image area related to the suspicious power equipment. This process effectively removes irrelevant pixel areas in the image and significantly reduces the amount of data for subsequent processing. Through the extracted information, the system can focus on processing image information directly related to fault identification, reduce noise interference, and improve identification accuracy.

[0089] Since the monitoring images have been preprocessed by target detection and information extraction before the input layer, the amount of data that needs to be processed by the input layer is greatly reduced, thereby reducing the computational complexity and improving the processing efficiency. At the same time, since the input vector only contains information that is highly relevant to fault identification (i.e., fault question text and area of ​​interest), the subsequent fault type identification process can be more focused and accurate, further improving the accuracy of identification.

[0090] In addition, by reducing the amount of data in the processing process, this technical solution not only improves processing efficiency, but also optimizes the use of system resources. This optimization is particularly important when resources such as computing power and storage are limited. In addition, due to the improvement of processing efficiency and the optimized use of resources, the overall operating cost of the system is also reduced accordingly, bringing economic benefits to the operation and maintenance management of power companies.

[0091] Some other embodiments of the present application provide a method for target detection based on multimodal optimization of images and texts, the method comprising: S601. The backend server obtains the actual voltage and current of each power device collected by the voltage and current sensors.

[0092] S602: The backend server determines whether there is a fault in the power equipment in the power distribution station according to the actual voltage and current of each power equipment and the power topology diagram of the power distribution station. If yes, proceed to S603; if no, return to S601.

[0093] S603, generating fault location information according to the voltage and current of each power device, generating a fault question text according to the fault location information, and acquiring a monitoring image taken by a camera according to the fault location information.

[0094] S604: The backend server uses the input layer to process the fault question text and the monitoring image to generate an input vector, and uses an image encoder to encode the input vector to generate an image encoding vector.

[0095] S605: The backend server calculates the similarity between the image coding vector and the text coding vectors of each description text in the preset description text set, selects the description text with the greatest similarity as the target description text output, and determines the fault type according to the target description text.

[0096] S606. The backend server generates a maintenance work order according to the fault type, and sends the maintenance work order to the procurement system; the procurement system generates a purchase list according to the maintenance work order.

[0097] After determining the fault type, the backend server can generate a maintenance work order based on the fault type, so that the maintenance personnel's maintenance terminal can repair the faulty power equipment after receiving the maintenance work order. In addition, the backend server also sends the maintenance work order to the procurement system, so that the procurement system determines the items to be purchased based on the fault type in the maintenance work order and generates a purchase list based on the items to be purchased.

[0098] In the above technical solution, after the background server determines the fault type of the faulty power equipment, a maintenance work order is generated according to the fault type, and the maintenance work order is sent to the procurement system, so that the procurement system generates a purchase list based on the maintenance work order, purchases the items required for maintenance in a timely manner, improves maintenance efficiency, and reduces the impact of power equipment failures in distribution stations on the powered areas.

[0099] Optionally, by automatically generating maintenance work orders and sending them to the procurement system, maintenance efficiency is significantly improved. In traditional methods, the generation of maintenance work orders and the preparation of procurement lists often require manual intervention, which is not only time-consuming and labor-intensive, but also prone to errors. This technical solution realizes automated processing, reduces manual intervention, and thus speeds up the progress of maintenance work.

[0100] Procurement of maintenance items in a timely manner is the key to ensuring smooth maintenance work. The above technical solution automates and standardizes the procurement process by enabling the procurement system to automatically generate a procurement list based on the maintenance work order, ensuring timely procurement and supply of maintenance items. This helps shorten the maintenance cycle and reduce the impact of power equipment failure on the power supply area.

[0101] In addition, the failure of power equipment often has different degrees of impact on the power supply area. Through the above technical solution, it is possible to quickly respond to the failure and start the maintenance process, and restore the normal operation of the power equipment in time, thereby improving the stability of power supply. In addition, automated processing reduces the possibility of manual intervention and errors, and reduces operation and maintenance costs. At the same time, by timely purchasing and supplying the items required for maintenance, maintenance delays and additional costs caused by shortages of items are avoided.

[0102] In a possible implementation, S502, using the input layer to process the fault question text and the region of interest to generate an input vector, specifically includes: S701, dividing the region of interest into blocks to obtain multiple image blocks, using the input layer to process the multiple image blocks to obtain the content vector of each image block; using the input layer to process the position of each image block to obtain the position vector of each image block.

[0103] For example, if the region of interest is a rectangular pixel region, the rectangular pixel region is divided into 16 image blocks. Each row includes 4 image blocks, each column includes 4 image blocks, and the positions of the 16 image blocks are marked as position 1, position 2, position 3, ..., position 16.

[0104] The input layer includes a content conversion layer and a position conversion layer. The content conversion layer is used to perform vector conversion processing on each image block to obtain the content vector of each image block. More specifically, for each block, multiple convolution kernels are used to perform multiple convolution processing on the image block to obtain the content vector of the image block. The content vector of each image block is obtained by traversing all image blocks. The parameters of the convolution kernel can be obtained through training.

[0105] The position conversion layer is used to convert the position of each image block into a vector to obtain the position vector of each image block. More specifically, a mapping function is used to map position 1, position 2, position 3, ..., position 16 into a position vector to obtain the position vector of each image block. The mapping function can also be obtained through training.

[0106] S702: Use the input layer to process the fault question text to generate a classification vector, and use the input layer to process the position of the classification vector to output a position vector corresponding to the classification vector.

[0107] The input layer also includes a word vector conversion module, which is used to convert the fault question text into a vector to obtain a classification vector. The parameters in the word vector conversion module can also be obtained through training. The position of the classification vector is initialized to position 0, and a mapping function is used to convert position 0 into a position vector to obtain a position vector corresponding to the classification component.

[0108] S703 . Generate an input vector according to the content vector of each image block, the position vector of each image block, the classification vector, and the position vector corresponding to the classification vector.

[0109] The content vector of each image block, the position vector of each image block, the classification vector, and the position vector corresponding to the classification vector are fused and concatenated to obtain an input vector.

[0110] In the above technical solution, by dividing the region of interest into multiple image blocks, each image block and the position of each image block are converted into a vector, the fault question text is converted into a classification vector, a position vector corresponding to the classification vector is generated, and an input vector is generated based on the above vector. In this way, the region of interest and the fault question text are converted into an input vector, and a classification vector is generated based on the fault question text, so that the vector contains more information, thereby improving the accuracy of fault identification.

[0111] Optionally, by dividing the region of interest into multiple image blocks and processing each image block and its position separately, more detailed information about the power equipment fault can be captured. The content vector of each image block reflects the characteristics of the local area, while the position vector provides spatial layout information. The combination of the two provides more comprehensive data support for subsequent fault identification. At the same time, the fault question text is processed to generate a classification vector, and the position vector corresponding to the classification vector is generated, further enriching the information dimension of the input vector. The classification vector represents the theme or category of the fault question text, while the position vector reflects the position or importance of the text in the overall processing flow, all of which help to improve the accuracy of fault identification.

[0112] Converting the region of interest and fault question text into vector representation makes the subsequent processing more flexible and efficient. Vector representation has the advantages of unified dimension and simple calculation, which is convenient for processing and analysis in machine learning models such as neural networks. In addition, vector representation is also convenient for similarity calculation, cluster analysis and other operations, providing more technical means and possibilities for fault type identification.

[0113] The input vector is generated based on the content vector, position vector, classification vector and the position vector corresponding to the classification vector of each image block, thus realizing the effective fusion of multi-source information. This fusion method not only retains the key features of the original data, but also generates new information through the combination and transformation of vectors, providing more powerful data support for fault identification. Through information fusion, the system can more comprehensively understand the situation of power equipment faults, including the type, location, severity and other aspects of the fault, so as to make more accurate fault identification judgments.

[0114] Although the above steps add the steps of fine processing of the region of interest and fault question text in the preprocessing stage, this increased processing complexity is effectively compensated in the subsequent stages. Since the input vector contains richer information, the subsequent fault identification model can converge to the correct result more quickly during processing, thereby improving the overall processing efficiency. At the same time, through refined processing and information fusion, the system can more accurately identify the fault area and type, reduce unnecessary false positives and false negatives, and further improve the efficiency of resource utilization.

[0115] In a possible implementation, S703, generating an input vector according to the content vector of each image block, the position vector of each image block, the classification vector, and the position vector corresponding to the classification vector, specifically includes: S801. Superimpose the content vector and the position vector of each image block to obtain a fusion vector of each image block, and superimpose the classification vector and the position vector corresponding to the classification vector to obtain a classification fusion vector.

[0116] The dimension of the content vector is the same as the dimension of the position vector. For each image block, the sum of the content vector and the position unit vector of the image block is calculated to obtain the fusion vector of each image block. The dimension of the classification vector is the same as the dimension of the position vector corresponding to the classification vector. The sum of the classification vector and the position vector corresponding to the classification vector is calculated to obtain the classification fusion vector.

[0117] For example, if there are 16 image blocks, the sum of the content vector and the position vector of the first image block is calculated to obtain the fusion vector of the first image block. The sum of the content vector and the position vector of the second image block is calculated to obtain the fusion vector of the second image block. And so on, 16 fusion vectors are obtained.

[0118] S802: Concatenate the fusion vector of each image block and the classification fusion vector to obtain an input vector.

[0119] Among them, after obtaining 16 fusion vectors, vector splicing is performed in the order of classification fusion vector, fusion vector of the first image block, fusion vector of the second image block, fusion vector of the third image block, ..., fusion vector of the sixteenth image block to obtain an input vector.

[0120] In the above technical solution, the content vector and position vector of each image block are fused so that the fused vector contains content and position information, the classification vector and the position vector corresponding to the classification vector are superimposed, and multiple vectors obtained by superposition calculation are concatenated and encoded to generate an input vector, so that the input vector contains content information, position information and classification information, thereby improving recognition accuracy.

[0121] Optionally, the content vector and position vector of each image block are superimposed to obtain the fusion vector of each image block. This step is crucial because it closely combines the specific content of the image block (such as texture, color and other features) with its position information in the image. Position information is crucial for understanding the contextual relationship in the image and judging the spatial distribution of the target. By fusing the two, the fusion vector not only contains the detailed features of the image block, but also its spatial position information, providing more comprehensive data support for subsequent target detection.

[0122] Next, the classification vector and the position vector corresponding to the classification vector are superimposed to obtain the classification fusion vector. The classification vector usually represents an abstract representation of the category to which the image block or the entire image belongs, while the position information of the classification vector indicates the approximate distribution of these categories in the image. By superimposing the classification vector and the position vector, the classification fusion vector can simultaneously capture the category information and its spatial layout in the image, providing strong support for identifying the target category and its position in the image.

[0123] Finally, the fusion vector and classification fusion vector of each image block are concatenated to generate the final input vector. This step combines the detailed information at the image block level with the category and position information at the image level to form a comprehensive representation that contains both local details and global information. The input vector therefore contains rich content information, location information, and classification information, providing comprehensive and accurate data input for the subsequent object detection model.

[0124] Through the implementation of the above technical solutions, the input vector contains richer and more comprehensive information, which enables the target detection model to understand the image content more accurately and identify the target object. Content information helps the model distinguish different targets, location information helps the model determine the spatial location of the target, and classification information provides contextual support at the image level. These factors work together to significantly improve the accuracy of target detection.

[0125] In a possible implementation, S105, the backend server calculates the similarity between the image coding vector and the text coding vector of each description text in the preset description text set, specifically including: S901. The background server extracts a classification coding vector corresponding to the classification fusion vector from the image coding vector.

[0126] The input vector is obtained by concatenating the classification fusion vector, the fusion vector of the first image block, the fusion vector of the second image block, the fusion vector of the third image block, ..., the fusion vector of the sixteenth image block. The dimension of the classification fusion vector is the same as that of the fusion vector of each image block. For example, if the dimension of the fusion vector of each image block is 512, the size of the input vector is 512×17.

[0127] After the encoder encodes the input vector, the size of the generated encoding vector is the same as the size of the input vector, that is, the size of the encoding vector is 512 × 17. The first 512 elements are extracted from the encoding vector as the classification encoding vector corresponding to the classification fusion vector.

[0128] S902: The backend server calculates the similarity between the classification encoding vector corresponding to the classification fusion vector and the text encoding vector of each description text in the preset description text set.

[0129] Among them, for each description text in the preset description text set, the cosine similarity calculation formula is used to calculate the similarity between the classification encoding vector corresponding to the classification fusion vector and the text encoding vector of the description text.

[0130] In the above technical scheme, the input vector is obtained by concatenating the fusion vector of each image block and the classification fusion vector, and then the encoder encodes the input vector to output an encoding vector with the same dimension as the input vector, and extracts the classification-related components from the encoding vector, so that the target description text can be obtained based on the components.

[0131] Optionally, by extracting classification-related components from the image encoding vector and calculating similarity with the text encoding vector, we can more accurately match the description text that matches the image content, effectively avoiding the information missing problem that may exist in traditional single-modality target detection methods and improving the accuracy of target description.

[0132] In practical applications, image and text data are often affected by many factors (such as illumination changes, occlusion, text ambiguity, etc.). By combining multimodal information and performing similarity calculations, we can overcome the interference caused by these factors to a certain extent and enhance the robustness of target detection.

[0133] In the above technical solution, the input vector is obtained by concatenating the fusion vector of each image block and the classification fusion vector, and then the encoder performs encoding processing, thereby achieving efficient feature extraction and fusion. This method avoids complex preprocessing steps and redundant calculation processes, and optimizes calculation efficiency.

[0134] Some other embodiments of the present application provide a method for target detection based on multimodal optimization of images and texts, the method comprising: S1001. Crawling training images and training description texts of each training image from the Internet, using an input layer to vectorize the training image to obtain an input vector of the training image, and using an image encoder to encode the input vector of the training image to obtain an encoded vector of the training image.

[0135] Here, crawler software is used to crawl training images and description text of each training image from the Internet, for example: crawling an image of a transformer and a description text of the image "This is an oil-immersed transformer".

[0136] The training image is divided into blocks to obtain multiple training image blocks, and each training image block is converted into a vector using a content conversion layer to obtain a content vector of each training image block. More specifically, multiple convolution kernels are used to perform multiple convolution processing on each training image block to obtain a content vector of each training image block.

[0137] The position conversion layer is used to convert the position of each training image block into a position vector. More specifically, the position of the training image block is represented as position 1, position 2, position 3, ..., and a mapping function is used to convert the position of each training image block into a position vector to obtain the position vector of each training image block.

[0138] Calculate the sum of the content vector and the position vector of each training image block to obtain the fusion vector of each training image block. Randomly initialize the training classification vector. Take position 0 as the position of the training classification vector, use the mapping function to convert position 0 to the position vector, and obtain the position vector corresponding to the training classification vector. Calculate the sum of the classification vector and the position vector corresponding to the classification vector to obtain the training classification fusion vector.

[0139] The training classification fusion vector and the fusion vector of each training image block are concatenated to obtain the input vector of the training image. The training image vector is encoded using an image encoder to obtain the encoding vector of the training image.

[0140] S1002. Use a word vector conversion module to vectorize the training description text to obtain an input vector of the training description text, and use a text encoder to encode the input vector of the training description text to obtain a text encoding vector.

[0141] Among them, the word vector conversion module can use an existing module, which will not be repeated here. The text encoder includes multiple encoding modules, each of which includes a multi-head attention mechanism module and a feedforward neural network module. The multi-head attention mechanism module and the feedforward neural network module can refer to the existing solution, which will not be repeated here. The input vector of the training description text is encoded by multiple encoding modules to generate the encoding vector of the training description text.

[0142] S1003. For each coding vector of a training image, calculate the paired similarity between the coding vector of the training image and the coding vector of the corresponding training description text, and calculate the unpaired similarity between the coding vector of the training image and the coding vector of each other training description text.

[0143] The training data pair includes training images and training description texts. Three pairs of training data pairs are used as examples below. The first set of training data pairs is training image 1 and training description text 1, the second set of training data pairs is training image 2 and training description text 2, and the third set of training data pairs is training image 3 and training description text 3.

[0144] For training image 1, the cosine similarity between the encoding vector of the training image generated based on training image 1 and the encoding vector of the training description text generated based on training description text 1 is calculated as the matching similarity. The cosine similarity between the encoding vector of the training image generated based on training image 1 and the encoding vector of the training description text generated based on training description text 2 is calculated as the non-matching similarity. The cosine similarity between the encoding vector of the training image generated based on training image 1 and the encoding vector of the training description text generated based on training description text 3 is calculated as the non-matching similarity.

[0145] For training image 2, the cosine similarity between the encoding vector of the training image generated based on training image 2 and the encoding vector of the training description text generated based on training description text 2 is calculated as the matching similarity. The cosine similarity between the encoding vector of the training image generated based on training image 2 and the encoding vector of the training description text generated based on training description text 1 is calculated as the non-matching similarity. The cosine similarity between the encoding vector of the training image generated based on training image 2 and the encoding vector of the training description text generated based on training description text 3 is calculated as the non-matching similarity.

[0146] For training image 3, the cosine similarity between the encoding vector of the training image generated based on training image 3 and the encoding vector of the training description text generated based on training description text 3 is calculated as the matching similarity. The cosine similarity between the encoding vector of the training image generated based on training image 3 and the encoding vector of the training description text generated based on training description text 1 is calculated as the non-matching similarity. The cosine similarity between the encoding vector of the training image generated based on training image 3 and the encoding vector of the training description text generated based on training description text 2 is calculated as the non-matching similarity.

[0147] S1004, calculating a loss value according to the unpaired similarity and the paired similarity, and adjusting parameters of the image encoder, parameters of the input layer, parameters of the text encoder, and parameters of the word vector conversion module according to the loss value.

[0148] Among them, the loss value is calculated according to the unpaired similarity and paired similarity, and the parameters of the image encoder, the input layer, the text encoder, and the word vector conversion module are adjusted in reverse order to minimize the loss function until the trained input layer and image encoder are obtained after the preset number of trainings. The loss value is calculated using the contrast loss function, so that the paired similarity can be maximized and the unpaired similarity can be minimized.

[0149] In the above technical solution, training images and training description texts are used for model training, the parameters of the image encoder, the parameters of the input layer, the parameters of the text encoder, and the parameters of the word vector conversion module are optimized, and the most similar vectors are screened from the encoding vectors of multiple description texts based on the image encoding vector output by the image encoder to obtain the most matching description text and improve recognition accuracy.

[0150] Optionally, the model is trained using training images and training description texts, and the relevant parameters are optimized so that the image encoding vector output by the image encoder can more accurately match the encoding vector of the description text. This training method enhances the model's ability to understand the correlation between images and texts, and improves recognition accuracy.

[0151] By calculating paired similarity and unpaired similarity, the model not only learns how to identify description text that matches the image, but also learns how to distinguish irrelevant description text. This training method enhances the generalization ability of the model, enabling the model to better match and identify when faced with new data.

[0152] By minimizing the loss value to optimize the model parameters, the model is constantly adjusted and improved during the training process. This optimization process is based on a large amount of training data, so it is possible to find a better combination of parameters and improve the overall performance of the model.

[0153] In addition, this technical solution supports learning in two modes: image and text, providing an effective way to implement multimodal learning. By combining image and text information, the model can understand complex scenes more comprehensively, providing a broader space for subsequent applications.

[0154] In summary, the above technical solution uses training images and training description texts to train the model, optimizes the parameters of the image encoder, the parameters of the input layer, the parameters of the text encoder, and the parameters of the word vector conversion module, thereby achieving a more accurate match between images and texts, improving the matching accuracy and generalization ability of the model, optimizing the model parameters, and supporting multimodal learning.

[0155] In addition, the embodiment of the present application also provides a target detection device based on image-text multimodal optimization, including: An acquisition module is used to acquire the actual voltage and current of each power device collected by the voltage and current sensors and the power topology diagram of the power distribution station; A processing module, used to determine whether there is a fault in the power equipment in the power distribution station according to the actual voltage and current of each power equipment and the power topology diagram of the power distribution station; The processing module is further used to generate fault location information according to the voltage and current of each power device, generate a fault question text according to the fault location information, and obtain a monitoring image taken by a camera according to the fault location information; The processing module is further used to use the input layer to process the fault question text and the monitoring image to generate an input vector; use the image encoder to encode the input vector to generate an image encoding vector; The processing module is also used to calculate the similarity between the image coding vector and the text coding vectors of each description text in the preset description text set; select the description text with the greatest similarity as the target description text output; and determine the fault type according to the target description text.

[0156] Figure 3 A schematic diagram of the structure of the backend server provided for this application. Figure 3As shown, the background server 110 provided in this embodiment includes: at least one processor 1101 and a memory 1102. Optionally, the device 110 further includes a communication component 1103. The processor 1101, the memory 1102 and the communication component 1103 are connected via a bus 1104.

[0157] In a specific implementation process, at least one processor 1101 executes the computer-executable instructions stored in the memory 1102, so that at least one processor 1101 executes the above method.

[0158] The specific implementation process of the processor 1101 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0159] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the invention can be directly implemented as a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0160] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.

[0161] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0162] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0163] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0164] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0165] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0166] The division of units is only a logical function division, and there may be other divisions in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0167] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0168] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0169] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0170] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0171] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A target detection method based on image-text multimodal optimization, characterized in that: The target detection method is applied to a background server, the background server is in communication connection with a camera located in a power distribution station, and voltage and current sensors are arranged on each power equipment in the power distribution station. The method includes: Acquire the actual voltage and current of each power device and the power topology diagram of the power distribution station collected by the voltage and current sensor; Determine whether there is a fault in the power equipment in the power distribution station according to the actual voltage and current of each power equipment and the power topology diagram of the power distribution station; If yes, generate fault location information according to the voltage and current of each power device, generate a fault question text according to the fault location information, and obtain a monitoring image taken by the camera according to the fault location information; Using an input layer to process the fault question text and the monitoring image to generate an input vector; using an image encoder to encode the input vector to generate an image encoding vector; Calculate the similarity between the image coding vector and the text coding vectors of each description text in the preset description text set; select the description text with the greatest similarity as the target description text output; and determine the fault type according to the target description text.

2. The target detection method according to claim 1, characterized in that: Determining whether there is a fault in the power equipment in the power distribution station according to the actual voltage and current of each power equipment and the power topology diagram of the power distribution station specifically includes: Calculate and obtain the theoretical voltage and current of each power device according to the power topology diagram of the power distribution station and the input voltage of the power distribution station; For each electric device, the difference between the theoretical voltage and current and the actual voltage and current of the electric device is calculated. If the difference is greater than a preset threshold, it is determined that a fault exists.

3. The target detection method according to claim 2, characterized in that: Generating fault location information according to the voltage and current of each power device specifically includes: For each power device, if the difference of the power device is greater than a preset threshold, the power device is used as a target power device; The power equipment adjacent to the target power equipment is acquired from the power topology map of the power distribution station, the target power equipment and the adjacent power equipment are regarded as suspicious power equipment, and the fault location information is generated according to the identification of the suspicious power equipment.

4. The target detection method according to claim 3, characterized in that: Generating a fault question text according to the fault location information, and acquiring a monitoring image taken by the camera according to the fault location information, specifically includes: Acquiring the voltage and current of the suspicious electric power equipment, and generating a fault question text for inquiring whether the suspicious electric power equipment has a fault according to the voltage and current of the suspicious electric power equipment; Obtain the camera identification corresponding to the suspicious power equipment, generate a shooting instruction according to the camera identification corresponding to the suspicious power equipment, the shooting instruction controls the camera corresponding to the suspicious power equipment to capture a monitoring image, and receive the monitoring image sent back by the camera corresponding to the suspicious power equipment.

5. The target detection method according to any one of claims 1 to 4, characterized in that: Calculating the similarity between the image encoding vector and the text encoding vectors of each description text in the preset description text set specifically includes: Extracting a classification coding vector corresponding to the classification fusion vector from the image coding vector; The similarity between the classification encoding vector corresponding to the classification fusion vector and the text encoding vector of each description text in the preset description text set is calculated.

6. The target detection method according to any one of claims 1 to 4, characterized in that: Before using the input layer to process the fault question text and the monitoring image to generate an input vector, the method further includes: Crawling training images and training description text of each training image from the Internet, using the input layer to vectorize the training image to obtain an input vector of the training image, and using the image encoder to encode the input vector of the training image to obtain an encoded vector of the training image; Using a word vector conversion module to vectorize the training description text to obtain an input vector of the training description text, and using a text encoder to encode the input vector of the training description text to obtain an encoded vector of the training description text; For each coding vector of a training image, calculating the paired similarity between the coding vector of the training image and the coding vector of the corresponding training description text, and calculating the unpaired similarity between the coding vector of the training image and the coding vector of each other training description text; A loss value is calculated according to the paired similarity and the unpaired similarity, and parameters of the image encoder, parameters of the input layer, parameters of the text encoder, and parameters of the word vector conversion module are optimized according to the loss value.

7. The target detection method according to any one of claims 1 to 4, characterized in that: After determining the fault type according to the target description text, the method further includes: A maintenance work order is generated according to the fault type, and the maintenance work order is sent to a procurement system; and the procurement system generates a purchase list according to the maintenance work order.

8. A target detection device based on image-text multimodal optimization, characterized in that: include: An acquisition module is used to acquire the actual voltage and current of each power device collected by the voltage and current sensors and the power topology diagram of the power distribution station; A processing module, used to determine whether there is a fault in the power equipment in the power distribution station according to the actual voltage and current of each power equipment and the power topology diagram of the power distribution station; The processing module is further used to generate fault location information according to the voltage and current of each power device, generate a fault question text according to the fault location information, and obtain a monitoring image taken by a camera according to the fault location information; The processing module is further used to use the input layer to process the fault question text and the monitoring image to generate an input vector; use the image encoder to encode the input vector to generate an image encoding vector; The processing module is also used to calculate the similarity between the image coding vector and the text coding vectors of each description text in the preset description text set; select the description text with the greatest similarity as the target description text output; and determine the fault type according to the target description text.

9. An electronic device, characterized in that: include: processor; as well as, A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.