Government affairs inspection method and device, storage medium and electronic equipment
By combining target image data and multimodal data from the government knowledge base, accurate government inspection results are generated, solving the problem of insufficient inspection accuracy in existing technologies and achieving highly accurate government inspection and anomaly cause analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AEROSPACE AGE LOW AERIAL TECHNOLOGY CO LTD
- Filing Date
- 2026-04-17
- Publication Date
- 2026-07-21
AI Technical Summary
Existing government inspection technologies are inaccurate, especially when distinguishing between normal industrial heat dissipation and abnormal overheating, resulting in a high false alarm rate and low inspection accuracy.
By acquiring target image data and spatiotemporal labels of candidate targets, and combining them with inspection object indication data in the government knowledge base, the government inspection results are generated using a multimodal target model, and a standardized report is generated.
It improves the accuracy of government inspections, reduces the false alarm rate, and can infer the causes of anomalies through multimodal reasoning chains, generating standardized reports that directly support government decision-making.
Smart Images

Figure CN122046174B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of anomaly detection technology, and in particular to a government inspection method, device, storage medium, and electronic device. Background Technology
[0002] Currently, with the integration of flight equipment (such as drones) and artificial intelligence technology, their application in the field of government inspection is becoming increasingly sophisticated. However, these technologies typically rely on pixel-level physical quantity detection for inspection, such as temperature threshold alarms. This makes it difficult to understand the underlying business implications of these physical quantities, such as distinguishing between "normal industrial heat dissipation" and "abnormal overheating," leading to a high false alarm rate and consequently low accuracy in government inspections. Therefore, there is currently no satisfactory solution to improve the accuracy of government inspections. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method, apparatus, storage medium, and electronic device for government inspections to solve the problems of low accuracy in government inspections caused by related technologies. In other words, embodiments of the present invention can place physical observation targets in a precise three-dimensional spatial context through multimodal data such as target image data and target government knowledge text, thereby obtaining government inspection results with higher accuracy and reducing false alarm rate, thus effectively improving the accuracy of government inspections. Moreover, embodiments of the present invention can not only detect physical anomalies (i.e., anomaly detection conclusions), but also infer their causes through multimodal reasoning chains, thereby transforming candidate target detection data into standardized government inspection reports that can directly support government decision-making.
[0004] According to one aspect of the present invention, a method for government inspection is provided, the method comprising: Acquire candidate target detection data, which includes target image data and spatiotemporal labels of the candidate targets; wherein, the spatiotemporal labels of the targets include the location information corresponding to the candidate targets; Based on the target spatiotemporal label, the target government knowledge text corresponding to the candidate target is obtained from the government knowledge base; wherein, the target government knowledge text includes at least one inspection object indication data, and an inspection object indication data includes at least one of the following: the location information, geographic entity data, environmental dynamic data and historical data of an inspection object; The target multimodal big model is invoked, and based on the target image data and the target government affairs knowledge text, the government affairs inspection results of the candidate targets are generated. The government affairs inspection results include a multimodal inference chain and anomaly detection conclusions. Based on the results of the government inspection, a standardized government inspection report is generated; and the standardized government inspection report is sent to the regulatory object corresponding to the candidate target.
[0005] According to another aspect of the present invention, a government affairs inspection device is provided, the device comprising: An acquisition unit is used to acquire candidate target detection data, the candidate target detection data including target image data and target spatiotemporal labels of the candidate targets; wherein, the target spatiotemporal labels include the location information corresponding to the candidate targets; The processing unit is configured to obtain the target government knowledge text corresponding to the candidate target from the government knowledge base based on the target spatiotemporal label; wherein, the target government knowledge text includes at least one patrol object indication data, and the patrol object indication data includes at least one of the following: the location information, geographic entity data, environmental dynamic data and historical data of a patrol object; The processing unit is also used to call the target multimodal large model, and generate the government inspection results of the candidate target based on the target image data and the target government knowledge text. The government inspection results include a multimodal inference chain and anomaly detection conclusions. The processing unit is further configured to generate a standardized government inspection report based on the government inspection results; and send the standardized government inspection report to the regulatory object corresponding to the candidate target.
[0006] According to another aspect of the present invention, an electronic device is provided, the electronic device including a processor and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to perform the methods mentioned above.
[0007] According to another aspect of the present invention, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods mentioned above.
[0008] This invention provides embodiments that can acquire candidate target detection data, including target image data and spatiotemporal tags for the candidate targets; wherein the spatiotemporal tags include location information corresponding to the candidate targets. Then, based on the target spatiotemporal tags, target government knowledge text corresponding to the candidate targets can be obtained from a government knowledge base; wherein the target government knowledge text includes at least one inspection object indication data, and one inspection object indication data includes at least one of the following: location information, geographic entity data, environmental dynamic data, and historical data of an inspection object. Based on this, a target multimodal large model can be invoked to generate government inspection results for the candidate targets based on the target image data and target government knowledge text. The government inspection results include a multimodal inference chain and anomaly detection conclusions; and a standardized government inspection report can be generated based on the government inspection results; and the standardized government inspection report can be sent to the regulatory object corresponding to the candidate targets. As can be seen, the embodiments of the present invention can place the physical observation target in a precise three-dimensional spatial context by using multimodal data such as target image data and target government affairs knowledge text, thereby obtaining government affairs inspection results with higher accuracy, reducing the false alarm rate, and thus effectively improving the accuracy of government affairs inspection. Furthermore, the embodiments of the present invention can not only detect physical anomalies (i.e., anomaly detection conclusions), but also infer their causes through multimodal reasoning chains, thereby transforming candidate target detection data into standardized government affairs inspection reports that can directly support government affairs decision-making. Attached Figure Description
[0009] Further details, features, and advantages of the invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 A flowchart illustrating a government affairs inspection method according to an exemplary embodiment of the present invention is shown; Figure 2 A schematic diagram of a government affairs inspection system according to an exemplary embodiment of the present invention is shown; Figure 3 A flowchart illustrating another government inspection method according to an exemplary embodiment of the present invention is shown; Figure 4 A schematic block diagram of a government affairs inspection device according to an exemplary embodiment of the present invention is shown; Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation
[0010] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.
[0011] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0012] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0013] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0014] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0015] It should be noted that the executing entity of the government inspection method provided in this embodiment of the invention can be one or more electronic devices, and this invention does not limit this; wherein, the electronic device can be a terminal (i.e., a client) or a server. Therefore, when the executing entity includes multiple electronic devices, and among these multiple electronic devices are at least one terminal and at least one server, the government inspection method provided in this embodiment of the invention can be jointly executed by the terminal and the server. Accordingly, the terminal mentioned herein may include, but is not limited to: smartphones, tablets, laptops, desktop computers, smartwatches, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The server mentioned herein can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc.
[0016] Optionally, embodiments of the present invention may also provide a government inspection system. In this case, the executing entity of the government inspection method proposed in the embodiments of the present invention may also be the government inspection system, that is, the electronic device used to execute the government inspection method may be one or more electronic devices in the government inspection system, etc.; embodiments of the present invention do not limit this.
[0017] Optionally, the government inspection method provided in this embodiment of the invention can be applied to any government inspection scenario (such as any low-altitude government inspection scenario), and this embodiment of the invention does not limit it. For example, it can be applied to the safety inspection scenario of power facilities, such as the candidate target including any power facility, and the inspection object including enterprises, communities, etc.; or, it can also be applied to the illegal burning behavior identification scenario, such as the candidate target including but not limited to at least one: straw accumulation traces, fire source (i.e., fire point), etc., and the inspection object including land ownership, etc.; or, it can also be applied to the forest fire prevention monitoring scenario, such as the candidate target including but not limited to at least one of the following: fire source, forest, etc., and the inspection object including forest or forest management department, etc.; or, it can also be applied to the river sewage discharge supervision scenario, such as the candidate target including any drainage outlet, and the inspection object including enterprises, communities, etc.; or, it can also be applied to the emergency search and rescue assistance scenario, such as the candidate target including but not limited to at least one of the following: trapped personnel, disaster-stricken facilities, etc., and the inspection object including locations (such as villages), geological disaster-prone areas, etc., etc.
[0018] Based on the above description, this embodiment of the invention proposes a government affairs inspection method, which can be executed by the aforementioned electronic device (terminal or server); or, the government affairs inspection method can be executed jointly by a terminal and a server, etc. For ease of explanation, the following description will use the execution of this government affairs inspection method by an electronic device as an example; such as Figure 1 As shown, this government inspection method may include the following steps S101-S104: S101, acquire candidate target detection data, which includes target image data and target spatiotemporal labels of the candidate targets; wherein, the target spatiotemporal labels include the location information corresponding to the candidate targets.
[0019] Optionally, the aforementioned candidate targets can be any government inspection targets, such as power facilities, drainage outlets, etc., and this embodiment of the invention does not limit this. Optionally, the target image data of a candidate target can be an image block of the corresponding candidate region cropped out, and one candidate target can correspond to one candidate region, that is, one candidate target can be a target in one candidate region.
[0020] Optionally, the location information (also known as geographic location tag) corresponding to the candidate target may include, but is not limited to, at least one of the following: GPS (Global Positioning System) coordinates (such as latitude and longitude, which may be in the WGS-84 coordinate system) and relative altitude of the area where the candidate target is located, from which image data was collected. This embodiment of the invention does not limit this. Optionally, the spatiotemporal tag of the target may also include, but is not limited to, at least one of the following: timestamp (i.e., collection time) of the image data collected in the area where the candidate target is located, flight speed, attitude angles (such as pitch, roll, and yaw, angles used to control the aircraft's attitude), gimbal angle (angles used to control the camera lens), etc. This embodiment of the invention does not limit this. Based on this, this embodiment of the invention can achieve unified timestamp and / or geographic location tag binding through the target spatiotemporal tag.
[0021] Optionally, the methods for obtaining candidate target detection data may include, but are not limited to, the following: The first acquisition method: The electronic device stores at least one target detection data in its own storage space. In this case, the electronic device can take each target detection data in the at least one target detection data as a candidate target detection data to acquire the candidate target detection data.
[0022] The second acquisition method involves an electronic device acquiring region-specific image data, which may include an initial visible light acquisition image and an initial infrared acquisition image. The initial visible light acquisition image can then undergo visible light image preprocessing to obtain a target visible light acquisition image; similarly, the initial infrared acquisition image can undergo infrared image preprocessing to obtain a target infrared acquisition image. Furthermore, using the target visible light acquisition image as a reference, the target infrared acquisition image can be pixel-level aligned to obtain an aligned infrared acquisition image. Based on this, target detection can be performed on the target visible light acquisition image and the aligned infrared acquisition image to obtain target detection results; and based on these results, candidate target detection data can be acquired. Optionally, the initial visible light and initial infrared images can be dual-light images acquired in the target area by a government inspection and acquisition device (such as a government inspection drone, also known as a flying device) equipped with a synchronously triggered high-definition visible light camera and an infrared thermal imaging camera (such as the DJI Zenmuse H30 series, a camera that provides thermal imaging capabilities), to construct regional image data. Based on this, after constructing the regional image data, the government inspection and acquisition device can send the regional image data to an electronic device so that the electronic device can receive the regional image data, thereby acquiring the regional image data. A single visible light image can be a single visible light image, that is, an image consistent with human vision; correspondingly, a single infrared image can be a thermal infrared image, that is, an image formed by receiving infrared electromagnetic radiation emitted or reflected by the object itself, and then performing photoelectric conversion and signal processing. Optionally, in other embodiments, the electronic device may store at least one image data to be processed in its own storage space. In this case, the electronic device may collect each image data to be processed as a region image data to acquire region image data, etc.; the embodiments of the present invention do not limit this.
[0023] Optionally, government inspection and data collection equipment (such as government inspection drones via flight control systems) can output flight parameter data in real time while collecting dual-light images (i.e., initial visible light image and initial infrared image). This means that one area of image data can correspond to one set of flight parameter data, also known as the flight parameter data corresponding to the area of image data. Optionally, a set of flight parameter data may include, but is not limited to, at least one of the following: GPS coordinates (latitude and longitude), relative altitude, flight speed, attitude angles (pitch, roll, yaw), gimbal angle, and other low-altitude flight parameters. This embodiment of the invention does not limit these parameters. For example, all area image data can be assigned a unified timestamp (i.e., collection time) and a geographic location tag (which can be determined through flight parameter data) at the moment of collection. Optionally, the area acquisition image data may also include, but is not limited to, at least one of the following: acquisition time (i.e., the time when the area acquisition image data is acquired, which may also be referred to as the time when the dual-light image in the area acquisition image data is acquired) and flight parameter data (i.e., the flight parameter data when the area acquisition image data is acquired, which may also be referred to as the flight parameter data when the dual-light image in the area acquisition image data is acquired); based on this, one dual-light image can correspond to one acquisition time and one flight parameter data, thereby forming a spatiotemporally aligned raw data stream.
[0024] It should be noted that the specific implementation process of visible light image preprocessing in this embodiment of the invention is not limited. For example, when performing visible light image preprocessing on the initial visible light acquisition image to obtain the target visible light acquisition image, white balance correction can be performed on the initial visible light acquisition image to eliminate color shift caused by atmospheric scattering, thereby obtaining a white balance-corrected initial visible light acquisition image, which can then be used as the target visible light acquisition image; alternatively, contrast adaptive enhancement can be performed on the white balance-corrected initial visible light acquisition image to improve the visibility of details in weak texture areas, thereby obtaining a contrast adaptively enhanced initial visible light acquisition image, which can then be used as the target visible light acquisition image, and so on.
[0025] Accordingly, the specific implementation process of infrared image preprocessing in this embodiment of the invention is not limited. For example, when performing infrared image preprocessing on the initial infrared acquisition image to obtain the target infrared acquisition image, the initial infrared acquisition image can be subjected to dynamic range normalization processing to obtain a normalized initial infrared acquisition image, which can be used as the target infrared acquisition image. In this case, it can adaptively adapt to the temperature range of different scenes; or, the normalized initial infrared acquisition image can be subjected to noise suppression processing to obtain the target infrared acquisition image, and so on. Optionally, the initial infrared image can be calibrated temperature data output by an infrared thermal imaging camera. Correspondingly, when performing dynamic range normalization on the initial infrared image to obtain a normalized initial infrared image, the maximum and minimum temperatures can be determined from the initial infrared image. Based on the determined maximum and minimum temperatures, the initial infrared image can be normalized to normalize the temperature of each pixel in the initial infrared image, thereby achieving dynamic range normalization of the initial infrared image to obtain a normalized initial infrared image.
[0026] Optionally, when performing pixel-level alignment of the target infrared image with the target visible light image as a reference to obtain an aligned infrared image, the electronic device can extract feature points from both the target visible light image and the target infrared image using the SIFT (Scale-invariant feature transform) algorithm. This yields a visible light feature point data set for the target visible light image and an infrared feature point data set for the target infrared image. Each feature point data set can include the position of a feature point (e.g., coordinate position, i.e., pixel position) and a feature vector (also called a descriptor). Then, based on the visible light feature point data set and the infrared feature point data set, a feature point matching pair set is determined (a feature point matching pair can also be called a pair of feature points), and a perspective transformation matrix can be determined based on the feature point matching pair set. Based on this, the target infrared image can be aligned to the visible light coordinate system based on the perspective transformation matrix. In other words, pixel-level alignment of the target infrared image can be performed based on the perspective transformation matrix to obtain the aligned infrared image. A feature point matching pair may include a visible light feature point (i.e., a feature point in the visible light feature point dataset) and an infrared feature point (i.e., a feature point in the infrared feature point dataset). Optionally, when determining the feature point matching pair set based on the visible light feature point dataset and the infrared feature point dataset, the distances of all feature point matching pairs formed by the visible light feature point dataset and the infrared feature point dataset (which may include feature point matching pairs formed by each visible light feature point and each infrared feature point respectively) can be calculated separately. The distance of a feature point matching pair can be the distance between the feature vector of the visible light feature point and the feature vector of the infrared feature point in the corresponding feature point matching pair (such as Euclidean distance or cosine distance, etc.), which can also be called the distance between feature point matching pairs; and feature point matching pairs with a distance less than a distance threshold are added to the feature point matching pair set; or, for For any visible light feature point in the visible light feature point dataset, the two closest infrared feature points to that visible light feature point can be determined from the infrared feature point dataset. If the ratio of the distance between any visible light feature point and its closest infrared feature point to the distance between any visible light feature point and its second closest infrared feature point (i.e., the other infrared feature point among the two closest infrared feature points) is less than a preset ratio threshold, then any visible light feature point and its closest infrared feature point can be considered as a feature point matching pair in the feature point matching pair set, and so on. This embodiment of the invention does not limit this. Optionally, both the distance threshold and the preset ratio threshold can be set based on experience or actual needs, and this embodiment of the invention does not limit this.
[0027] Optionally, when determining the perspective transformation matrix based on a set of feature point matching pairs, the set of feature point matching pairs can be used to construct perspective transformation constraints (i.e., establish two linear equations for each feature point matching pair), and the perspective transformation constraints can be solved using singular value decomposition to obtain the perspective transformation matrix; alternatively, four feature point matching pairs can be randomly selected from the set of feature point matching pairs, and a temporary perspective transformation matrix can be calculated using singular value decomposition on the selected feature point matching pairs. This temporary perspective transformation matrix can then be used to calculate the reprojection error between the projected positions and actual positions of all feature point matching pairs in the set (e.g., using...). A temporary perspective transformation matrix is used to project the positions of infrared feature points in a feature point matching pair to projected positions (the actual positions can be the positions of visible light feature points in the corresponding feature point matching pair). Feature point matching pairs with reprojection errors less than a preset error threshold are recorded as inliers. Sampling and calculation are repeated, and finally, the feature point matching pair set is updated using all inliers with the maximum number of inliers (i.e., the largest number of inliers in a single sampling). This triggers the execution of constructing perspective transformation constraints using the feature point matching pair set, and solving the perspective transformation constraints through singular value decomposition to obtain the perspective transformation matrix, etc. This embodiment of the invention does not limit this. Optionally, the perspective transformation matrix can be a 3×3 matrix.
[0028] Optionally, when performing pixel-level alignment of the target infrared acquisition image based on the perspective transformation matrix to obtain an aligned infrared acquisition image, for any pixel in the aligned infrared acquisition image, the inverse of the perspective transformation matrix can be determined, and the pixel coordinates of any pixel in the target infrared acquisition image can be determined using the inverse of the perspective transformation matrix. Thus, the pixel value of any pixel can be determined according to its pixel coordinates in the target infrared acquisition image. For example, the pixel value at the pixel coordinates in the target infrared acquisition image can be used as the pixel value of any pixel, and so on. Based on this, embodiments of the present invention can achieve pixel-level alignment of the target infrared acquisition image with the target visible light acquisition image as a reference to obtain an aligned infrared acquisition image through feature point matching alignment.
[0029] Optionally, in other embodiments, a deep learning alignment method based on spatial transformation networks can be used to perform pixel-level alignment of the target infrared image with the target visible light acquisition image as a reference, resulting in an aligned infrared image; alternatively, a pixel-by-pixel distortion alignment method based on dense optical flow estimation can be used to perform pixel-level alignment of the target infrared image with the target visible light acquisition image as a reference, and so on; the present invention does not limit this. Based on this, embodiments of the present invention can accurately align the target infrared acquisition image to the visible light coordinate system, so that the pixels in the aligned infrared acquisition image correspond one-to-one with the pixels in the target visible light acquisition image, thereby achieving dual-light image registration.
[0030] Optionally, when performing target detection on the target visible light image and the aligned infrared image to obtain the target detection result, a target detection model can be used to perform target detection on the target visible light image and the aligned infrared image to obtain the target detection result. For example, the target visible light image and the aligned infrared image can be concatenated (e.g., 3 visible light channels + 1 infrared channel = 4 channel input) to obtain a model input tensor, which can then be input into the target detection model to output the target detection result. Optionally, the target detection result may include, but is not limited to, at least one of the following: a target bounding box containing any candidate region, category labels, and confidence scores, etc., which are not limited in this embodiment of the invention. Optionally, the target detection model can be any detection model, which is not limited in this embodiment of the invention. For example, a lightweight real-time object detection model (such as an RT-DETR (Real-Time Detection Transformer) model with approximately 20M parameters) can be deployed in the edge computing layer. The Transformer can be a self-attention mechanism. Alternatively, the object detection model can be a detection model accelerated by TensorRT quantization (a model inference optimization technique) to reduce model accuracy and effectively reduce inference latency, such as ensuring inference latency of less than 50ms on the edge computing layer. Based on this, the object detection model can extract multi-scale visual features through the backbone network, capture global contextual information using the Transformer encoder (to correlate infrared thermal features with visible light texture features), and then generate prediction results (i.e., object detection results) through the decoder.
[0031] In one implementation, when acquiring candidate target detection data based on target detection results, a preset confidence threshold can be determined, and background regions with confidence scores less than or equal to the preset confidence threshold can be filtered out to obtain candidate targets (i.e., targets indicated by candidate regions with confidence scores greater than the preset confidence threshold, or targets contained in the corresponding candidate regions). Then, target image data of the candidate targets (i.e., image blocks of candidate regions where candidate targets are located) can be cropped from the target visible light acquisition image and the aligned infrared acquisition image, respectively, to add the target image data of the candidate targets to the candidate target detection data, and the target spatiotemporal labels of the candidate targets can be added to the candidate target detection data. Optionally, the preset confidence threshold can be set according to experience or actual needs, and this embodiment of the invention does not limit this. Optionally, when there are multiple candidate targets, multiple candidate targets can be cropped separately to obtain candidate target detection data corresponding to each candidate target.
[0032] In another embodiment, the electronic device can also acquire multiple flight parameter data within a target time range, where the multiple flight parameter data are time-series data within the target time range; wherein, the acquisition time of the regional image data is within the target time range. Optionally, the target time range can be divided according to a preset duration; optionally, the preset duration can be set according to experience or actual needs, and this embodiment of the invention does not limit this; for example, the preset duration can be 2 seconds, in which case the continuous flight parameter data can be segmented, that is, every 2 seconds of data can be sliced into 1 feature vector, and so on. Based on this, the multiple flight parameter data within the target time range can include flight parameter data at each acquisition time (i.e., the time of each acquisition of dual-light images) within the target time range. Then, time-series feature extraction can be performed on the multiple flight parameter data to obtain flight parameter time-series features; for example, time-series feature extraction can be performed on the multiple flight parameter data through a lightweight LSTM (Long Short-Term Memory) encoder to obtain flight parameter time-series features, thereby encoding them into feature vectors reflecting the observation scale and platform stability, and so on.
[0033] Therefore, when obtaining candidate target detection data based on target detection results, candidate target detection data can be obtained based on both the target detection results and the temporal features of flight parameters. Correspondingly, the target image data and spatiotemporal labels of the candidate targets can be added to the candidate target detection data based on the target detection results, and the temporal features of flight parameters can be added to the candidate target detection data, so that the temporal features of flight parameters serve as the flight parameter feature vectors corresponding to the candidate targets, and so on. In this case, the candidate target detection data also includes the flight parameter feature vectors corresponding to the candidate targets.
[0034] Optionally, if no candidate target is detected (i.e., the confidence scores of all candidate regions are less than or equal to a preset confidence threshold), the electronic device may also upload the flight log data corresponding to the regional image data to a monitoring system or the cloud, etc., which is not limited in this embodiment of the present invention; Optionally, the flight log data may include, but is not limited to, at least one of the following: thumbnails of the target visible light acquisition image and / or aligned infrared acquisition image, flight parameter data corresponding to the regional image data (i.e., flight parameter data acquired when acquiring regional image data), etc., which is not limited in this embodiment of the present invention.
[0035] The third acquisition method: Electronic devices can receive candidate target detection data sent by the edge computing layer to acquire candidate target detection data. For example, the cloud intelligence layer in the electronic device can acquire candidate target detection data when it receives candidate target detection data sent by the edge computing layer, and so on. Optionally, the government inspection system may include, but is not limited to, at least one of the following: a low-altitude perception layer (also known as the edge side, which can be located on flight equipment, i.e., it can be realized through flight equipment, which can realize data acquisition through camera equipment (such as cameras) and flight control systems, such as image acquisition and flight parameter data acquisition), an edge computing layer (also known as the edge side or edge-side), and a cloud intelligence layer (also known as the cloud side), such as... Figure 2 As shown; optionally, the edge computing layer and the cloud intelligence layer may be located on the same electronic device or on different electronic devices, and the embodiments of the present invention do not limit this; for example, both the edge computing layer and the cloud intelligence layer may be located on the electronic device that provides cloud services, or the edge computing layer may be located on a lightweight electronic device (such as a laptop) while the cloud intelligence layer may be located on a large electronic device (such as a cloud server), and so on.
[0036] Optionally, in other embodiments, the preprocessing, pixel-level alignment, and target detection of the regional image data can also be performed in real time through the edge computing layer; alternatively, in other embodiments, the edge computing layer can also be located in the airborne or ground station edge computing unit, in which case a large amount of background data without abnormalities can be excluded through lightweight inference on the edge side, thereby significantly reducing cloud bandwidth pressure and computing load, etc.; the present invention does not limit this.
[0037] S102, based on the target spatiotemporal label, obtain the target government knowledge text corresponding to the candidate target from the government knowledge base; wherein, the target government knowledge text includes at least one inspection object indication data, and the inspection object indication data includes at least one of the following: the location information, geographic entity data, environmental dynamic data and historical data of an inspection object.
[0038] Optionally, the government knowledge base may include a knowledge base for any government inspection scenario; that is, the specific content of the government knowledge base is not limited in this embodiment of the invention. For example, the government knowledge base may include inspection object instruction data for all inspection objects within a specified area. The specified area may be set according to experience or actual needs, and this embodiment of the invention does not limit this. Optionally, an inspection object may be an enterprise or legal person, a community, a forest or forest management department, etc.; this embodiment of the invention does not limit this. Optionally, inspection object instruction data may be stored in the government knowledge base in structured text form or unstructured text form; this embodiment of the invention does not limit this.
[0039] It should be noted that the data (i.e., government knowledge) in the government knowledge base can be obtained through legal and compliant channels from regulatory departments (i.e., government departments) and / or inspection targets, that is, it can be obtained with authorization from regulatory departments and / or inspection targets, and the embodiments of this invention only use the obtained government knowledge for government inspections and not for any other purpose; optionally, the data in the government knowledge base can also be stored after declassification processing, and / or through encrypted storage, etc. Based on this, the embodiments of this invention can achieve the legal and compliant construction of the government knowledge base through legal sharing of government data, hierarchical classification and desensitization, full-process security control and usage restriction, etc. Accordingly, the target government knowledge text corresponding to the candidate targets obtained from the government knowledge base is authorized data, that is, it is obtained under legal and compliant conditions.
[0040] Optionally, the electronic device can obtain the target government knowledge text corresponding to the candidate target from the government knowledge base based on the location information corresponding to the candidate target and / or the timestamp in the target spatiotemporal tag. For example, the electronic device can determine the patrol object indication data of at least one patrol object located within the target area from the government knowledge base based on the location information corresponding to the candidate target, thereby obtaining at least one patrol object indication data and thus obtaining the target government knowledge text corresponding to the candidate target from the government knowledge base. Optionally, the target area can refer to the area centered on the target location (such as GPS coordinates) indicated by the location information corresponding to the candidate target. This embodiment of the invention does not limit the shape and size of the target area. For example, the target area can be a circular area with the target location as its center and a preset radius, or it can be a square area with the target location as its center and a preset side length, etc. Optionally, the preset radius and preset side length can be set according to experience or actual needs, and this embodiment of the invention does not limit them.
[0041] Optionally, the electronic device can use the location information of candidate target objects to determine the patrol object indication data of at least one patrol object located within the target area from the government knowledge base; and / or, it can also use the timestamp in the target spatiotemporal label (i.e., the timestamp corresponding to the candidate target) to obtain the real-time environmental data (such as real-time wind direction, wind speed, etc. under the corresponding timestamp) corresponding to each patrol object in at least one patrol object, so as to add the real-time environmental data corresponding to each patrol object to the patrol object indication data of the corresponding patrol object, thereby serving as the environmental dynamic data of the corresponding patrol object, and so on. For example, at this time, the patrol object indication data of all patrol objects can be stored in the government knowledge base in structured text form.
[0042] Optionally, when the inspection target indication data is stored in the knowledge database in the form of unstructured text, the government knowledge base may include, but is not limited to, at least one of the following: preset spatial geographic information, preset inspection target information, environmental dynamic data, etc., which are not limited in this embodiment of the invention; for example, the preset spatial geographic information may include, but is not limited to, at least one of the following: high-precision GIS (Geographic Information System) map, three-dimensional real scene model, underground pipeline network map and building structure map within a specified area, etc., the preset inspection target information may include the inspection target information of all inspection targets within a specified area (such as the inspection target information of an inspection target may include, but is not limited to, at least one of the following: the equipment list of the corresponding inspection target (such as the list of power equipment, etc.), historical safety supervision records (such as whether there are historical high temperature warnings, accidents, inspection records, etc.), the introduction information of the inspection target, etc.), the environmental dynamic data may include, but is not limited to, real-time wind speed, wind direction, etc.; optionally, the government knowledge base may be connected to a meteorological interface to obtain real-time wind speed and wind direction through the meteorological interface for subsequent risk diffusion simulation. Based on this, electronic devices can also obtain target spatial geographic information, target patrol object information, and target environmental dynamic data corresponding to candidate targets from the government knowledge base based on target spatiotemporal tags (such as location information and / or timestamps corresponding to candidate targets). They can then perform structured transformation on the target spatial geographic information, target patrol object information, and target environmental dynamic data to obtain at least one patrol object indication data. This allows for the retrieval of target government knowledge text corresponding to candidate targets from the government knowledge base, and so on. Specifically, target spatial geographic information may include spatial geographic information within the target area (such as GIS maps, 3D reality models, etc.), target patrol object information may include patrol object information for each patrol object within the target area, and target environmental dynamic data may include real-time environmental data within the target area at the timestamp corresponding to the candidate target (e.g., real-time environmental data corresponding to each patrol object within the target area, or real-time environmental data within each grid area within the target area).
[0043] Optionally, electronic devices can directly call the meteorological interface to obtain real-time environmental data. For example, they can obtain the target spatial geographic information and target inspection object information corresponding to the candidate target from the government knowledge base, and call the meteorological interface to obtain target environmental dynamic data; or, they can access the meteorological interface through the government knowledge base to obtain real-time environmental data, and in this case, they can obtain target environmental dynamic data from the government knowledge base, etc. The embodiments of the present invention do not limit this.
[0044] Optionally, the location information of a patrol object can be used to indicate the location of the corresponding patrol object. It should be noted that the specific representation of the location information of the patrol object in this embodiment of the invention is not limited; it can be coordinates, or it can be a text description (e.g., 200m (meters) on xx road, etc.). Optionally, the specific content of the geographic entity data of a patrol object in this embodiment of the invention is not limited. For example, when a patrol object is an enterprise and it is applied to a power facility safety patrol scenario, the geographic entity data of a patrol object may include, but is not limited to, the object identifier of the corresponding patrol object (e.g., object name or object number, such as enterprise name), the included equipment (e.g., power equipment, etc.); or, when applied to a forest fire monitoring scenario, and a patrol object is a forest, the geographic entity data of a patrol object may include, but is not limited to, the object identifier of the corresponding patrol object (e.g., xx forest, etc.), vegetation type, etc. Optionally, the environmental dynamic data of a patrol object may include real-time environmental data of the grid area where the corresponding patrol object is located, or it may also include real-time environmental data of the monitoring station closest to the corresponding patrol object, etc.; this embodiment of the invention does not limit this. Optionally, the historical data of a patrol object may include any historical record, such as "2025-12-01 Temperature 45℃ (degrees Celsius)" etc. For example, the indication data of a patrol object may be: [Location] 200m on XX Road, [Company] Company A, [Equipment] 110kV (kilovolt) cable, [History] 2025-12-01 Temperature 45℃, etc.
[0045] S103, invoke the target multimodal large model, based on the target image data and target government affairs knowledge text, to generate government affairs inspection results for candidate targets. The government affairs inspection results include multimodal inference chains and anomaly detection conclusions.
[0046] It should be noted that the embodiments of the present invention do not limit the specific model structure of the multimodal large model (such as the target multimodal large model). That is to say, a multimodal large model can be any multimodal large model. Optionally, a multimodal large model can also be called a multimodal fusion model.
[0047] Optionally, the target image data may include the target visible light image and the target infrared image of the candidate target. A multimodal large model includes a visual feature encoder (also called a visual encoder), that is, the target multimodal large model may include a visual feature encoder. Based on this, when calling the target multimodal large model to generate the government inspection results of the candidate target based on the target image data and the target government knowledge text, the electronic device may call the visual feature encoder in the target multimodal large model to extract features from the target visible light image and the target infrared image respectively, to obtain the image features of the target visible light image and the image features of the target infrared image; wherein, the visual feature encoder includes a visible light adaptation layer and an infrared adaptation layer, the image features of the target visible light image are extracted through the visible light adaptation layer, and the image features of the target infrared image are extracted through the infrared adaptation layer. In this embodiment of the invention, the visible light adaptation layer supports the enhancement of edge and texture region responses (that is, it can be used to enhance the edge and texture region responses of the visible light image), and the infrared adaptation layer supports the enhancement of feature weights for regions with significant temperature gradients (that is, it can be used to enhance feature weights for regions with significant temperature gradients in the infrared image). Optionally, the visual feature encoder may also include a visual shared backbone network (such as a ViT (Vision Transformer) backbone network). Based on this, the electronic device can call the visual shared backbone network in the target multimodal large model to extract features from the target visible light image and the target infrared image, thereby obtaining the initial image features of the target visible light image and the initial image features of the target infrared image. It can also call the visible light adaptation layer in the target multimodal large model to extract features from the initial image features of the target visible light image, thereby obtaining the image features of the target visible light image. Furthermore, it can call the infrared adaptation layer in the target multimodal large model to extract features from the initial image features of the target infrared image, thereby obtaining the image features of the target infrared image. This allows the general visual features (i.e., the initial image features) to be distributed and aligned through the modality-specific attention adaptation layers (i.e., the visible light adaptation layer and the infrared adaptation layer).
[0048] Optionally, when calling the visual shared backbone network in the target multimodal large model to extract features from the target visible light image and the target infrared image to obtain the initial image features of the target visible light image and the target infrared image, the visual shared backbone network in the target multimodal large model can be called separately to extract features from the target visible light image and the target infrared image to obtain the initial image features of the target visible light image and the target infrared image; or, the target visible light image and the target infrared image can be stitched together to obtain a stitched image, and the visual shared backbone network in the target multimodal large model can be called to extract features from the stitched image to obtain the initial image features of the target visible light image and the target infrared image, etc.; the embodiments of the present invention do not limit this.
[0049] Then, the electronic device can determine the target fused visual feature vector based on the image features of the target visible light image and the target infrared image. For example, the electronic device can fuse the image features of the target visible light image and the target infrared image (e.g., feature concatenation of the two image features, or fusion of the two image features through at least one fully connected layer, etc.) to obtain an initial fused visual feature vector. It can then call the visual text projection layer (also called the cross-modal projection layer) in the target multimodal large model to project the initial fused visual feature vector onto the text encoding space to obtain the target fused visual feature vector; or, it can call the visual text projection layer in the target multimodal large model to project the image features of the target visible light image and the target infrared image onto the text encoding space respectively to obtain the projection features of the target visible light image and the projection features of the target infrared image, and fuse the projection features of the target visible light image and the projection features of the target infrared image to obtain the target fused visual feature vector, etc. This embodiment of the invention does not limit the scope of the invention. The dimension of the target fusion visual feature vector is the same as the dimension of the text feature vector output by the text encoder, that is, the dimension of the target fusion visual feature vector is the same as the dimension of the text feature vector of the target government knowledge text.
[0050] In this embodiment of the invention, a multimodal large model may further include a text encoder, such as a target multimodal large model that may include a text encoder. Accordingly, the electronic device can invoke the text encoder in the target multimodal large model to extract features from the target government knowledge text, obtaining a text feature vector of the target government knowledge text. For example, the target government knowledge text can be input into the text encoder in the target multimodal large model, thereby outputting the text feature vector of the target government knowledge text. The text feature vector can also be called a semantic feature vector.
[0051] Optionally, the electronic device can also acquire candidate region description information corresponding to the candidate target (such as structured description information of the candidate region, which may include, but is not limited to, at least one of the following: the target bounding box, category label, and confidence score of the candidate region where the candidate target is located), and convert the candidate region description information into descriptive prompt text (i.e., natural language prompt text), thereby calling the text encoder in the target multimodal large model to extract features from the target government knowledge text based on the descriptive prompt text, and obtain the text feature vector of the target government knowledge text. For example, the text encoder in the target multimodal large model can be called separately to extract features from the descriptive prompt text and the target government knowledge text to obtain the initial text feature vector of the descriptive prompt text and the initial text feature vector of the target government knowledge text, and then the initial text feature vector of the descriptive prompt text and the initial text feature vector of the target government knowledge text can be fused to obtain the text feature vector of the target government knowledge text; or, the descriptive prompt text and the target government knowledge text can be concatenated first to obtain concatenated text, and the concatenated text can be input into the text encoder in the target multimodal large model to output the text feature vector of the target government knowledge text, etc.; the embodiments of the present invention do not limit this.
[0052] Furthermore, the electronic device can invoke the joint reasoning module in the target multimodal large model to generate the government inspection results of the candidate target based on the target fused visual feature vector and the text feature vector of the target government knowledge text. Based on this, a multimodal large model may also include a joint reasoning module, such as the target multimodal large model may include a joint reasoning module. Optionally, a joint reasoning module may include, but is not limited to, at least one of the following: a cross-modal attention layer (e.g., used for information interaction, information fusion, and semantic association reasoning between features of different modalities), a multimodal feature fusion layer (e.g., used to fuse the encoding results of each modality to obtain a unified joint representation), a joint semantic encoding layer (e.g., used to perform high-level semantic joint encoding and reasoning upon receiving the fused features), and a reasoning decoding layer (used to output the final result), etc., which are not limited in this embodiment of the invention; for example, the functions implemented by any two layers may also be integrated into one unit, or the reasoning decoding layer may also be located outside the joint reasoning module, etc. For example, embodiments of the present invention can establish semantic alignment between the target fused visual feature vector and the target government affairs knowledge text text feature vector by establishing a cross-modal attention layer, thereby achieving semantic alignment between visual features (such as infrared high-temperature areas, such as smoke, flames, etc.) and target government affairs knowledge text (such as equipment operation specifications, etc.).
[0053] In one implementation, the electronic device can input the target fused visual feature vector and the target government affairs knowledge text text text into the joint reasoning module in the target multimodal big model, so as to output the government affairs inspection results of the candidate target through the joint reasoning module in the target multimodal big model.
[0054] In another implementation, the candidate target detection data may further include flight parameter feature vectors corresponding to the candidate targets. In this case, the electronic device may also call the low-altitude parameter modality processing module in the target multimodal large model to perform linear projection on the flight parameter feature vectors to obtain flight parameter projection feature vectors. The dimension of the flight parameter projection feature vectors is the same as the dimension of the target fused visual feature vectors. Accordingly, when calling the joint inference module in the target multimodal large model to generate the administrative inspection results of the candidate targets based on the target fused visual feature vectors and the text feature vectors of the target government affairs knowledge text, the joint inference module in the target multimodal large model can be called to generate the administrative inspection results of the candidate targets based on the target fused visual feature vectors, the text feature vectors of the target government affairs knowledge text, and the flight parameter projection feature vectors. For example, the target fused visual feature vectors, the text feature vectors of the target government affairs knowledge text, and the flight parameter projection feature vectors can be input into the joint inference module in the target multimodal large model to output the administrative inspection results of the candidate targets through the joint inference module in the target multimodal large model. Optionally, the low-altitude parameter modality processing module may be one or more linear projection layers; based on this, in this embodiment of the invention, the flight parameter feature vector can be expanded through the linear projection layer into a flight parameter projection feature vector with the same dimension as the target fused visual feature vector, so as to be embedded in the spatial context and participate in cross-modal fusion.
[0055] Optionally, the electronic device can also determine the weight values of each weight to be adjusted in the weight reorganization based on the target spatiotemporal label, thereby calling the joint inference module in the target multimodal large model to generate the administrative inspection results of the candidate target based on the target fused visual feature vector, the text feature vector of the target government knowledge text, and the weight values of each weight to be adjusted; for example, it can also call the joint inference module in the target multimodal large model to generate the administrative inspection results of the candidate target based on the target fused visual feature vector, the text feature vector of the target government knowledge text, the flight parameter projection feature vector, and the weight values of each weight to be adjusted, etc. Optionally, the weight reorganization to be adjusted may include, but is not limited to, the modal weighting coefficients of the multimodal fusion layer, including, but not limited to, at least one of the following: the weights corresponding to the target image data (such as the weights of the target fused visual feature vector), the weights corresponding to the target government knowledge text (such as the weights of the text feature vectors of the target government knowledge text), and the weights corresponding to the flight parameter data (such as the weights of the flight parameter projection feature vectors), etc., which are not limited in this embodiment of the invention; wherein, each weight in the modal weighting coefficients of the multimodal fusion layer can be used for feature fusion (i.e., weighted summation).
[0056] Optionally, when determining the weight values of each weight to be adjusted in the weight reorganization based on the target spatiotemporal label, the electronic device can determine the weight values of each weight to be adjusted in the weight reorganization based on the target spatiotemporal label according to a preset weight dynamic adjustment strategy; optionally, the preset weight dynamic adjustment strategy can be set according to experience or actual needs, and the embodiments of the present invention do not limit it in this regard. For example, the target spatiotemporal label may include relative altitude (i.e., the relative altitude when image data is collected in the area where the candidate target is located, or it can also be represented as flight altitude). Then, according to a preset weight dynamic adjustment strategy, the weight value of each weight to be adjusted in the weight reassembly can be determined based on the relative altitude. For example, when the relative altitude is less than or equal to the first altitude threshold (e.g., 50 meters), the first weight reassembly can be used to determine the weight value of each weight to be adjusted (i.e., each weight value in the first weight reassembly is used as the weight value of each weight to be adjusted). When the relative altitude is greater than or equal to the second altitude threshold (e.g., 100 meters), the second weight reassembly can be used to determine the weight value of each weight to be adjusted. When the relative altitude is greater than the first altitude threshold and less than the second altitude threshold, the third weight reassembly (or default weight reassembly) can be used to determine the weight value of each weight to be adjusted (i.e., the weight value of each weight to be adjusted can be the default value at this time), and so on. Optionally, the first altitude threshold, the first weight reassembly, the second altitude threshold, the second weight reassembly, and the third weight reassembly can all be set according to experience or actual needs, and the embodiments of the present invention do not limit this. Optionally, the weight value of the target image data corresponding to the weight in the first weight reassembly can be greater than the weight value of the target image data corresponding to the weight in the second weight reassembly; and / or, the weight value of the target image data corresponding to the weight in the first weight reassembly can be greater than the weight value of the target image data corresponding to the weight in the third weight reassembly, and the weight value of the target image data corresponding to the weight in the third weight reassembly can be greater than the weight value of the target image data corresponding to the weight in the second weight reassembly, and so on. Based on this, embodiments of the present invention can introduce low-altitude flight parameters as spatial constraints to dynamically adjust the inference weights (i.e., the weight reassemblies to be adjusted involved in the inference process) to improve model performance and thus improve inference accuracy; for example, when the relative altitude is less than or equal to the first altitude threshold, the model can enhance its sensitivity to the discrimination of local details in the image; while when the relative altitude is greater than or equal to the second altitude threshold, the focus can be on macroscopic risk pattern recognition, and so on.
[0057] In summary, the embodiments of the present invention can embed a government affairs rule engine (such as a text encoder) into a multimodal large model, thereby jointly verifying physical observations (i.e., target image data, such as temperature anomalies) and business logic (i.e., target government affairs knowledge text, such as cable joint temperature > 70℃ and no maintenance history). This effectively eliminates environmental interference and ultimately generates a parseable causal chain (i.e., a multimodal inference chain). For example, the multimodal inference chain can be: infrared image shows excessively high temperature in the joint area - visible light confirms that the location is a cable T-joint - government affairs database shows no maintenance record for the equipment in the past 6 months - combined with the current wind speed of 1.2 m / s, it is determined that the risk of heat accumulation is high. Based on this, the embodiments of the present invention can achieve dynamic aggregation of multimodal evidence and causal inference through a hierarchical routing mechanism in a multimodal large model.
[0058] Optionally, the anomaly detection result may be located within the multimodal inference chain (such as in the last inference step of the multimodal inference chain, such as the above-mentioned high risk of heat accumulation) or outside the multimodal inference chain (such as a conclusion identifier used to indicate a high risk of heat accumulation). This embodiment of the present invention does not limit this. Optionally, a conclusion identifier may be a conclusion name (such as high risk of heat accumulation) or a conclusion number. This embodiment of the present invention does not limit this.
[0059] Optionally, the results of government inspections may include qualitative diagnosis results of anomalies, which may include the aforementioned multimodal reasoning chain and anomaly detection conclusions (also known as anomaly type labels). In other words, the embodiments of the present invention can realize qualitative diagnosis of anomalies to output anomaly type labels that conform to industry standards, infer the root cause (such as "overheating of power equipment - poor contact", "abnormal temperature rise of municipal facilities - external heat source interference", etc.), thereby attaching a multimodal reasoning chain (also known as a multimodal evidence chain).
[0060] Optionally, the results of government inspections may include, but are not limited to, at least one of the following: quantitative risk assessment results, results of responsibility tracing and handling recommendations, etc., which are not limited in this embodiment of the invention. Optionally, the quantitative risk assessment results may include, but are not limited to, at least one of the following: risk level (such as low, medium, high, critical, etc.), risk assessment confidence level, etc., which are not limited in this embodiment of the invention. Optionally, the quantitative risk assessment results may be determined based on quantitative risk assessment indicator data, which may include, but are not limited to, at least one of the following: comprehensive temperature value (such as comprehensive outdoor temperature, etc.), equipment type (such as the type of power equipment, which can be inferred from enterprise information (i.e., inspection object information)), material hazard (such as that which can be inferred from enterprise information), surrounding population density (such as that which can be inferred from GIS), etc., which are not limited in this embodiment of the invention. In this case, a large model (such as a multimodal large model) may have a built-in risk assessment matrix to output quantitative risk assessment results. Optionally, the target government knowledge text may include quantitative risk assessment indicator data, which can be input into a multimodal large model (such as a target multimodal large model) through the target government knowledge text, or the quantitative risk assessment indicator data can be input into a multimodal large model or other large models separately to output quantitative risk assessment results, etc.; this embodiment of the invention does not limit this. Optionally, when the risk level is greater than a preset risk level (such as intermediate level), the quantitative risk assessment result may also include a risk diffusion map, which may be simulated based on dynamic environmental data, etc. Optionally, the preset risk level may be set according to experience or actual needs, this embodiment of the invention does not limit this.
[0061] Optionally, the results of responsibility tracing and handling recommendations may include, but are not limited to, at least one of the following: the identifier of the responsible unit (such as the name or number of the responsible unit), the identifier of the regulatory department (such as the name or number of the regulatory department), and preliminary handling recommendations, etc., which are not limited in this embodiment of the present invention; Optionally, each output field in the results of responsibility tracing and handling recommendations may be accompanied by a data source identifier to support audit traceability, etc.
[0062] As can be seen, the embodiments of the present invention can output structured, in-depth interpretation information through the model, which far exceeds that of traditional detection boxes.
[0063] S104. Based on the results of the government inspection, generate a standardized government inspection report; and send the standardized government inspection report to the regulatory objects corresponding to the candidate targets.
[0064] In one implementation, electronic devices can add the results of government inspections to a standardized government inspection report to generate such a report.
[0065] In another implementation, the electronic device can generate a standardized government inspection report based on the results of a government inspection, according to a preset report template; that is, it can drive a report template engine to automatically generate a standardized government inspection report based on the results of the government inspection. Based on this, a standardized government inspection report can be generated from the structured analysis results (i.e., the government inspection results). Optionally, both the preset report template and the report template engine can be set according to experience or actual needs, and this embodiment of the invention does not limit this. Optionally, the standardized government inspection report may include, but is not limited to, at least one of the following: textual descriptions (such as multimodal inference chains, anomaly detection conclusions, etc.), key evidence images (such as target image data or portraits of inspection objects within the target area), risk diffusion maps, etc., and this embodiment of the invention does not limit this.
[0066] It should be noted that the embodiments of the present invention do not limit the format of the standardized government inspection report. For example, the standardized government inspection report can be in PDF (Portable Document Format) and / or JSON (JavaScript Object Notation, a lightweight data exchange format), which can generate a standardized government inspection report in PDF format and / or JSON format, etc.
[0067] Optionally, the electronic device can send the standardized government inspection report to the corresponding regulatory object (such as the regulatory department and / or the inspection object) according to a preset push strategy. Optionally, the preset push strategy can be set based on experience or actual needs; this embodiment of the invention does not limit this, and may trigger multi-channel alerts such as SMS and in-app notifications, etc. Optionally, the government inspection system may also include a government cloud platform (i.e., the electronic device may also include a government cloud platform), in which case the report can be automatically pushed to the dedicated terminal of the relevant regulatory object based on the association results of the responsible entity through the government cloud platform, etc.
[0068] In summary, the embodiments of the present invention can encode flight control parameters such as altitude, attitude, and speed (i.e., flight parameter data) into spatial context features (i.e., flight parameter projection feature vectors), enabling multimodal large models to have adaptive observation scale capabilities, breaking through the limitations of traditional static image analysis, and realizing adaptive reasoning for low-altitude dynamic perception. Correspondingly, the embodiments of the present invention can innovatively encode low-altitude flight parameters (such as altitude, attitude, and speed) into spatial context features, enabling the model to have dynamic scale perception capabilities. When observing at close range at low altitudes, it enhances the sensitivity to discriminative details (such as equipment joint corrosion and local heat accumulation), and when conducting wide-area inspections at medium and high altitudes, it focuses on macroscopic risk pattern recognition (such as the spatial distribution pattern of heat sources). This mechanism enables the system to output qualitative diagnostic conclusions that conform to industry standards, realizing a cognitive leap from "seeing hot spots" to "understanding risks". Furthermore, embodiments of the present invention can automatically associate inspection object information (such as enterprise files), equipment specifications, historical records and other government knowledge based on GPS coordinates, and perform cross-modal alignment with infrared / visible light observations. Through a rule engine, an explainable causal chain of "physical anomaly - cause inference - responsibility tracing" is generated, thereby effectively realizing causal reasoning driven by government knowledge. Furthermore, when the edge side and the cloud side (also known as the cloud-based intelligent layer or cloud) are not located on the same electronic device, embodiments of the present invention can complete dual-light image registration and preliminary screening of candidate regions (i.e., generating candidate target detection data) on the edge side to reduce transmission load. Then, the large model in the cloud can perform deep inference based on simplified input, avoiding computational redundancy caused by full data processing. Moreover, the cloud uses a modality-specific attention adaptation layer to achieve deep alignment of infrared (temperature gradient) and visible light (texture edge) features, which can avoid the accuracy loss of traditional registration methods, thereby improving accuracy and realizing a lightweight fusion architecture of end (i.e., low-altitude perception layer)-edge-cloud collaboration. Based on this, embodiments of the present invention can optimize system resource utilization efficiency and meet timeliness requirements. In other words, while ensuring analysis depth, it can meet the strict requirements of response timeliness in low-altitude patrol scenarios.
[0069] Furthermore, embodiments of the present invention can automatically aggregate multimodal evidence into structured knowledge output through joint reasoning, including anomaly qualitative diagnosis, quantitative risk assessment, association of responsible parties, and disposal suggestions, which significantly shortens the chain from data collection to decision response, improves the efficiency and accuracy of government inspections, and enables decision support capabilities to directly meet the needs of government management.
[0070] This invention provides embodiments that can acquire candidate target detection data, including target image data and spatiotemporal tags for the candidate targets; wherein the spatiotemporal tags include location information corresponding to the candidate targets. Then, based on the target spatiotemporal tags, target government knowledge text corresponding to the candidate targets can be obtained from a government knowledge base; wherein the target government knowledge text includes at least one inspection object indication data, and one inspection object indication data includes at least one of the following: location information, geographic entity data, environmental dynamic data, and historical data of an inspection object. Based on this, a target multimodal large model can be invoked to generate government inspection results for the candidate targets based on the target image data and target government knowledge text. The government inspection results include a multimodal inference chain and anomaly detection conclusions; and a standardized government inspection report can be generated based on the government inspection results; and the standardized government inspection report can be sent to the regulatory object corresponding to the candidate targets. As can be seen, the embodiments of the present invention can place the physical observation target in a precise three-dimensional spatial context by using multimodal data such as target image data and target government affairs knowledge text, thereby obtaining government affairs inspection results with higher accuracy, reducing the false alarm rate, and thus effectively improving the accuracy of government affairs inspection. Furthermore, the embodiments of the present invention can not only detect physical anomalies (i.e., anomaly detection conclusions), but also infer their causes through multimodal reasoning chains, thereby transforming candidate target detection data into standardized government affairs inspection reports that can directly support government affairs decision-making.
[0071] Based on the above description, this embodiment of the invention also proposes a more specific method for government affairs inspection. Accordingly, this method can be executed by the aforementioned electronic device (terminal or server); or, it can be executed jointly by a terminal and a server, and so on. For ease of explanation, the following description will use the execution of this method by an electronic device as an example; please refer to [link to relevant documentation]. Figure 3 This government inspection method may include the following steps S301-S308: S301, Obtain a multimodal alignment training data set. A multimodal alignment training data set includes a training image data and a corresponding training knowledge data text. A training knowledge data text includes a government affairs knowledge text.
[0072] Optionally, the electronic device may store a multimodal alignment training data set in its own storage space. In this case, the electronic device can obtain the multimodal alignment training data set from its own storage space; or, it can obtain a download link for the multimodal alignment training data set and download the multimodal alignment training data set using the download link to achieve the acquisition of the multimodal alignment training data set, etc.; this embodiment of the invention does not limit this. Optionally, a multimodal alignment training data set can also be referred to as a two-light image-text pair.
[0073] Optionally, a training image dataset may include a visible light image and an infrared image, wherein the images in the image dataset are images acquired in the same area at the same acquisition time.
[0074] Optionally, a training knowledge data text may also include, but is not limited to, at least one of the following: flight parameter feature vectors, expert annotation conclusions (also known as data annotation results, such as manually annotated data), etc., and this embodiment of the invention does not limit this. Optionally, a data annotation result may include, but is not limited to, at least one of the following: multimodal inference chain annotations and anomaly detection conclusion annotations, etc., and this embodiment of the invention does not limit this.
[0075] S302, the first multimodal large model is invoked to extract features from each multimodal alignment training data in the multimodal alignment training data set, and the feature extraction results of each multimodal alignment training data are obtained; wherein, the feature extraction result of a multimodal alignment training data includes the feature vector of the training image data and the feature vector of the training knowledge data text in the corresponding multimodal alignment training data.
[0076] Optionally, the first multimodal large model can be set according to experience or actual needs, and the embodiments of the present invention do not limit it; that is, the embodiments of the present invention do not limit the model structure of the first multimodal large model.
[0077] It should be understood that when calling the first multimodal large model to extract features from each multimodal alignment training data in the multimodal alignment training dataset and obtain the feature extraction results of each multimodal alignment training data, for any multimodal alignment training data in the multimodal alignment training dataset, the electronic device can call each feature extraction module (such as a visual feature encoder and a text encoder) in the first multimodal large model to extract features from the training image data and training knowledge data text in any multimodal alignment training data (such as extracting features from the training image data through the visual feature encoder and extracting features from the training knowledge data text through the text encoder) and obtain the feature extraction results of any multimodal alignment training data.
[0078] S303, based on the feature extraction results of each multimodal aligned training data, calculate the model loss value of the first multimodal large model; and optimize the model parameters of the visual text projection layer in the first multimodal large model in the direction of reducing the model loss value of the first multimodal large model, so as to reduce the distance between the feature vector of the training image data and the feature vector of the training knowledge data text in the same multimodal aligned training data obtained when the optimized first multimodal large model performs feature extraction.
[0079] The visual text projection layer can be a projection layer between the visual feature encoder and the text encoder. Accordingly, in this embodiment of the invention, the visual text projection layer can be trained using a multimodal alignment training dataset, thereby enabling the multimodal large model to possess basic image-text semantic alignment capabilities. This effectively improves the image-text semantic alignment capabilities of the multimodal large model, thus enhancing model performance. Optionally, the training process of the visual text projection layer can also be referred to as modal alignment pre-training.
[0080] Optionally, when calculating the model loss value of the first multimodal large model based on the feature extraction results of each multimodal aligned training data, the electronic device can calculate the model positive sample pair loss value based on the feature extraction results of each multimodal aligned training data, and determine the model loss value of the first multimodal large model based on the model positive sample pair loss value. Optionally, when calculating the model positive sample pair loss value based on the feature extraction results of each multimodal aligned training data, for any multimodal aligned training data in the multimodal aligned training data set, the positive sample pair loss value under any multimodal aligned training data can be calculated, and the weighted summation result (such as the mean operation result or the summation operation result, etc.) between the positive sample pair loss values under each multimodal aligned training data can be used as the model positive sample pair loss value; wherein, the positive sample pair loss value under any multimodal aligned training data can be: the distance (such as Euclidean distance or cosine distance, etc.) between the feature vector of the training image data and the feature vector of the training knowledge data text in any multimodal aligned training data.
[0081] Optionally, when determining the model loss value of the first multimodal large model based on the model positive sample pair loss value, the model positive sample pair loss value can be used as the model loss value of the first multimodal large model; or, the model negative sample pair loss value can also be determined, and the model loss value of the first multimodal large model can be determined using the model positive sample pair loss value and the model negative sample pair loss value. For example, the model loss value of the first multimodal large model can be the sum of the negatives of the model positive sample pair loss value and the model negative sample pair loss value, or it can be the sum of the reciprocals of the model positive sample pair loss value and the model negative sample pair loss value, etc.; the embodiments of the present invention do not limit this. Optionally, when determining the model's negative sample pair loss value, multiple negative sample pairs can be identified from the multimodal aligned training dataset, and the inter-pair distance of each negative sample pair in the multiple negative sample pairs can be calculated (the inter-pair distance of a negative sample pair can be the distance between the feature vector of the training image data and the feature vector of the training knowledge data text in the corresponding negative sample pair), so that the weighted sum of the inter-pair distances of each negative sample pair is used as the model's negative sample pair loss value. A negative sample pair may include training image data from the first selected multimodal alignment training data and training knowledge data text from the second selected multimodal alignment training data. The first selected multimodal alignment training data and the second selected multimodal alignment training data may be any two different multimodal alignment training data in the multimodal alignment training data set. It should be noted that the specific method for determining multiple negative sample pairs is not limited in the embodiments of the present invention. For example, a negative sample pair may be formed by combining the training image data from one multimodal alignment training data set with the training knowledge data text from each multimodal alignment training data set other than the corresponding multimodal alignment training data set. Alternatively, M pairs of multimodal alignment training data may be randomly selected from the multimodal alignment training data set, and each pair of multimodal alignment training data may be cross-combined to form two negative sample pairs, where M is a positive integer, and so on.
[0082] Based on this, embodiments of the present invention can train a visual text projection layer to reduce the distance between the feature vectors of the training image data and the feature vectors of the training knowledge data text in the same multimodal alignment training data, and / or increase the distance between the feature vectors of the training image data and the feature vectors of the training knowledge data text in different multimodal alignment training data, thereby achieving image-text semantic alignment of the same multimodal alignment training data, and further achieving image-text semantic alignment between the same two-light image-text pair.
[0083] S304, based on the optimized first multimodal large model, determine the target multimodal large model.
[0084] In one implementation, the electronic device can continue to train the optimized first multimodal large model using the multimodal aligned training dataset until a first model convergence condition is met (such as the number of iterations reaching a first iteration threshold or the model loss value being less than a first model loss threshold). The multimodal large model that meets the first model convergence condition is then used as the target multimodal large model. Optionally, both the first iteration threshold and the first model loss threshold can be set based on experience or actual needs; this embodiment of the invention does not limit this.
[0085] In another implementation, the electronic device can acquire a training patrol data set. A training patrol data set may include, but is not limited to, a patrol image, corresponding government affairs text, and data annotation results. Based on the optimized first multimodal large model, a second multimodal large model can be determined; that is, the multimodal large model that meets the convergence condition of the first model is used as the second multimodal large model. Then, the second multimodal large model can be invoked to determine the government affairs patrol probability prediction data for each training patrol data set, based on the patrol image and government affairs text included in each training patrol data set. Based on the government affairs patrol probability prediction data and data annotation results for each training patrol data set, the model loss value of the second multimodal large model is calculated. Based on this, the parameters to be optimized in the second multimodal large model can be optimized in the direction of reducing the model loss value, resulting in an optimized second multimodal large model. Based on the optimized second multimodal large model, a target multimodal large model is determined. The parameters to be optimized may include, but are not limited to, at least one of the following: model parameters of the cross-modal attention layer and model parameters of the language generation head.
[0086] Optionally, the electronic device can obtain the training and inspection data set from its own storage space; alternatively, it can obtain a download link for the training and inspection data set and download the training and inspection data set using the download link to achieve the acquisition of the training and inspection data set, etc.; this embodiment of the invention does not limit this. Optionally, a training and inspection data set may be the same as or different from a multimodal alignment training data set; this embodiment of the invention does not limit this.
[0087] Optionally, when calling the second multimodal large model to determine the government inspection probability prediction data for each training patrol data based on the patrol image data and government knowledge text included in each training patrol data set, for any training patrol data in the training patrol data set, the patrol image data and government knowledge text in any training patrol data can be input into the second multimodal large model to output the government inspection probability prediction data for any training patrol data through the second multimodal large model; or, a training patrol data may also include a patrol flight parameter feature vector, then the patrol image data, government knowledge text, and patrol flight parameter feature vector in any training patrol data can be input into the second multimodal large model to output the government inspection probability prediction data for any training patrol data through the second multimodal large model, and so on; the embodiments of the present invention do not limit this.
[0088] Among them, the government inspection probability prediction data of any training inspection data may include the predicted probability of each word (i.e., token) in the data annotation results of any training inspection data, and the government inspection probability prediction data of any training inspection data can be used to generate the government inspection results corresponding to any training inspection data. Optionally, when calculating the model loss value of the second multimodal large model based on the government inspection probability prediction data and data annotation results of each training inspection data (the data annotation result of one training inspection data is the data annotation result in the corresponding training inspection data), the cross-entropy loss can be calculated based on the government inspection probability prediction data and data annotation results of each training inspection data to obtain the model loss value of the second multimodal large model; alternatively, the government inspection results corresponding to each training inspection data can be determined based on the government inspection probability prediction data of each training inspection data, and the consistency verification of the government inspection results corresponding to any training inspection data and the data annotation results in any training inspection data can be performed to obtain the consistency verification score of any training inspection data. Then, the model loss value of the second multimodal large model can be determined based on the sum of the consistency verification scores of each training inspection data. For example, the reciprocal or negative number of the sum of the consistency verification scores of each training inspection data can be used as the model loss value of the second multimodal large model, etc.; the embodiments of the present invention do not limit this. Optionally, the electronic device can call any large language model to perform consistency verification on the government inspection results corresponding to any training inspection data and the data annotation results in any training inspection data, and obtain the consistency verification score of any training inspection data, etc.
[0089] For example, the electronic device can also calculate the model loss value of the second multimodal large model based on the government inspection probability prediction data, government knowledge text, and data annotation results of each training inspection data. In this case, the government inspection probability prediction data of any training inspection data can also include the prediction probability of each word in the government knowledge text of any training inspection data, and so on. It should be noted that the calculation of the model loss value of the second multimodal large model based on the government inspection probability prediction data, government knowledge text, and data annotation results of each training inspection data can be the same as the implementation method of calculating the model loss value of the second multimodal large model based on the government inspection probability prediction data and data annotation results of each training inspection data. The embodiments of the present invention will not be described in detail here.
[0090] Optionally, the language generation head may include, but is not limited to, at least one of the following: a joint semantic encoding layer and an inference decoding layer, etc., which are not limited in this embodiment of the invention. Based on this, this embodiment of the invention can freeze the visual backbone parameters, etc., and fine-tune the cross-modal attention layer and / or language generation head in the multimodal large model by training the patrol data set to improve the model performance; correspondingly, this embodiment of the invention can introduce a spatial context loss function (that is, loss calculation can be performed through government knowledge text, patrol flight parameter feature vectors, etc.) to force the multimodal large model to learn to dynamically adjust the extraction granularity of visual features according to altitude, and prevent the model from generating false alarms when the data quality is poor (jitter and blur), which can improve the robustness of the government patrol system under complex weather conditions.
[0091] Optionally, the electronic device can continue to use the training inspection dataset to train the optimized second multimodal large model until the second model convergence condition is met (such as the number of iterations reaching the second iteration threshold or the model loss value being less than the second model loss threshold, etc.), thereby using the multimodal large model that has met the second model convergence condition as the third multimodal large model. Optionally, both the second iteration threshold and the second model loss threshold can be set according to experience or actual needs, and this embodiment of the invention does not limit this.
[0092] Optionally, the electronic device can use a third multimodal large model as the target multimodal large model. Alternatively, the electronic device can also acquire a set of instruction fine-tuning data (e.g., from its own storage space or downloaded via a link to the instruction fine-tuning data set). One instruction fine-tuning data set may include, but is not limited to, at least one of the following: scene description data (including images and / or text), problem description information, inference chains, and inference conclusions. It can also call the third multimodal large model and, based on the scene description data and problem description information in each instruction fine-tuning data set, determine the probability prediction data for each instruction fine-tuning data set. The probability prediction data for one instruction fine-tuning data set may include the predicted probabilities of each word in the inference chain and inference conclusion of the corresponding instruction fine-tuning data set. Then, based on the probability prediction data, inference chain, and inference conclusion of each instruction fine-tuning data set, the model loss value of the third multimodal large model can be calculated. This allows for optimization of the parameters to be optimized in the third multimodal large model in the direction of reducing the model loss value, resulting in an optimized third multimodal large model. Finally, based on the optimized third multimodal large model, the target multimodal large model can be determined. It should be noted that the calculation of the model loss value of the third multimodal large model based on the probabilistic prediction data, inference chain, and inference conclusions of each instruction fine-tuning data can be carried out in the same way as the calculation of the model loss value of the second multimodal large model based on the probabilistic prediction data of government inspections and data annotation results of each training inspection data. This embodiment of the invention will not be elaborated upon here. Based on this, the multimodal large model can learn human logical thinking in answering questions through scene description data and problem description information, thereby improving model performance.
[0093] Accordingly, the electronic device can continue to use instructions to fine-tune the data set to train the optimized third multimodal large model until the third model convergence condition is met (such as the number of iterations reaching the third iteration threshold or the model loss value being less than the third model loss threshold, etc.). The multimodal large model that meets the third model convergence condition is then used as the target multimodal large model, and so on. Optionally, the third iteration threshold and the third model loss threshold can both be set according to experience or actual needs; this embodiment of the invention does not limit this.
[0094] S305, acquire candidate target detection data, which includes target image data and target spatiotemporal labels of the candidate targets; wherein, the target spatiotemporal labels include the location information corresponding to the candidate targets.
[0095] S306, Based on the target spatiotemporal label, obtain the target government knowledge text corresponding to the candidate target from the government knowledge base; wherein, the target government knowledge text includes at least one inspection object indication data, and the inspection object indication data includes at least one of the following: the location information, geographic entity data, environmental dynamic data and historical data of an inspection object.
[0096] S307, invoke the target multimodal large model, based on the target image data and target government affairs knowledge text, to generate government affairs inspection results for candidate targets. The government affairs inspection results include multimodal inference chains and anomaly detection conclusions.
[0097] S308 generates a standardized government inspection report based on the results of government inspections and sends the standardized government inspection report to the regulatory objects corresponding to the candidate targets.
[0098] In summary, the embodiments of this invention can deeply integrate multimodal data such as infrared / visible light images, low-altitude flight parameters (altitude, attitude, speed), and government knowledge text collected by government inspection and data collection equipment. This places the physical observation target in a precise three-dimensional spatial context, enabling accurate identification of thermal anomalies and semantic depth based on "spatial location-semantic attributes-causal logic." Based on this, a joint reasoning mechanism is constructed, enabling the government inspection system not only to detect physical anomalies but also to infer their causes, associate responsible parties, assess risk levels, and provide intuitive risk diffusion simulations. This transforms raw data into "knowledge" that can directly support government decision-making. Furthermore, the embodiments of this invention design an analysis architecture optimized for low-altitude mobile platforms, which can fully utilize the dynamic, multi-angle, real-time dual-light image data and low-altitude flight parameter data with precise pose information unique to government inspection and data collection equipment. This can effectively improve the level of intelligent inspection in complex urban environments.
[0099] For example, when applied to power facility safety inspection scenarios, after a drone detects an abnormally high temperature at a cable joint, it can automatically link enterprise files and maintenance records through government inspection methods, integrate infrared temperature, visible light structure, and meteorological data, infer and determine the risk of poor contact, and output a structured report containing the responsible unit and disposal suggestions, directly pushing it to the power operation and maintenance terminal; or, when applied to the identification of illegal burning, it can combine heat source detection, straw accumulation traces, land ownership and agricultural registration information, eliminate interference from normal heat sources such as mechanical operations, accurately identify illegal burning, and generate a risk report based on wind direction to simulate the smoke diffusion path and push it to the environmental protection department; or, when applied to forest fire monitoring scenarios, it can integrate vegetation type Historical fire risk and meteorological data can be used to identify early fire thermal anomalies, distinguish between natural heat sources and actual fire points, and predict the spread trend by combining topography and wind direction, providing fire early warning and response decision support. Alternatively, in river sewage discharge monitoring, abnormal hot discharge outlets can be identified through water body infrared thermal characteristics, and compliance can be determined by linking sewage discharge permits with enterprise information, generating alarms for illegal discharges, and simulating the pollution spread range by combining water flow direction to assist in precise law enforcement. Or, in emergency search and rescue assistance scenarios, trapped personnel can be quickly located after a disaster by analyzing thermal anomaly distribution, and the risks of rescue routes can be assessed by combining building structure and geographic information, generating rescue assistance maps containing vital sign hotspots and safety alerts to improve search and rescue efficiency and safety, etc. All these scenarios can be based on a unified multimodal fusion architecture, achieving direct transformation from physical observation to government decision-making, demonstrating the universality and practical value of this invention.
[0100] In this embodiment of the invention, after obtaining a multimodal alignment training dataset, a first multimodal large model is invoked to extract features from each multimodal alignment training data in the dataset, obtaining feature extraction results for each multimodal alignment training data. Based on the feature extraction results of each multimodal alignment training data, the model loss value of the first multimodal large model is calculated. Following the direction of reducing the model loss value of the first multimodal large model, the model parameters of the visual text projection layer in the first multimodal large model are optimized to obtain an optimized first multimodal large model. This reduces the distance between the feature vectors of the training image data and the feature vectors of the training knowledge data text in the same multimodal alignment training dataset obtained when the optimized first multimodal large model performs feature extraction. Furthermore, based on the optimized first multimodal large model, a target multimodal large model can be determined. Based on this, candidate target detection data can be obtained, including target image data and spatiotemporal labels of the candidate targets. Based on the target spatiotemporal labels, the corresponding government affairs knowledge text of the candidate targets can be obtained from the government affairs knowledge base. This allows the use of a target multimodal large model to generate government affairs inspection results for the candidate targets based on the target image data and the target government affairs knowledge text. The government affairs inspection results include a multimodal inference chain and anomaly detection conclusions. Furthermore, a standardized government affairs inspection report can be generated based on the government affairs inspection results and sent to the regulatory objects corresponding to the candidate targets. It is evident that this embodiment of the invention can ensure that the target multimodal large model can understand the correlation between low-altitude dynamic parameters and government affairs semantics through a specific training data construction and fine-tuning process, thereby effectively improving the model performance of the target multimodal large model and thus effectively improving the accuracy of government affairs inspection results. Moreover, this embodiment of the invention can effectively eliminate environmental interference factors by introducing a strong semantic context provided by the government affairs knowledge base, jointly reasoning with physical observations and government affairs information, significantly improving the accuracy and reliability of anomaly detection, thereby effectively reducing the false alarm rate and greatly improving the reliability of identification.
[0101] Based on the description of the relevant embodiments of the above-mentioned government affairs inspection method, this embodiment of the invention also proposes a government affairs inspection device, which can be a computer program (including program code) running on an electronic device; such as Figure 4 As shown, the government affairs inspection device may include an acquisition unit 401 and a processing unit 402. The government affairs inspection device can perform... Figure 1 or Figure 3 The government inspection method shown, that is, the government inspection device can operate the above-mentioned unit: The acquisition unit 401 is used to acquire candidate target detection data, which includes target image data and target spatiotemporal labels of the candidate targets; wherein, the target spatiotemporal labels include the location information corresponding to the candidate targets; Processing unit 402 is used to obtain target government knowledge text corresponding to the candidate target from the government knowledge base based on the target spatiotemporal label; wherein, the target government knowledge text includes at least one patrol object indication data, and a patrol object indication data includes at least one of the following: location information, geographic entity data, environmental dynamic data and historical data of a patrol object; The processing unit 402 is further configured to call the target multimodal large model, and generate the government inspection results of the candidate target based on the target image data and the target government affairs knowledge text, wherein the government inspection results include a multimodal inference chain and anomaly detection conclusions; The processing unit 402 is further configured to generate a standardized government inspection report based on the government inspection results; and send the standardized government inspection report to the regulatory object corresponding to the candidate target.
[0102] In one implementation, the target image data includes a visible light image and an infrared image of the candidate target; when the processing unit 402 calls the target multimodal large model and generates the government inspection results of the candidate target based on the target image data and the target government affairs knowledge text, it can be specifically used for: The visual feature encoder in the target multimodal large model is invoked to extract features from the target visible light image and the target infrared image respectively, to obtain the image features of the target visible light image and the image features of the target infrared image; wherein, the visual feature encoder includes a visible light adaptation layer and an infrared adaptation layer, the image features of the target visible light image are extracted through the visible light adaptation layer, and the image features of the target infrared image are extracted through the infrared adaptation layer; Based on the image features of the target's visible light image and the image features of the target's infrared image, a target fusion visual feature vector is determined; The text encoder in the target multimodal large model is invoked to extract features from the target government knowledge text, thereby obtaining the text feature vector of the target government knowledge text; The joint reasoning module in the target multimodal large model is invoked to generate the government inspection results of the candidate target based on the target fused visual feature vector and the text feature vector of the target government knowledge text.
[0103] In another embodiment, the candidate target detection data further includes the flight parameter feature vector corresponding to the candidate target; the processing unit 402 can also be used for: The low-altitude parameter modality processing module in the target multimodal large model is invoked to perform linear projection on the flight parameter feature vector to obtain the flight parameter projection feature vector. The dimension of the flight parameter projection feature vector is the same as the dimension of the target fused visual feature vector. The step of invoking the joint inference module in the target multimodal large model to generate the government inspection results of the candidate target based on the target's fused visual feature vector and the target's government affairs knowledge text text includes: The joint inference module in the target multimodal large model is invoked to generate the administrative inspection results of the candidate target based on the target fused visual feature vector, the text feature vector of the target government affairs knowledge text, and the flight parameter projection feature vector.
[0104] In another embodiment, when acquiring candidate target detection data, the acquisition unit 401 may specifically be used for: Acquire regional image data, which includes an initial visible light image and an initial infrared image; The initial visible light acquisition image is preprocessed to obtain the target visible light acquisition image; and the initial infrared acquisition image is preprocessed to obtain the target infrared acquisition image. Using the target visible light acquisition image as a reference, the target infrared acquisition image is pixel-level aligned to obtain an aligned infrared acquisition image; Target detection is performed on the target visible light acquisition image and the aligned infrared acquisition image to obtain target detection results; and candidate target detection data is obtained based on the target detection results.
[0105] In another embodiment, the acquisition unit 401 can also be used for: Acquire multiple flight parameter data within a target time range, wherein the multiple flight parameter data are time-series data within the target time range; wherein the acquisition time of the regional image data is within the target time range; Temporal features are extracted from the multiple flight parameter data to obtain the flight parameter temporal features; When acquiring candidate target detection data based on the target detection result, the acquisition unit 401 can be specifically used for: Based on the target detection results and the time-series characteristics of the flight parameters, candidate target detection data are obtained.
[0106] In another embodiment, the acquisition unit 401 can also be used for: Obtain a multimodal alignment training dataset. A multimodal alignment training dataset includes a training image dataset and the corresponding training knowledge data text. A training knowledge data text includes a government affairs knowledge text. Processing unit 402 can also be used for: The first multimodal large model is invoked to extract features from each multimodal alignment training data in the multimodal alignment training dataset, thereby obtaining the feature extraction results of each multimodal alignment training data. The feature extraction result of a multimodal alignment training data includes the feature vector of the training image data and the feature vector of the training knowledge data text in the corresponding multimodal alignment training data. Based on the feature extraction results of each multimodal alignment training data, the model loss value of the first multimodal large model is calculated; and in accordance with the direction of reducing the model loss value of the first multimodal large model, the model parameters of the visual text projection layer in the first multimodal large model are optimized to obtain the optimized first multimodal large model, so as to reduce the distance between the feature vector of the training image data and the feature vector of the training knowledge data text in the same multimodal alignment training data obtained when the optimized first multimodal large model performs feature extraction; Based on the optimized first multimodal large model, the target multimodal large model is determined.
[0107] In another implementation, when determining the target multimodal large model based on the optimized first multimodal large model, the processing unit 402 may specifically be used to: Obtain a training patrol data set. Each training patrol data set includes a patrol image data set, the corresponding government affairs knowledge text for the patrol image data set, and the data annotation results. Based on the optimized first multimodal large model, a second multimodal large model is determined; The second multimodal large model is invoked to determine the government inspection probability prediction data of each training inspection data based on the inspection image data and government knowledge text included in each training inspection data set. Based on the government inspection probability prediction data and data annotation results of each training inspection data, the model loss value of the second multimodal large model is calculated; Following the direction of reducing the model loss value of the second multimodal large model, the parameters to be optimized in the second multimodal large model are optimized to obtain the optimized second multimodal large model; and based on the optimized second multimodal large model, the target multimodal large model is determined; wherein, the parameters to be optimized include at least one of the following: model parameters of the cross-modal attention layer and model parameters of the language generation head.
[0108] According to one embodiment of the present invention, Figure 4 Each unit in the illustrated government inspection device can be individually or entirely merged into one or more other units, or one or more of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of the present invention. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of the present invention, any government inspection device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.
[0109] According to another embodiment of the present invention, it is possible to perform operations such as those described above by running on a general-purpose electronic device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 1 or Figure 3 The computer program (including program code) involved in each step of the corresponding method shown, to construct such... Figure 4 The document describes a government affairs inspection device and a government affairs inspection method for implementing embodiments of the present invention. The computer program can be recorded on, for example, a computer storage medium, loaded onto the aforementioned electronic device via the computer storage medium, and run therein.
[0110] Based on the description of the method and apparatus embodiments above, an exemplary embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the electronic device to perform the method according to an embodiment of the present invention.
[0111] An exemplary embodiment of the present invention also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of the present invention.
[0112] An exemplary embodiment of the present invention also provides a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a method according to an embodiment of the present invention.
[0113] refer to Figure 5The present invention will now be described in the form of a structural block diagram of an electronic device 500 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0114] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0115] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, output unit 507, storage unit 508, and communication unit 509. Input unit 506 can be any type of device capable of inputting information to electronic device 500. Input unit 506 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 507 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 508 may include, but is not limited to, disk and optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0116] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the government inspection method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. In some embodiments, the computing unit 501 can be configured to perform the government inspection method by any other suitable means (e.g., by means of firmware).
[0117] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0118] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0119] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0120] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0121] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0122] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0123] Furthermore, it should be understood that the above-disclosed embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, any equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.
Claims
1. A method for government inspection, characterized in that, include: Acquire candidate target detection data, which includes target image data and spatiotemporal labels of the candidate targets; wherein, the spatiotemporal labels of the targets include the location information corresponding to the candidate targets, and the candidate target detection data also includes the flight parameter feature vectors corresponding to the candidate targets; Based on the target spatiotemporal label, target government knowledge text corresponding to the candidate target is obtained from the government knowledge base; wherein, the target government knowledge text includes at least one patrol object indication data, and a patrol object indication data includes at least one of the following: location information, geographic entity data, environmental dynamic data, and historical data of a patrol object; wherein, the at least one patrol object indication data includes patrol object indication data of at least one patrol object within the target area, and the target area refers to the area centered on the target location indicated by the location information corresponding to the candidate target; the target government knowledge text supports the exclusion of environmental interference. The target multimodal large model is invoked, and based on the target image data and the target government affairs knowledge text, the government affairs inspection results of the candidate targets are generated. The government affairs inspection results include a multimodal inference chain and anomaly detection conclusions. The flight parameter feature vector can be input into the target multimodal large model to reflect the observation scale. Based on the results of the government inspection, a standardized government inspection report is generated; and the standardized government inspection report is sent to the regulatory object corresponding to the candidate target. The target multimodal large model is obtained by training a second multimodal large model; the method further includes: Obtain a training patrol data set. A training patrol data set includes a patrol image data set, the corresponding government affairs knowledge text for the patrol image data set, a patrol flight parameter feature vector, and data annotation results. The second multimodal large model is invoked to determine the government inspection probability prediction data of each training inspection data based on the inspection image data, government knowledge text and inspection flight parameter feature vector included in each training inspection data set; Based on the government inspection probability prediction data and data annotation results of each training inspection data, the model loss value of the second multimodal large model is calculated, so as to introduce spatial context loss through government knowledge text and inspection flight parameter feature vector; Following the direction of reducing the model loss value of the second multimodal large model, the parameters to be optimized in the second multimodal large model are optimized to obtain the optimized second multimodal large model; and based on the optimized second multimodal large model, the target multimodal large model is determined; wherein, the parameters to be optimized include at least one of the following: model parameters of the cross-modal attention layer and model parameters of the language generation head.
2. The method according to claim 1, characterized in that, The target image data includes the visible light image and infrared image of the candidate target; the step of calling the target multimodal large model, based on the target image data and the target government affairs knowledge text, to generate the government affairs inspection results of the candidate target includes: The visual feature encoder in the target multimodal large model is invoked to extract features from the target visible light image and the target infrared image respectively, to obtain the image features of the target visible light image and the image features of the target infrared image; wherein, the visual feature encoder includes a visible light adaptation layer and an infrared adaptation layer, the image features of the target visible light image are extracted through the visible light adaptation layer, and the image features of the target infrared image are extracted through the infrared adaptation layer; Based on the image features of the target's visible light image and the image features of the target's infrared image, a target fusion visual feature vector is determined; The text encoder in the target multimodal large model is invoked to extract features from the target government knowledge text, thereby obtaining the text feature vector of the target government knowledge text; The joint reasoning module in the target multimodal large model is invoked to generate the government inspection results of the candidate target based on the target fused visual feature vector and the text feature vector of the target government knowledge text.
3. The method according to claim 2, characterized in that, The method further includes: The low-altitude parameter modality processing module in the target multimodal large model is invoked to perform linear projection on the flight parameter feature vector to obtain the flight parameter projection feature vector. The dimension of the flight parameter projection feature vector is the same as the dimension of the target fused visual feature vector. The step of invoking the joint inference module in the target multimodal large model to generate the government inspection results of the candidate target based on the target's fused visual feature vector and the target's government affairs knowledge text text includes: The joint inference module in the target multimodal large model is invoked to generate the administrative inspection results of the candidate target based on the target fused visual feature vector, the text feature vector of the target government affairs knowledge text, and the flight parameter projection feature vector.
4. The method according to any one of claims 1-3, characterized in that, The acquisition of candidate target detection data includes: Acquire regional image data, which includes an initial visible light image and an initial infrared image; The initial visible light acquisition image is preprocessed to obtain the target visible light acquisition image; and the initial infrared acquisition image is preprocessed to obtain the target infrared acquisition image. Using the target visible light acquisition image as a reference, the target infrared acquisition image is pixel-level aligned to obtain an aligned infrared acquisition image; Target detection is performed on the target visible light acquisition image and the aligned infrared acquisition image to obtain target detection results; and candidate target detection data is obtained based on the target detection results.
5. The method according to claim 4, characterized in that, The method further includes: Acquire multiple flight parameter data within a target time range, wherein the multiple flight parameter data are time-series data within the target time range; wherein the acquisition time of the regional image data is within the target time range; Temporal features are extracted from the multiple flight parameter data to obtain the flight parameter temporal features; The step of obtaining candidate target detection data based on the target detection results includes: Based on the target detection results and the time-series characteristics of the flight parameters, candidate target detection data are obtained.
6. The method according to any one of claims 1-3, characterized in that, The method further includes: Obtain a multimodal alignment training dataset. A multimodal alignment training dataset includes a training image dataset and the corresponding training knowledge data text. A training knowledge data text includes a government affairs knowledge text. The first multimodal large model is invoked to extract features from each multimodal alignment training data in the multimodal alignment training dataset, thereby obtaining the feature extraction results of each multimodal alignment training data. The feature extraction result of a multimodal alignment training data includes the feature vector of the training image data and the feature vector of the training knowledge data text in the corresponding multimodal alignment training data. Based on the feature extraction results of each multimodal alignment training data, the model loss value of the first multimodal large model is calculated; and in accordance with the direction of reducing the model loss value of the first multimodal large model, the model parameters of the visual text projection layer in the first multimodal large model are optimized to obtain the optimized first multimodal large model, so as to reduce the distance between the feature vector of the training image data and the feature vector of the training knowledge data text in the same multimodal alignment training data obtained when the optimized first multimodal large model performs feature extraction; Based on the optimized first multimodal large model, the second multimodal large model is determined.
7. A government affairs inspection device, characterized in that, The device includes: The acquisition unit is used to acquire candidate target detection data, which includes target image data and spatiotemporal labels of the candidate targets; wherein, the spatiotemporal labels of the targets include the location information corresponding to the candidate targets, and the candidate target detection data also includes the flight parameter feature vectors corresponding to the candidate targets; The processing unit is configured to obtain target government knowledge text corresponding to the candidate target from the government knowledge base based on the target spatiotemporal label; wherein, the target government knowledge text includes at least one patrol object indication data, and the patrol object indication data includes at least one of the following: location information, geographic entity data, environmental dynamic data, and historical data of a patrol object; wherein, the at least one patrol object indication data includes patrol object indication data of at least one patrol object within the target area, the target area refers to the area centered on the target location indicated by the location information corresponding to the candidate target, and the target government knowledge text supports the exclusion of environmental interference; The unit is also used to invoke the target multimodal large model, and based on the target image data and the target government affairs knowledge text, generate the government affairs inspection results of the candidate target. The government affairs inspection results include a multimodal inference chain and anomaly detection conclusions. The flight parameter feature vector can be input into the target multimodal large model to reflect the observation scale. The processing unit is further configured to generate a standardized government inspection report based on the government inspection results; and send the standardized government inspection report to the regulatory object corresponding to the candidate target; The target multimodal large model is obtained by training the second multimodal large model; the processing unit is also used for: Obtain a training patrol data set. A training patrol data set includes a patrol image data set, the corresponding government affairs knowledge text for the patrol image data set, a patrol flight parameter feature vector, and data annotation results. The second multimodal large model is invoked to determine the government inspection probability prediction data of each training inspection data based on the inspection image data, government knowledge text and inspection flight parameter feature vector included in each training inspection data set; Based on the government inspection probability prediction data and data annotation results of each training inspection data, the model loss value of the second multimodal large model is calculated, so as to introduce spatial context loss through government knowledge text and inspection flight parameter feature vector; Following the direction of reducing the model loss value of the second multimodal large model, the parameters to be optimized in the second multimodal large model are optimized to obtain the optimized second multimodal large model; and based on the optimized second multimodal large model, the target multimodal large model is determined; wherein, the parameters to be optimized include at least one of the following: model parameters of the cross-modal attention layer and model parameters of the language generation head.
8. An electronic device, characterized in that, include: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.