Water affair inspection method and equipment based on visual language large model, and storage medium
Through the water inspection method based on the visual language large model, the problems of low efficiency and low accuracy of traditional water inspection are solved, and efficient and accurate water inspection results are generated.
Patent Information
- Application Number
- CN202510565193.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional water inspections rely on manual patrols and traditional computer vision systems, which are inefficient and inaccurate, and lack the ability to understand the semantics of complex scenarios and the integration of water expertise.
A water conservancy inspection method based on a visual language big model is adopted. By converting the image encoding format and generating question sentences, the visual language big model is used to process the images and sentences to obtain the water conservancy inspection results.
It improves the efficiency and accuracy of water inspections, can effectively identify key inspection points of water inspections, and generate detailed or brief inspection reports.
Smart Images

Figure CN120708018A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of environmental monitoring, and in particular to a water inspection method, equipment and storage medium based on a visual language large model. Background Art
[0002] Waterworks inspections, as a crucial tool for water resource management, are crucial for timely identifying and resolving water issues. The evolution of waterworks inspection technology is closely tied to urbanization and technological innovation. Traditional waterworks inspections rely primarily on manual on-site inspections or fixed camera surveillance. Workers regularly conduct on-site inspections along rivers, reservoirs, and other water bodies, documenting their conditions and enabling the detection and reporting of anomalies. However, traditional waterworks inspections rely primarily on manual visual judgment and subjective experience, resulting in low efficiency and accuracy. Summary of the Invention
[0003] The main purpose of this application is to provide a water inspection method, equipment and storage medium based on a visual language large model, aiming to improve the efficiency and accuracy of water inspection.
[0004] In a first aspect, the present application provides a water affairs inspection method, comprising:
[0005] Acquire a plurality of first images taken in the water area to be inspected;
[0006] Converting the encoding format of the plurality of first images to obtain a plurality of second images suitable for a large visual language model; and
[0007] Based on multiple keywords of water affairs inspection in a preset knowledge base, generating question sentences for asking questions to the visual language model;
[0008] The plurality of second images and the question statements are input into the visual language model for processing to obtain water inspection results.
[0009] In a second aspect, the present application also provides a computer device, which includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of the water inspection method described above are implemented.
[0010] In a third aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the water inspection method described above are implemented.
[0011] The present application provides a water conservancy inspection method, device and storage medium based on a visual language large model. The present application converts the encoding format of multiple first images taken of the water area to be inspected to obtain multiple second images suitable for the visual language large model, and generates question statements for asking questions to the visual language large model based on multiple keywords of interest to water conservancy inspection in a preset knowledge base. The multiple second images and question statements are then input into the visual language large model for processing to obtain water conservancy inspection results. Through the multiple keywords of interest to water conservancy inspection, the visual language large model can be guided to focus on the key detection points of water conservancy inspection, so that the water conservancy inspection results can be obtained efficiently and accurately, thereby greatly improving the efficiency and accuracy of water conservancy inspection. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1 A schematic flow chart of the steps of a water affairs inspection method provided in an embodiment of the present application;
[0014] Figure 2 for Figure 1 A schematic diagram of the sub-step flow of the water affairs inspection method in FIG.
[0015] Figure 3 A schematic diagram of a scenario for implementing the water affairs inspection method provided in this embodiment;
[0016] Figure 4 A schematic block diagram of a water affairs inspection system provided in an embodiment of the present application;
[0017] Figure 5 for Figure 4 Schematic block diagram of the submodules of the water affairs inspection system;
[0018] Figure 6 A schematic block diagram of a computer device provided in an embodiment of the present application.
[0019] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0020] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0021] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.
[0022] Existing water inspection technology has significant shortcomings, primarily in manual inspection methods and technical means. Manual inspections are not only labor-intensive and limited in frequency, but also suffer from inconsistent judgment criteria, fragmented data records, and difficulty ensuring inspection quality in inclement weather.
[0023] While traditional computer vision systems have seen some improvements, they are still overly specialized, lacking semantic understanding of complex scenarios, resulting in poor adaptability and high false positive rates. Dedicated deep learning models, while improving recognition accuracy, are highly data-dependent, lack comprehensive analytical capabilities, have limited generalization capabilities, and have long development cycles, making it difficult to efficiently incorporate water sector expertise.
[0024] The application of existing multimodal technologies in the water sector is still in its infancy. General models lack domain expertise, have low system integration, and lack interpretability of results. They often have high hardware requirements and unfriendly user interfaces, making it difficult to meet the actual work needs of water professionals.
[0025] Therefore, how to improve the efficiency and accuracy of water inspections has become an urgent problem that needs to be solved.
[0026] Based on this, embodiments of the present application provide a water inspection method, device, and storage medium based on a visual language large model. The water inspection method can be applied to computer devices, including terminal devices or servers. The terminal devices can be electronic devices such as mobile phones, tablets, laptops, desktop computers, personal digital assistants, and wearable devices. The server can be a single server or a server cluster consisting of multiple servers.
[0027] The present application provides a water inspection method, device, and storage medium based on a visual language large-scale model. These methods relate to the following technical fields and industrial applications: environmental monitoring and protection, environmental risk warning, water resource management, large-scale modeling, and artificial intelligence visual analysis. Application scenarios include intelligent urban river inspection systems, intelligent identification systems for illegal sewage outlet discharges, and river construction management and violation monitoring systems. The system is applicable to the following industries: municipal water management departments, environmental pollution prevention and control companies, and aerial remote sensing monitoring service providers.
[0028] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.
[0029] Please refer to Figure 1 , Figure 1 A schematic flow chart of the steps of a water inspection method provided in an embodiment of the present application.
[0030] like Figure 1 As shown, the water affairs inspection method includes steps S101 to S104.
[0031] Step S101: Acquire a plurality of first images taken in the water area to be inspected.
[0032] The waters to be inspected can include areas with water bodies such as rivers, lakes, and ditches. The waters to be inspected can be user-defined, for example, a user can set a section of a river as the waters to be inspected. Waters to be inspected can also be automatically set when set conditions are met. For example, if an abnormal water area is detected, the abnormal water area can be automatically set as the waters to be inspected.
[0033] The multiple first images may be captured by a single camera or by multiple cameras. The camera may be movable or fixed. When the multiple first images are captured by multiple cameras, the multiple cameras may be of the same or different categories.
[0034] In one embodiment, the multiple first images are obtained by taking pictures with a shooting device, which may include but is not limited to a camera carried by a drone, a fixed camera, a camera carried by an unmanned boat, a mobile phone, a video camera, and other devices.
[0035] For example, the plurality of first images may be captured by a camera carried by a drone. Specifically, the drone may fly over the water area to be inspected and use the camera carried by the drone to capture the water area to be inspected during flight, thereby obtaining the plurality of first images.
[0036] For example, the multiple first images may be captured by a fixed camera. Specifically, there may be one or more fixed cameras, capable of capturing images of the waters to be inspected around the clock and at all times. The fixed camera may be positioned within the surrounding environment of the waters to be inspected, capturing the waters to be inspected at fixed or rotatable angles to generate the multiple first images.
[0037] Exemplarily, the multiple first images may be captured by different types of camera devices. For example, the multiple first images may be captured by both a camera carried by a drone and a fixed camera, wherein some of the first images may be captured by the camera carried by the drone, and others may be captured by the fixed camera.
[0038] It should be noted that the multiple first images may be sent by the camera to the computer device, thereby enabling the computer device to obtain the multiple first images captured by the camera in the waters to be inspected. Alternatively, the multiple first images captured by the camera in the waters to be inspected may be uploaded to a storage server, and then the computer device may obtain the multiple first images through the storage server.
[0039] In one embodiment, scheduled inspections can be set for the waters to be inspected, thereby acquiring multiple first images captured within the waters to be inspected. For example, the waters to be inspected can be inspected once daily or on multiple occasions. Alternatively, the waters to be inspected can be inspected starting at a preset time, such as at 8:00 AM every weekday.
[0040] In one embodiment, after obtaining multiple first images taken in the water area to be inspected, it also includes: detecting the multiple first images, and selecting multiple candidate images that meet preset conditions from the multiple first images based on the detection results; preprocessing the multiple candidate images, and using the preprocessed candidate images as new first images, wherein the preprocessing includes at least one of denoising, enhancement, and cropping.
[0041] It should be noted that detecting multiple first images may include detecting whether each first image satisfies a preset condition. The preset condition may be determined based on at least one of image quality parameters, resolution, shooting angle, color depth, and number of pixels of the first image. For example, satisfying the preset condition includes at least one of the following: an image quality parameter greater than a preset image quality parameter, a resolution greater than a preset resolution, a shooting angle within a preset angle range, a color depth greater than a preset color depth, and a number of pixels greater than a preset number of pixels.
[0042] It should be noted that by selecting multiple candidate images that meet preset conditions from the multiple first images based on the detection results of the multiple first images and then preprocessing the multiple candidate images, the overall image quality of the multiple first images can be optimized, making the subsequent water conservancy inspection results more accurate, thereby improving the accuracy and reliability of water conservancy inspections. At the same time, since the number of first images is reduced, the efficiency of the subsequent water conservancy inspection results can be effectively improved, thereby improving the efficiency of water conservancy inspections.
[0043] In one embodiment, when detecting multiple first images, if a first image that does not meet a preset condition is detected, the first image that does not meet the preset condition is marked, and the next image after the marked first image is obtained for detection. The marked first image may not be detected again subsequently, thereby avoiding repeated detection.
[0044] In one embodiment, after acquiring multiple first images taken in the water area to be inspected, the method further includes: detecting the multiple first images, and selecting multiple first images that meet preset conditions from the multiple first images based on the detection results. It should be noted that by detecting the multiple first images and thus screening out the multiple first images that meet the preset conditions, invalid images or low-quality images can be filtered out, and high-quality images can be retained. Therefore, the error rate of the subsequent processing of vision-language models (VLMs) can be reduced, and the obtained water conservancy inspection results are more efficient and accurate, thereby improving the efficiency and accuracy of water conservancy inspections.
[0045] In one embodiment, after acquiring multiple first images captured in the waters to be inspected, the method further includes preprocessing the multiple first images and using the preprocessed first images as new first images, wherein the preprocessing includes at least one of denoising, enhancement, and cropping. It should be noted that preprocessing the multiple first images improves the effectiveness of subsequent processing of the first images by the large visual language model, thereby making subsequent water inspection results more accurate, thereby improving the accuracy of water inspections.
[0046] Step S102: Convert the encoding format of the plurality of first images to obtain a plurality of second images suitable for the large visual language model.
[0047] The second image being suitable for the visual language model may refer to the second image having an encoding format suitable for the visual language model. The encoding format of the second image can be determined based on the actual situation of the visual language model. For example, the encoding format of the first image can be PNG or GIF, and the encoding format of the second image can be Base32 or Base64. Examples of visual language models include GPT-4V, Gemini, Qwen-VL series, Doubao-Vision, etc.
[0048] It should be noted that by converting the encoding format of multiple first images, multiple second images suitable for the visual language large model are obtained, which ensures the compatibility of subsequent second images when they are processed by the visual language large model. It also facilitates network transmission and interface calls, thereby improving the reliability of water inspections.
[0049] Exemplarily, the encoding format of multiple first images is converted to obtain multiple second images applicable to the large visual language model, including: obtaining an image encoding format applicable to the large visual language model; and converting the encoding format of multiple first images according to the image encoding format applicable to the large visual language model to obtain multiple second images applicable to the large visual language model.
[0050] Exemplarily, the image encoding format applicable to the large visual language model is base64, and the encoding format of the multiple first images is converted to base64, thereby obtaining the second image applicable to the large visual language model.
[0051] Step S103: Based on multiple keywords of concern to water inspection in a preset knowledge base, generate a question statement for asking questions to the visual language model.
[0052] The preset knowledge base may be, for example, a water affairs professional knowledge base, which may include various indicators related to water affairs monitoring. Multiple keywords concerned by water affairs inspections can be used to represent various indicators related to water affairs monitoring, thereby ensuring comprehensive monitoring of water affairs inspections.
[0053] It should be noted that the method of generating question statements based on the multiple keywords may include rule-based methods and artificial intelligence-based methods. For example, context prediction in a deep learning model may be used to generate question statements from multiple keywords, and the question statements may include question words or question symbols.
[0054] For example, the question statement may be: whether the water surface contains floating objects, the number of floating objects, whether the floating objects are water bottles or garbage belts, whether it includes sewage outlets, whether there are illegal constructions, and the location information of illegal constructions.
[0055] In one embodiment, if Figure 2 As shown, step S103 includes: sub-step S1031 to sub-step S1032.
[0056] Sub-step S1031: Acquire multiple first keywords and multiple second keywords of concern to water affairs inspection in a preset knowledge base.
[0057] The first keyword is used to identify the inspection items of interest to water conservancy inspections, and the second keyword is used to define the inspection rules for the inspection items. For example, the first keyword includes at least one of the following: floating objects on the water surface, sewage outlet status, water quality, illegal construction conditions, and slope conditions. The second keyword includes at least one of the following: detection rules for floating objects on the water surface, detection rules for sewage outlet status, detection rules for water quality, detection rules for illegal construction conditions, and detection rules for slope conditions.
[0058] Sub-step S1032: Based on the multiple first keywords and the multiple second keywords, generate a question sentence for asking a question to the visual language large model.
[0059] Since multiple first keywords and multiple second keywords can include multiple detection items that water inspections are concerned about and detection rules for multiple detection items, the multiple detection items and detection rules for multiple detection items can be used to guide the visual language big model to focus on key detection points of water inspections, thereby obtaining a question statement that includes multiple indicators related to water monitoring and corresponding detection rules. During the subsequent processing by the visual language big model, the question statement can efficiently and accurately obtain the water inspection results, thereby greatly improving the efficiency and accuracy of water inspections.
[0060] It should be noted that the method of generating a question statement based on the multiple first keywords and the multiple second keywords can also include a rule-based method and an artificial intelligence-based method. For example, the context prediction in the deep learning model can be used to merge the multiple first keywords and the multiple second keywords, and add question words or question symbols to generate the question statement.
[0061] In one embodiment, multiple first keywords, multiple second keywords, and multiple third keywords of concern to water inspection are obtained from a preset knowledge base; the first keywords are used to characterize the detection items of concern to water inspection, and the second keywords are used to define the detection rules of the detection items; the third keywords are used to provide reference data for abnormal detection items; based on the multiple first keywords, multiple second keywords, and multiple third keywords, question statements are generated for asking questions to the visual language large model.
[0062] It should be noted that the reference data for abnormal detection items includes, for example, standard parameters and thresholds for water quality and pollutants, thereby supporting the determination of abnormal situations. By combining these multiple third keywords with multiple first and second keywords, the visual language model can also derive abnormal detection items, thereby enriching the water inspection results, greatly improving the monitoring effectiveness of water inspections, and enhancing the user experience.
[0063] Step S104: Input the plurality of second images and the question statement into the visual language model for processing to obtain the water inspection result.
[0064] It should be noted that multiple second images and question statements are input into the visual language model for processing. Since the question statement includes multiple keywords that water conservancy inspection focuses on, it can guide the visual language model to focus on the key detection points of water conservancy inspection, thereby identifying the key detection points that may exist in multiple second images, and can efficiently and accurately obtain water conservancy inspection results, thereby greatly improving the efficiency and accuracy of water conservancy inspection.
[0065] In one embodiment, the water inspection results include a first water inspection result and a second water inspection result. The first water inspection result is a detailed report, while the second water inspection result is a simplified report. The first water inspection result is a detailed report for professional reference. The second water inspection result is a key information report for quick understanding of the situation.
[0066] In one embodiment, the visual language model may include a multimodal large language model (MLLMs), which includes a visual encoder, a multimodal projector, and a large language model (LLM). The plurality of second images are encoded and compressed by a pre-trained visual encoder (such as SimCLR, CLIP, or SigLIP) to obtain a visual representation of the plurality of second images. The visual representation of the plurality of second images is passed to the multimodal projector and aligned with the text representation of the question sentence to obtain projected visual tags and text tags. The projected visual tags and text tags are input into the pre-trained large language model (LLM) for processing together to generate water inspection results.
[0067] In one embodiment, multiple second images and question statements are input into the visual language large model for processing to obtain water inspection results, including: detecting the network connection status; when the network connection status is disconnected, reconnecting the network; when the network connection status is connected, inputting the multiple second images and question statements into the visual language large model for analysis to obtain the analysis results returned by the visual language large model, and parsing the analysis results to obtain the water inspection results.
[0068] It should be noted that multiple second images and question statements can be carried in the request data, and by sending the request data to the visual language large model, multiple second images and question statements can be input into the visual language large model. Before sending the request data, the network connection status can be detected to ensure stable communication with the application interface service of the visual language large model. If the network connection fails (the network connection status is not connected), the network connection is retried; if the connection is normal (the network connection status is connected), the data transmission is continued, so that the visual language large model can perform operations such as analysis on multiple second images and question statements. By detecting the network connection status, the stability in a complex network environment is greatly improved.
[0069] In one embodiment, multiple second images and question statements are input into the visual language model for processing to obtain water inspection results, including: inputting multiple second images and question statements into the visual language model for analysis; checking the response status returned by the visual language model; and judging whether the analysis is successful based on the response status.
[0070] If the response status is abnormal, an error is reported and the multiple second images and question sentences are input into the visual language model for processing. If the response status is normal, the analysis results returned by the visual language model are obtained and parsed to obtain the water inspection results.
[0071] In one embodiment, multiple second images and question statements are input into the visual language model for processing to obtain water inspection results, including: inputting multiple second images and question statements into the visual language model for analysis; obtaining the analysis results returned by the visual language model; and parsing the analysis results to obtain water inspection results.
[0072] It's important to note that the analysis results returned by the visual language model require parsing and verification to ensure the integrity and validity of the analyzed data (water inspection results). For example, the analysis results returned by the visual language model may be in JSON format, which requires parsing and verification using a specialized JSON parsing mechanism. The parsed water inspection results can be structured data, facilitating subsequent query and analysis.
[0073] In one embodiment, parsing the analysis results to obtain a water inspection result includes: performing a first type of analysis on the analysis results to obtain a first water inspection result; and performing a second type of analysis on the analysis results to obtain a second water inspection result. The first type of analysis is a detailed analysis, while the second type of analysis is a simplified analysis. The first water inspection result is a water inspection result that contains detailed information, while the second water inspection result is a water inspection result that contains simplified information. The first water inspection result includes a detailed report of the complete analysis results for reference by professionals. The second water inspection result includes key information from the analysis results, facilitating quick understanding of the situation.
[0074] In one embodiment, after parsing the analysis results to obtain the water inspection results, it also includes: detecting whether there are abnormal detection items in the water inspection results; if there are abnormal detection items in the water inspection results, marking the abnormal detection items and determining the corresponding target water area; and continuously monitoring multiple first images taken of the target water area.
[0075] It should be noted that after receiving the water inspection results, it is necessary to evaluate whether they have found any key issues, such as floating garbage, water pollution, illegal construction, and other abnormal detection items. If abnormalities are found, they will be marked and monitoring of the target water area with abnormalities will continue. If no abnormalities are found or the issues have been confirmed to be recorded, the report will continue to be generated.
[0076] In one embodiment, the method further includes: if the water service inspection results contain abnormal detection items, grading and classifying the detected abnormal detection items to facilitate understanding the severity of the problem. The severity grading and classification can be intuitively displayed in the form of a chart, heat map, etc.
[0077] In one embodiment, the method also includes: if there are abnormal detection items in the water inspection results, multiple indicator parameters in the water inspection results are compared with corresponding multiple historical indicator parameters, so as to facilitate the discovery of the changing trends of each indicator parameter, thereby facilitating the understanding of the changes in related issues.
[0078] In one embodiment, the method further includes: if there are abnormal detection items in the water inspection results, generating and outputting an early warning signal, thereby prompting the user that there are abnormal detection items, so as to facilitate timely resolution of the problem.
[0079] In one embodiment, the method further includes: outputting the water inspection results to a display device, such as a terminal device or a display screen. The water inspection results can be intuitively displayed on the display device in the form of a chart, a heat map, or the like.
[0080] Please refer to Figure 3 , Figure 3A schematic diagram of a scenario for implementing the water inspection method provided in this embodiment.
[0081] like Figure 3 As shown, the drone 10 captures multiple first images of the water area to be inspected and sends the multiple first images to the computer device 20. The computer device 20 converts the encoding format of the multiple first images to obtain multiple second images suitable for the visual language model. It also generates question statements for asking questions to the visual language model based on multiple keywords of interest in water conservancy inspections in a preset knowledge base. The computer device 20 then sends the multiple second images and question statements to the cloud server 30. The cloud server 30 inputs the multiple second images and question statements into the visual language model for processing and obtains analysis results. The cloud server 30 returns the analysis results to the computer device 20. Based on the analysis results, the computer device 20 obtains the water conservancy inspection results.
[0082] The water conservancy inspection method provided in the above embodiment converts the encoding format of multiple first images taken of the water area to be inspected to obtain multiple second images suitable for the visual language large model. Based on multiple keywords of interest to water conservancy inspections in a preset knowledge base, a question statement for questioning the visual language large model is generated. The multiple second images and the question statement are then input into the visual language large model for processing to obtain water conservancy inspection results. By using the multiple keywords of interest to water conservancy inspections, the visual language large model can be guided to focus on key inspection points of water conservancy inspections, thereby efficiently and accurately obtaining water conservancy inspection results, thereby greatly improving the efficiency and accuracy of water conservancy inspections.
[0083] Please refer to Figure 4 , Figure 4 A schematic block diagram of a water inspection system provided in an embodiment of the present application.
[0084] like Figure 4 As shown, the water inspection system 200 includes:
[0085] A photographing device 210 is used to photograph a plurality of first images in the water area to be inspected;
[0086] Computer device 220 is configured to receive a plurality of first images, convert the encoding formats of the plurality of first images to obtain a plurality of second images suitable for use with a large visual language model; and generate, based on a plurality of keywords of interest to water conservancy inspection in a preset knowledge base, a question for questioning the large visual language model.
[0087] The cloud server 230 is configured to receive the plurality of second images and the question sentences, and input the plurality of second images and the question sentences into the visual language model for analysis to obtain analysis results;
[0088] The computer device 220 is also used to generate water inspection results based on the analysis results.
[0089] The camera 210 includes a camera device, which may include a camera-equipped drone, a fixed camera, an unmanned boat equipped with a camera, a mobile phone, a video camera, etc. The cloud server 230 stores a large visual language model. The computer 220 is in communication with the camera 210 and the cloud server 230.
[0090] It should be noted that by generating question statements including multiple keywords of concern to water inspections through computer equipment 220, the visual language model in the cloud server 230 can be guided to focus on key detection points of water inspections, thereby enabling efficient and accurate water inspection results to be obtained, thereby greatly improving the efficiency and accuracy of water inspections.
[0091] In one embodiment, Figure 5 As shown, the water inspection system 200 further includes a user interaction device 240, which is in communication with the computer device 220. The computer device 220 is further configured to send water inspection results to the user interaction device 240, which is configured to display the water inspection results.
[0092] In one embodiment, the user interaction device 240 is further used to display abnormal detection items in the water inspection results.
[0093] In one embodiment, the user interaction device 240 is further configured to output an early warning signal, where the early warning signal is generated based on a detection event indicating that there is an abnormality in the water inspection result.
[0094] In one embodiment, the user interaction device 240 is further configured to output a first water inspection result or a second water inspection result. The first water inspection result is a water inspection result that includes detailed information, while the second water inspection result is a water inspection result that includes simplified information. The first water inspection result includes a detailed report of the complete analysis results for reference by professionals. The second water inspection result includes key information from the analysis results, facilitating quick understanding of the situation.
[0095] It should be noted that technical personnel in the relevant field can clearly understand that for the convenience and conciseness of description, the specific working processes of the modules and units of the water inspection system described above can refer to the corresponding processes in the aforementioned water inspection method embodiment, and will not be repeated here.
[0096] See also Figure 6 , Figure 6 A schematic block diagram of a computer device provided in an embodiment of the present application.
[0097] like Figure 6 As shown, the computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory may include a storage medium and an internal memory, and the storage medium may be non-volatile or volatile.
[0098] The storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any water affairs inspection method.
[0099] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0100] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, the processor can execute any water inspection method.
[0101] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0102] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0103] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:
[0104] Acquire a plurality of first images taken in the water area to be inspected;
[0105] Converting the encoding format of the plurality of first images to obtain a plurality of second images suitable for a large visual language model; and
[0106] Based on multiple keywords of water affairs inspection in a preset knowledge base, generating question sentences for asking questions to the visual language model;
[0107] The plurality of second images and the question statements are input into the visual language model for processing to obtain water inspection results.
[0108] In one embodiment, after acquiring the plurality of first images taken in the water area to be inspected, the processor is further configured to:
[0109] Detecting the plurality of first images, and selecting a plurality of candidate images that meet preset conditions from the plurality of first images according to the detection results;
[0110] Preprocessing is performed on the plurality of candidate images, and the preprocessed candidate images are used as new first images, wherein the preprocessing includes at least one of denoising, enhancing, and cropping.
[0111] In one embodiment, the satisfying of the preset condition includes at least one of the following:
[0112] The image quality parameter is greater than the preset image quality parameter, the resolution is greater than the preset resolution, the shooting angle is within the preset angle range, the color depth is greater than the preset color depth, and the number of pixels is greater than the preset number of pixels.
[0113] In one embodiment, the plurality of first images are captured by a camera carried by a drone, and / or the plurality of first images are captured by a fixed camera.
[0114] In one embodiment, when generating a question statement for asking the visual language model based on a plurality of keywords of interest to water affairs inspection in a preset vocabulary, the processor is configured to implement:
[0115] Acquire multiple first keywords and multiple second keywords of concern to water affairs inspection in a preset knowledge base; the first keywords are used to characterize the inspection items of concern to water affairs inspection, and the second keywords are used to define the inspection rules of the inspection items;
[0116] Based on the plurality of the first keywords and the plurality of the second keywords, a question sentence for asking the visual language macro model is generated.
[0117] In one embodiment, the first keyword includes at least one of the following: floating objects on the water surface, sewage outlet status, water quality, illegal construction status, and slope status;
[0118] The second keyword includes at least one of the following: detection rules for floating objects on the water surface, detection rules for sewage outlet status, detection rules for water quality conditions, detection rules for illegal construction conditions, and detection rules for slope conditions.
[0119] In one embodiment, when inputting the plurality of second images and the question statement into the visual language model for processing to obtain the water affairs inspection result, the processor is configured to implement:
[0120] Detect network connection status;
[0121] When the network connection state is a disconnected state, reconnecting to the network;
[0122] When the network connection status is connected, multiple second images and the question statements are input into the visual language model for analysis to obtain the analysis results returned by the visual language model, and the analysis results are parsed to obtain the water inspection results.
[0123] In one embodiment, after parsing the analysis result to obtain the water inspection result, the processor is further configured to:
[0124] Detecting whether there are any abnormalities in the water affairs inspection results;
[0125] If there are abnormal detection items in the water affairs inspection results, the abnormal detection items are marked and the corresponding target water areas are determined;
[0126] The plurality of first images taken of the target water area are continuously monitored.
[0127] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the computer equipment described above can refer to the corresponding process in the aforementioned water inspection method embodiment, and will not be repeated here.
[0128] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0129] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions. The method implemented when the program instructions are executed can refer to the various embodiments of the water inspection method of the present application.
[0130] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc., equipped on the computer device.
[0131] Furthermore, the computer-usable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.
[0132] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0133] It should also be understood that the term "and / or" used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system that includes the element.
[0134] The serial numbers of the embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments. The above description is only a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A water inspection method based on a visual language large model, characterized in that: include: Acquire a plurality of first images taken in the water area to be inspected; Performing encoding format conversion on the plurality of first images to obtain a plurality of second images suitable for a large visual language model; as well as Based on multiple keywords of water affairs inspection in a preset knowledge base, generating question sentences for asking questions to the visual language model; The plurality of second images and the question statements are input into the visual language model for processing to obtain water inspection results.
2. The water affairs inspection method according to claim 1, characterized in that: After acquiring a plurality of first images taken in the water area to be inspected, the method further includes: Detecting the plurality of first images, and selecting a plurality of candidate images that meet preset conditions from the plurality of first images according to the detection results; Preprocessing is performed on the plurality of candidate images, and the preprocessed candidate images are used as new first images, wherein the preprocessing includes at least one of denoising, enhancing, and cropping.
3. The water affairs inspection method according to claim 2, characterized in that: The predetermined condition includes at least one of the following: The image quality parameter is greater than the preset image quality parameter, the resolution is greater than the preset resolution, the shooting angle is within the preset angle range, the color depth is greater than the preset color depth, and the number of pixels is greater than the preset number of pixels.
4. The water affairs inspection method according to claim 1, characterized in that: The plurality of first images are captured by a camera carried by a drone, and / or the plurality of first images are captured by a fixed camera.
5. The water affairs inspection method according to any one of claims 1 to 4, characterized in that: The generation of question statements for asking questions to the visual language model based on multiple keywords of concern to water affairs inspection in the preset vocabulary includes: Acquire multiple first keywords and multiple second keywords of concern to water affairs inspection in a preset knowledge base; the first keywords are used to characterize the inspection items of concern to water affairs inspection, and the second keywords are used to define the inspection rules of the inspection items; Based on the plurality of the first keywords and the plurality of the second keywords, a question sentence for asking the visual language macro model is generated.
6. The water affairs inspection method according to claim 5, characterized in that: The first keyword includes at least one of the following: floating objects on the water surface, sewage outlet status, water quality, illegal construction status, and slope status; The second keyword includes at least one of the following: detection rules for floating objects on the water surface, detection rules for sewage outlet status, detection rules for water quality conditions, detection rules for illegal construction conditions, and detection rules for slope conditions.
7. The water affairs inspection method according to any one of claims 1 to 4, characterized in that: The step of inputting the plurality of second images and the question statements into the visual language model for processing to obtain water affairs inspection results includes: Detect network connection status; When the network connection state is a disconnected state, reconnecting to the network; When the network connection status is connected, multiple second images and the question statements are input into the visual language model for analysis to obtain the analysis results returned by the visual language model, and the analysis results are parsed to obtain the water inspection results.
8. The water affairs inspection method according to claim 7, characterized in that: After the analysis result is parsed to obtain the water affairs inspection result, the method further includes: Detecting whether there are any abnormalities in the water affairs inspection results; If there are abnormal detection items in the water affairs inspection results, the abnormal detection items are marked and the corresponding target water areas are determined; The plurality of first images taken of the target water area are continuously monitored.
9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the water affairs inspection method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the water affairs inspection method according to any one of claims 1 to 8 is implemented.