Visual management method and system based on large model and AR live-action map

Through visual big models, AR real-life maps are generated and natural language processing models are used to analyze instructions, the problem of data integration and operation complexity in traditional video surveillance systems in park management is solved, and fast and flexible monitoring content adjustment and efficient management are achieved.

CN120372037APending Publication Date: 2025-07-25JILIN VISIBLE AGRI TECH DEV CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510240937.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Traditional video surveillance systems lack multi-dimensional data integration and processing capabilities in park management, and are complex in operations, making it difficult to quickly and in real time to visually display the park's situation. They lack flexibility and intelligence, and cannot respond efficiently to management needs.

Method used

The pre-constructed visual big model is used to splice, identify and annotate multi-dimensional image data, generate the initial AR real-life map, and analyze preset instructions through natural language processing the big model, adjust the AR real-life map to achieve visual management.

Benefits of technology

It achieves rapid and real-time updates and displays of the target area, and can flexibly adjust monitoring content according to management needs, improving management efficiency and intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372037A_ABST
    Figure CN120372037A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data visualization, in particular to a visual management method and system based on a large model and an AR live-action map, and the method comprises the steps: carrying out the splicing, recognition and marking of multi-dimensional image data through a visual large model, obtaining a plurality of processing results, generating an initial AR live-action map, and displaying the initial AR live-action map through a visual terminal; and analyzing the preset instruction by using the natural language processing large model to obtain a target analysis result, adjusting at least one element in the initial AR live-action map according to the target analysis result to obtain a target AR live-action map, and updating display of the visual terminal to realize visual management. According to the method and the device, comprehensive and visual display of the target area is realized, the monitoring picture of the target area can be quickly updated and displayed in real time, the monitoring content can be flexibly adjusted according to the management requirement, and the management efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the continuous improvement of the management requirements of modern parks, the video surveillance system plays a crucial role in park management. Although the traditional video surveillance system can provide real-time image data, there are still some obvious deficiencies in practical applications. First, the traditional surveillance system lacks effective multi-dimensional data integration and processing capabilities, resulting in the difficulty of intuitively and comprehensively displaying the actual situation of the target area in the surveillance screen. Second, the operation and scheduling of the surveillance system are usually relatively complex, and managers need to spend a lot of time on manual adjustment and viewing, making it difficult to quickly and real-time understand the park situation. In addition, the existing surveillance systems often lack flexibility and intelligence when facing complex instructions and cannot efficiently respond to management requirements.

[0003] Therefore, there is an urgent need to provide a technical solution to solve the above problems. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a visualization management method and system based on a large model and an AR real-world map.

[0005] In a first aspect, the present invention provides a visualization management method based on a large model and an AR real-world map. The technical solution of this method is as follows:

[0006] Using a pre-constructed visual large model, splice, identify, and label the multi-dimensional image data of the obtained target area to obtain multiple processing results of the multi-dimensional image data, and generate an initial AR real-world map, and display the multiple processing results through a visualization terminal;

[0007] Using a pre-constructed natural language processing large model, parse a preset instruction to obtain a target parsing result, and adjust at least one element in the initial AR real-world map according to the target parsing result to obtain a target AR real-world map, and update the display of the visualization terminal to achieve visualization management.

[0008] The beneficial effects of a visualization management method based on a large model and an AR real-world map of the present invention are as follows:

[0009] The method of the present invention realizes a comprehensive and intuitive display of the target area, can not only quickly and real-time update and display the surveillance screen of the target area, but also flexibly adjust the surveillance content according to management requirements, improving management efficiency.

[0010] On the basis of the above solution, a visualization management method based on a large model and an AR real-world map of the present invention can also be improved as follows.

[0011] In an optional manner, it further includes:

[0012] Collect the multi-dimensional image data through a variety of data collection devices set within the target area;

[0013] Among them, the variety of data collection devices includes at least one of: AR panoramic eagle eye, AR pan-tilt head, and AR dome camera.

[0014] In the above optional method, by applying a variety of data collection devices, the comprehensiveness and diversity of multi-dimensional image data collection are ensured, the quality and coverage of multi-dimensional image data acquisition are improved, thereby enhancing the monitoring ability and monitoring accuracy.

[0015] In an optional method, using a pre-constructed large vision model, the steps of splicing, recognizing, and annotating the multi-dimensional image data of the target area obtained to obtain multiple processing results of the multi-dimensional image data and generating an initial AR live map include:

[0016] Use the SIFT algorithm or SURF algorithm in the large vision model to extract and match the image features in the multi-dimensional image data to obtain the paired feature points after matching;

[0017] According to the paired feature points, use perspective transformation or thin plate spline interpolation in the large vision model to align the images from different perspectives in the multi-dimensional image data to obtain an initial spliced image;

[0018] Use multi-band fusion or Poisson fusion in the large vision model to process the color difference and brightness difference at the boundary of the initial spliced image to obtain a seamless target spliced image;

[0019] Detect different target objects in the target spliced image, identify the category, current position, and current quantity of each target object to obtain target object information;

[0020] Classify each pixel in the target spliced image to identify the boundaries and attributes of different scene areas to obtain scene area information;

[0021] According to the target object information and the scene area information, annotate different target objects and different scene areas to obtain multiple processing results including category labels, attribute information, and spatial coordinates, and combine the target spliced image and the multiple processing results to generate the initial AR live map.

[0022] In the above optional method, the image processing steps using the large vision model ensure the efficient splicing and alignment of multi-dimensional image data, improve the visual quality of the images, achieve the precise recognition and annotation of target objects and scene areas, thereby improving the accuracy and practicality of the AR live map.

[0023] In an alternative approach, the step of parsing a preset instruction using a pre - built large - model of natural language processing to obtain a target parsing result includes:

[0024] Using the large - model of natural language processing to perform text pre - processing on the preset instruction, segmenting the text in the preset instruction into individual words or phrases, assigning part - of - speech tags to each word or phrase, and removing meaningless words from the text to obtain a first parsing result;

[0025] Performing syntactic analysis and semantic analysis on the preset instruction, extracting the text structure and text meaning of the preset instruction to obtain a second parsing result;

[0026] Based on the first parsing result and the second parsing result, performing intent recognition on the preset instruction to determine the target operation and target parameters of the preset instruction, and obtaining the target parsing result.

[0027] In the above - mentioned alternative approach, the instruction parsing steps using the large - model of natural language processing ensure the accurate understanding and execution of the preset instruction, improve the intelligence and response speed, enable management requirements to be efficiently transformed into specific operations, thereby enhancing management flexibility and user experience.

[0028] In a second aspect, the present invention provides a visualization management system based on a large - model and an AR real - scene map. The technical solution of this system is as follows:

[0029] The visualization management system based on a large - model and an AR real - scene map includes: a first processing module and a second processing module;

[0030] The first processing module is used to: use a pre - built visual large - model to splice, recognize, and label the multi - dimensional image data of the target area obtained, obtain multiple processing results of the multi - dimensional image data, generate an initial AR real - scene map, and display the multiple processing results through a visualization terminal;

[0031] The second processing module is used to: use a pre - built large - model of natural language processing to parse a preset instruction to obtain a target parsing result, and according to the target parsing result, adjust at least one element in the initial AR real - scene map to obtain a target AR real - scene map, and update the display of the visualization terminal to achieve visualization management.

[0032] The beneficial effects of a visualization management system based on a large - model and an AR real - scene map of the present invention are as follows:

[0033] The system of the present invention realizes a comprehensive and intuitive display of the target area. It can not only quickly and real-time update and display the monitoring images of the target area, but also flexibly adjust the monitoring content according to management requirements, improving management efficiency.

[0034] Based on the above solution, a visualization management system of the present invention based on a large model and an AR real scene map can also be improved as follows.

[0035] In an optional manner, it further includes: a third processing module;

[0036] The third processing module is used to: collect the multi-dimensional image data through a variety of data collection devices arranged in the target area;

[0037] Among them, the variety of data collection devices includes at least one of an AR panoramic eagle eye, an AR pan-tilt, and an AR dome camera.

[0038] In the above optional manner, by applying a variety of data collection devices, the comprehensiveness and diversity of the multi-dimensional image data collection are ensured, the quality and coverage of the multi-dimensional image data acquisition are improved, thereby enhancing the monitoring ability and monitoring accuracy.

[0039] In an optional manner, the first processing module is specifically used to:

[0040] Use the SIFT algorithm or SURF algorithm in the visual large model to extract and match the image features in the multi-dimensional image data, and obtain the matched feature point pairs;

[0041] According to the feature point pairs, use the perspective transformation or thin plate spline interpolation in the visual large model to align the images from different perspectives in the multi-dimensional image data, and obtain an initial stitched image;

[0042] Use the multi-band fusion or Poisson fusion in the visual large model to process the color difference and brightness difference at the boundary of the initial stitched image, and obtain a seamless target stitched image;

[0043] Detect different target objects in the target stitched image, identify the category, current position and current quantity of each target object, and obtain target object information;

[0044] Classify each pixel in the target stitched image, identify the boundaries and attributes of different scene areas, and obtain scene area information;

[0045] Based on the target object information and the scene area information, information annotation is performed on different target objects and different scene areas to obtain multiple processing results including category labels, attribute information, and spatial coordinates, and the target stitching image and the multiple processing results are combined to generate the initial AR live map.

[0046] In the above optional manner, the image processing steps using the vision large model ensure the efficient stitching and alignment of multi-dimensional image data, improve the visual quality of the images, and achieve the accurate recognition and annotation of target objects and scene areas, thereby improving the accuracy and practicality of the AR live map.

[0047] In an optional manner, the second processing module is specifically configured to:

[0048] Using the natural language processing large model, perform text preprocessing on the preset instruction, segment the text in the preset instruction into individual words or phrases, assign part-of-speech labels to each word or each phrase, and remove the meaningless words in the text to obtain a first parsing result;

[0049] Perform syntactic analysis and semantic analysis on the preset instruction, extract the text structure and text meaning of the preset instruction to obtain a second parsing result;

[0050] According to the first parsing result and the second parsing result, perform intent recognition on the preset instruction to determine the target operation and target parameters of the preset instruction to obtain the target parsing result.

[0051] In the above optional manner, the instruction parsing steps using the natural language processing large model ensure the accurate understanding and execution of the preset instruction, improve the intelligence and response speed, enable the management requirements to be efficiently transformed into specific operations, thereby enhancing the management flexibility and user experience.

[0052] In a third aspect, the technical solution of an electronic device of the present invention is as follows:

[0053] It includes a memory, a processor, and a program stored on the memory and running on the processor. When the processor executes the program, it implements the steps of the visualization management method based on the large model and the AR live map of the present invention.

[0054] In a fourth aspect, the technical solution of a computer-readable storage medium provided by the present invention is as follows:

[0055] Instructions are stored in the computer-readable storage medium. When the computer-readable storage medium reads the instructions, it causes the computer-readable storage medium to execute the steps of the visualization management method based on the large model and the AR live map of the present invention.

[0056] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Description of the Drawings

[0057] The drawings are only used to illustrate the embodiments and are not considered to limit the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0058] Figure 1 is a schematic flow chart of an embodiment of a visualization management method based on a large model and an AR real - scene map of the present invention;

[0059] Figure 2 is a schematic structural diagram of an embodiment of a visualization management system based on a large model and an AR real - scene map of the present invention;

[0060] Figure 3 is a schematic structural diagram of an embodiment of an electronic device of the present invention. Detailed Embodiments

[0061] The following will describe the exemplary embodiments of the present invention in more detail with reference to the drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein.

[0062] Figure 1 shows a schematic flow chart of an embodiment of a visualization management method based on a large model and an AR real - scene map provided by the present invention, which is executed by a controller. As Figure 1 shown, it includes the following steps:

[0063] S1. Using a pre - constructed visual large model, splice, identify and label the multi - dimensional image data of the acquired target area to obtain multiple processing results of the multi - dimensional image data, and generate an initial AR real - scene map, and display the multiple processing results through a visualization terminal. In S1:

[0064] 1) The pre - constructed visual large model refers to a trained large - scale machine learning model dedicated to processing and understanding visual information. The visual large model has functions such as image feature extraction, recognition, splicing, and annotation, and can extract useful information from multi - dimensional image data and perform corresponding processing. The visual large model is mainly used to perform a series of complex processing operations on the acquired multi - dimensional image data, including feature extraction and matching, image splicing, boundary processing, target detection and recognition, scene classification, etc.

[0065] The pre-built visual large model specifically includes:

[0066] Feature extraction and matching algorithms: such as SIFT algorithm and SURF algorithm, which are used to extract image features from multi-dimensional image data and perform matching.

[0067] Image alignment algorithms: such as perspective transformation and thin-plate spline interpolation, etc., which are used to align images from different perspectives and generate an initial stitched image.

[0068] Image fusion algorithms: such as multi-band fusion and Poisson fusion, etc., which are used to process the color and brightness differences at the boundaries of the stitched images and generate a seamless target stitched image.

[0069] Object detection and recognition function: used to detect and recognize different target objects in the image, including category recognition, as well as determination of position and quantity.

[0070] Scene classification function: used to classify different scene regions in the image, and identify the boundaries and attributes of different scene regions.

[0071] 2) Multi-dimensional image data refers to various types of image data obtained from different angles, different times, and different devices. Multi-dimensional image data can include two-dimensional images, three-dimensional point clouds, panoramic images, etc. Multi-dimensional image data is the basic data processed by the visual large model. By stitching, recognizing, and annotating multi-dimensional image data, an initial AR real-scene map can be generated, providing support for subsequent visualization management.

[0072] Multi-dimensional image data specifically includes:

[0073] Two-dimensional images: such as images directly captured by various shooting devices.

[0074] Three-dimensional point cloud data: three-dimensional space data obtained by a laser scanner or a depth camera.

[0075] Panoramic images: such as wide-angle or panoramic images collected by devices such as AR panoramic eagles-eye, AR pan-tilt heads, and AR dome cameras.

[0076] 3) Multiple processing results refer to various information obtained after the visual large model processes multi-dimensional image data, including target object information, scene region information, etc. These processing results are the basis for generating the initial AR real-scene map and are also important data supports for visualization management.

[0077] Multiple processing results specifically include:

[0078] Target object information: including the category, current position, and current quantity of the target object.

[0079] Scene region information: including the boundaries and attributes of different scene regions.

[0080] Information annotation: including category labels, attribute information, and spatial coordinates.

[0081] 4) The initial AR real-world map refers to the preliminary augmented reality (AR) map generated by processing multi-dimensional image data through a visual large model. The initial AR real-world map contains information annotations of target objects and scene areas and is the basis for visual management.

[0082] The initial AR real-world map specifically includes:

[0083] Target stitched image: an image after seamless stitching processing.

[0084] Processing result information: including target object information, scene area information, and information annotation.

[0085] Specifically, the controller uses a pre-built visual large model to stitch, recognize, and annotate the multi-dimensional image data of the acquired target area, obtains multiple processing results of the multi-dimensional image data, generates an initial AR real-world map, and displays the multiple processing results through a visualization terminal.

[0086] S2. Use a pre-built natural language processing large model to parse a preset instruction to obtain a target parsing result, and adjust at least one element in the initial AR real-world map according to the target parsing result to obtain a target AR real-world map, and update the display of the visualization terminal to achieve visual management. In S2:

[0087] 1) The pre-built natural language processing large model refers to a trained large machine learning model that can understand and process human language. The natural language processing large model is trained based on a large-scale dataset, has powerful language understanding and processing capabilities, can parse the input natural language instruction, and output the corresponding structured result. The natural language processing large model is responsible for performing text preprocessing, syntactic analysis, and semantic analysis on the input natural language instruction, thereby extracting the intention of the instruction, and outputting a target parsing result through parsing the instruction for subsequent adjustment operations on the elements in the AR real-world map.

[0088] The natural language processing large model specifically includes:

[0089] Text preprocessing module: responsible for splitting the input preset instruction into words or phrases and removing meaningless words.

[0090] Syntactic analysis module: analyzes the syntactic structure of the preset instruction.

[0091] Semantic analysis module: understands the semantic content of the preset instruction.

[0092] Intention Recognition Module: Determine the operation target and related parameters of the preset instruction.

[0093] 2) The preset instruction refers to a natural language command or instruction input by the user for operating and adjusting elements in the AR live map. The preset instruction needs to be parsed by the natural language processing large model to achieve dynamic management of the map. The preset instruction is the starting point for triggering the initial adjustment of the AR live map. By parsing the preset instruction, the user's operation requirements can be clarified to guide the execution of corresponding operations.

[0094] The preset instructions specifically include:

[0095] Operation type: Such as adding, deleting, moving, modifying, etc.

[0096] Operation object: Such as a certain element or a group of elements in the initial AR live map.

[0097] Operation parameters: Such as specific position coordinates, attribute information, etc.

[0098] 3) The target parsing result is a structured result generated by the natural language processing large model based on the preset instruction, including the parsing information of the preset instruction, such as the target operation and target parameters. The target parsing result directly guides the adjustment of elements in the AR live map. Through clear operation instructions and parameters, the accuracy and effectiveness of map adjustment are ensured.

[0099] The target parsing result specifically includes:

[0100] Target operation: Such as adding a marker, deleting a marker, modifying a marker, moving an object, changing attributes, etc.

[0101] Target parameters: Such as specific position, type, color, size and other attributes.

[0102] 4) The elements in the initial AR live map refer to various visual objects and information contained in the AR live map before any adjustment operations. These elements are generated by the vision large model through processing multi-dimensional image data. The elements in the initial AR live map provide the basic data for map adjustment.

[0103] The elements in the initial AR live map specifically include:

[0104] Target objects: Such as crops, agricultural machinery, irrigation equipment, personnel, etc.

[0105] Scene areas: Such as farmland, greenhouse, agricultural product processing area, storage area, etc.

[0106] Information annotations: Such as category labels (e.g., crop types such as "wheat", "corn"), attribute information (e.g., growth stage, health status), spatial coordinates, etc.

[0107] 5) The target AR real - scene map refers to a new map generated by adjusting the elements in the initial AR real - scene map according to the target parsing result. The target AR real - scene map reflects the execution result of the preset instruction. The target AR real - scene map presents the state after the adjustment of the map elements. Through the updated map, the effect after executing the preset instruction can be visually seen.

[0108] The target AR real - scene map specifically includes:

[0109] Updated target objects: such as crops, agricultural machinery, irrigation equipment, and personnel after position movement or attribute change.

[0110] Updated scene areas: such as modified farmland layouts, greenhouse layouts, agricultural product processing area layouts, and newly added annotation information.

[0111] Updated information annotations: such as new category labels (e.g., updated crop type like "soybean"), new attribute information (e.g., new growth stage, health status), and new spatial coordinates.

[0112] Specifically, the controller further uses a pre - built large natural language processing model to parse the preset instruction, obtain the target parsing result, and adjust at least one element in the initial AR real - scene map according to the target parsing result to obtain the target AR real - scene map, and update the display of the visualization terminal to achieve visual management.

[0113] The technical solution of this embodiment realizes a comprehensive and intuitive display of the target area. It can not only quickly and real - time update and display the monitoring screen of the target area, but also flexibly adjust the monitoring content according to management needs, improving management efficiency.

[0114] In an alternative way, it further includes:

[0115] Collect multi - dimensional image data through a variety of data collection devices set in the target area;

[0116] Among them, the variety of data collection devices includes at least one of an AR panoramic eagle eye, an AR pan - tilt head, and an AR dome camera.

[0117] In this embodiment, the AR panoramic eagle eye is an image collection device with a wide - angle field of view and high resolution. Through multi - lens combination or fisheye lens technology, it can capture images of a large - range area and, combined with augmented reality (AR) technology, superimpose virtual information on the real scene.

[0118] The AR panoramic eagle eye has the following characteristics:

[0119] Wide-area coverage: The AR panoramic eagle eye can cover a large target area at one time, reducing dead angles and blind spots.

[0120] High resolution: Provide high-definition images to ensure that details are clearly visible.

[0121] Augmented reality function: It can superimpose virtual information such as geographical coordinates and target markers on real-time videos.

[0122] Fixed installation: Installed in a fixed position, such as high places like tall buildings and lamp posts, to obtain the best view.

[0123] The AR pan-tilt is an image acquisition device with an adjustable angle, equipped with a high-power zoom lens, which can rotate horizontally and vertically through remote control to achieve dynamic tracking and detail capture of the target area, and combines augmented reality technology for information superposition.

[0124] The AR pan-tilt has the following characteristics:

[0125] Flexible rotation: The AR pan-tilt can achieve 360-degree horizontal and 180-degree vertical rotation to ensure full coverage of the target area.

[0126] High-power zoom: Provide high-power optical zoom function to magnify details of distant targets.

[0127] Dynamic tracking: It can automatically or manually track moving targets to ensure that the targets are always in the field of view.

[0128] Augmented reality function: Support superimposing virtual information such as path prediction and target trajectory on real-time videos.

[0129] The AR dome camera is an image acquisition device with a spherical shell, equipped with a high-performance camera and a zoom lens, which can achieve omnidirectional image acquisition, and combines augmented reality technology for information display and analysis.

[0130] The AR dome camera has the following characteristics:

[0131] Omnidirectional coverage: The AR dome camera can achieve 360-degree dead-angle-free image acquisition to ensure omnidirectional monitoring of the target area.

[0132] Built-in high-performance camera: Provide high-definition images and stable video streams to ensure image quality.

[0133] Automatic cruise: Support preset cruise paths and can automatically scan the target area periodically.

[0134] Augmented reality function: It can superimpose virtual information such as area division and target recognition on real-time videos.

[0135] Specifically, the controller further collects multi-dimensional image data through a variety of data collection devices set within the target area.

[0136] In an alternative approach, the steps of using a pre-built large vision model to splice, recognize, and label the multi-dimensional image data of the obtained target area, obtaining multiple processing results of the multi-dimensional image data, and generating an initial AR real-world map include:

[0137] Using the SIFT algorithm or SURF algorithm in the large vision model, extract and match the image features in the multi-dimensional image data to obtain pairs of matched feature points;

[0138] According to the pairs of feature points, use perspective transformation or thin plate spline interpolation in the large vision model to align the images from different perspectives in the multi-dimensional image data to obtain an initial spliced image;

[0139] Use multi-band fusion or Poisson fusion in the large vision model to process the color and brightness differences at the boundaries of the initial spliced image to obtain a seamless target spliced image;

[0140] Detect different target objects in the target spliced image, identify the category, current position, and current quantity of each target object to obtain target object information;

[0141] Classify each pixel in the target spliced image to identify the boundaries and attributes of different scene regions to obtain scene region information;

[0142] According to the target object information and scene region information, annotate different target objects and different scene regions to obtain multiple processing results including category labels, attribute information, and spatial coordinates, and combine the target spliced image and the multiple processing results to generate an initial AR real-world map.

[0143] In this embodiment, first, use the SIFT algorithm (Scale-Invariant Feature Transform) or SURF algorithm (Speeded-Up Robust Features) in the large vision model to extract and match the image features in the multi-dimensional image data. The SIFT algorithm can effectively identify key points in the image and generate corresponding feature descriptors, which are invariant to scale and rotation. The SURF algorithm is an accelerated optimization based on SIFT.

[0144] By using the SIFT algorithm or SURF algorithm, it is possible to detect and describe the feature points of the images from different perspectives in the multi-dimensional image data, and then match these feature descriptors to obtain a set of matched feature point pairs, which are the basis for subsequent image alignment and splicing.

[0145] After obtaining the matched feature point pairs, these feature point pairs are used to align the images from different perspectives in the multi-dimensional image data. The large vision model can adopt perspective transformation or thin plate spline (TPS) to achieve image alignment.

[0146] Perspective transformation is mainly used to handle the geometric deformation of images caused by different camera perspectives. Perspective transformation calculates a transformation matrix to accurately map the feature points of one image to the corresponding feature points of another image, thereby correcting the perspective distortion of the image.

[0147] Thin plate spline interpolation is a more complex non-linear transformation method, suitable for situations where there are large deformations in the image. Thin plate spline interpolation not only considers the matching of feature points, but also smooths the overall deformation of the image through an interpolation algorithm.

[0148] By adopting the method of perspective transformation or thin plate spline interpolation, the large vision model can accurately align the images from different perspectives and generate a preliminary stitched image (i.e., the initial stitched image). This step ensures that the positional relationship between the images is accurate during the subsequent stitching process.

[0149] After the image alignment is completed, color differences and brightness differences will appear at the boundaries of the generated initial stitched image. The large vision model uses multi-band blending or Poisson blending technology to process the boundaries of the stitched image.

[0150] Multi-band blending decomposes the initial stitched image into components of different frequencies and fuses the low-frequency and high-frequency parts separately. The low-frequency part contains the overall brightness information of the image, and the high-frequency part contains the details and textures of the image. By processing these frequency components separately, multi-band blending can effectively eliminate the color and brightness differences at the stitching boundaries while retaining the details of the initial stitched image.

[0151] Poisson blending optimizes the fusion process of the initial stitched image by solving a Poisson equation. Poisson blending is based on the gradient field of the image, smoothly transitions colors and brightness in the fusion area, thereby achieving seamless stitching, and can handle the problem of inconsistent lighting, making the initial stitched image look more natural.

[0152] Through the large vision model, a seamless stitched image (i.e., the target stitched image) can be generated. There are no obvious color and brightness differences in the transition area of the target stitched image from different perspectives, ensuring visual consistency.

[0153] After completing the image stitching, the large vision model then uses deep learning-based object detection algorithms, such as YOLO (You Only Look Once), Faster R-CNN, etc., to detect different target objects in the target stitched image, and can efficiently identify various objects from the target stitched image, and give the category, current position (usually represented in the form of a bounding box) and current quantity of each target object, so as to obtain the target object information.

[0154] The steps to obtain the target object information specifically include:

[0155] Input the target stitched image into a pre-trained large vision model based on deep learning, detect the target objects in the target stitched image, and output the category label of each target object (for example, crops, agricultural machinery, irrigation equipment, personnel, etc.) and the corresponding bounding box coordinates.

[0156] According to the bounding box coordinates of the object, calculate the center coordinates of the target object, so as to obtain its current position in the target stitched image.

[0157] Count the quantity of each target object to obtain the current quantity of the target object.

[0158] Finally, output the target object information. The form of the target object information is a list of information. Each target object in the information list has a corresponding category (for example, crop types such as "wheat", "corn"), position and quantity.

[0159] After completing the target object detection, the large vision model uses deep learning-based semantic segmentation algorithms, such as FCN (Fully Convolutional Networks), DeepLab, etc., to classify each pixel in the target stitched image, identify the boundaries and attributes of different scene regions, so as to be able to classify each pixel in the target stitched image and accurately divide different scene regions.

[0160] The specific steps to obtain the scene region information include:

[0161] Input the target stitched image into a pre-trained large vision model based on deep learning semantic segmentation algorithm, classify each pixel in the target stitched image, and output the scene category to which each pixel belongs (for example, farmland, greenhouse, agricultural product processing area, irrigation area, etc.) to obtain the pixel classification result.

[0162] Based on the pixel classification results, generate scene area information including the boundaries and attributes of different scene areas. The form of the scene area information is also a list of information, and each scene area in the information list has corresponding boundaries and attributes. The scene area information is crucial for subsequent information annotation and the generation of the AR real-world map.

[0163] After completing the scene area classification, the visual large model will combine the results of object detection and scene area classification, annotate the stitched image with information, and generate an AR real-world map.

[0164] For information annotation, on the target stitched image, add text labels and icons according to the object information and scene area information, and output the target stitched image after information annotation.

[0165] Among them, for each object, annotate its category, location, and quantity. For the scene area, annotate the category and attributes of the scene area.

[0166] For example, after performing object detection and scene area classification on agricultural vehicles, it is possible to annotate the location information of the agricultural vehicles (such as the eastern area of the farmland), the current crop growth status and the number of irrigation devices in the scene area involved by the agricultural vehicles, etc.

[0167] For the generation of the initial AR real-world map, combine the target stitched image after information annotation with AR technology, and use AR SDKs (such as ARCore, ARKit) to generate the initial AR real-world map. The initial AR real-world map can overlay virtual information in the real scene, providing a richer visual experience. Through AR technology, it can be viewed on mobile devices or AR glasses, and real-time object information and scene area information can be obtained. For example, users can see the crops, greenhouse sheds in the farmland, as well as the annotation information of irrigation devices and agricultural machinery through AR devices.

[0168] Specifically, the controller further uses the visual large model algorithm to extract and match image features, align images from different perspectives to form a stitched image, and then process the boundary differences to achieve seamless stitching. Then, detect and identify the object and scene area information, and perform annotation, and finally generate an initial AR real-world map with rich details.

[0169] In an optional manner, the step of using a pre-built natural language processing large model to parse a preset instruction to obtain a target parsing result includes:

[0170] Use the natural language processing large model to perform text preprocessing on the preset instruction, segment the text in the preset instruction into individual words or phrases, assign part-of-speech tags to each word or each phrase, and remove meaningless words in the text to obtain a first parsing result;

[0171] Performing syntactic analysis and semantic analysis on the preset instruction, extracting the text structure and text meaning of the preset instruction, and obtaining a second parsing result;

[0172] According to the first analysis result and the second analysis result, the preset instruction is intended to be recognized, the target operation and target parameters of the preset instruction are determined, and the target analysis result is obtained.

[0173] In this embodiment, the word segmentation tool in the natural language processing large model is used to segment the text in the preset instruction, and the continuous text is segmented into separate words or phrases. For example, for the input instruction "move the sprinkler equipment to the northern area of the farmland", the word segmentation result is: ["will", "sprinkler equipment", "move", "to", "farmland", "of", "northern area"].

[0174] After the word segmentation operation is completed, the natural language processing model will assign part-of-speech tags to each word or phrase. For example, the part-of-speech tags in the above word segmentation results can be: ["will / adverb", "sprinkler equipment / noun", "move / verb", "to / preposition", "farmland / noun", "of / particle", "northern region / noun"].

[0175] In the text preprocessing stage, meaningless words in the text, such as auxiliary words, adverbs, etc., which have no substantial impact on the core meaning of the instruction, are finally removed to generate the first parsing result. For example, after removing meaningless words in the text, the core words of the preset instruction can be: ["sprinkler equipment", "mobile", "farmland", "northern area"].

[0176] After obtaining the first parsing result, the natural language processing model will further perform syntactic and semantic analysis on the preset instructions to gain a deeper understanding of the structure and meaning of the instructions.

[0177] The purpose of syntactic analysis is to parse the grammatical structure of the preset instructions and clarify the relationship between each word. For example, in the instruction "Move the sprinkler equipment to the northern area of the farmland", the syntactic analysis will recognize that "sprinkler equipment" is the subject of the sentence, "move" is the predicate verb, and "to the northern area of the farmland" is a prepositional phrase, which is the destination of the move. Through syntactic analysis, the natural language processing large model can clarify the hierarchical structure of different parts of the sentence, such as subject, predicate, object, attributive, etc.

[0178] Based on syntactic analysis, semantic analysis interprets the actual meaning of a sentence. The goal of semantic analysis is to understand what operation the user wishes to perform and the objects and locations involved. For example, action: move; object: sprinkler equipment; destination: the northern area of the farmland. Semantic analysis ensures that the natural language processing large model not only understands the structure of the sentence but also the actual intention of the instruction, that is, the user wishes to move the camera to a specific location.

[0179] Through syntactic analysis and semantic analysis, a second parsing result can be obtained, which contains the grammatical structure and semantic information of the preset instruction.

[0180] After completing syntactic analysis and semantic analysis, the natural language processing large model performs intention recognition based on the first parsing result and the second parsing result. The purpose of intention recognition is to clarify the specific operation and relevant parameters to be executed by the preset instruction input by the user.

[0181] Furthermore, the natural language processing large model classifies the instruction according to the parsed semantic information. For example, user instructions can contain the following types of intentions:

[0182] Operation type: such as "move", "scale", "rotate", etc.; target object: such as "sprinkler equipment", "mark", "area", etc.; parameter information: such as destination "the northern area of the farmland", specific coordinates, scale ratio, etc.

[0183] For example, the intention of the preset instruction "Move the sprinkler equipment to the northern area of the farmland" can be classified as: operation type: move (determined by the verb "move"); target object: sprinkler equipment (determined by the noun "sprinkler equipment"); parameter information: the northern area of the farmland (determined by the noun phrase "the northern area of the farmland").

[0184] Based on intention classification, the natural language processing large model can extract the specific operation and parameters for subsequent execution. For example:

[0185] Target operation: Move the camera to the specified location.

[0186] Target parameters (including object and destination), where object: "sprinkler equipment"; destination: "the northern area of the farmland".

[0187] Finally, the natural language processing large model generates a target parsing result, which contains the parsing information of the user instruction, clarifying the operation type, the objects involved, and the relevant parameters of the user instruction.

[0188] In the above process of intention recognition, the large natural language processing model not only recognizes that the operation the user wishes to perform is "move", but also recognizes that the object involved in the operation is "sprinkler equipment", and the destination of the movement is "the northern area of the farmland". The information identified is crucial for subsequent adjustment of the elements in the initial AR real-scene map.

[0189] Specifically, the controller further preprocesses the preset instruction using the large natural language processing model, including text segmentation, part-of-speech tagging, removal of non-significant words, and then syntactic and semantic analysis. Finally, it recognizes the instruction intention, determines the target operation and parameters, and obtains the target parsing result.

[0190] It should be noted that after obtaining the target parsing result, the actual operation can be executed according to the operation type, object, and parameters in the target parsing result. For example, the target parsing result indicates to move the "sprinkler equipment" to the "northern area of the farmland", so the following operations are performed:

[0191] First, parse the location description of "the northern area of the farmland". If there are preset geographical coordinates or area markers, map the "northern area of the farmland" to specific coordinate points or the center of the predefined area. Once the destination is determined, trigger the movement function of the sprinkler equipment in the AR real-scene map, and adjust the position of the virtual sprinkler equipment to the specified position, that is, the "northern area of the farmland". Finally, update and display the adjusted AR real-scene map (i.e., the target AR real-scene map) on the visualization terminal.

[0192] In the initial AR real-scene map, using natural language processing technology can not only achieve intuitive dynamic adjustment (such as changing the position of the sprinkler equipment through simple instructions and immediately observing the map changes), but also directly modify and update the label information in the map. Therefore, it is not limited to physical perspective adjustment, but can also act on deep operations of map data. For example, the user can quickly update the label information by issuing preset instructions such as "change the label of this area to 'irrigation area'" or "mark this area as 'high-yield field'", without manually operating the interface, greatly improving the convenience and efficiency of interaction, making the operation of the AR real-scene map more intelligent and user-friendly. Whether adjusting the perspective or editing information, it can be quickly achieved through natural language instructions. This versatility is not only applicable to observation and navigation, but also can efficiently manage, edit, and update information, meeting the diverse application scenario requirements, such as agricultural management, irrigation control, and crop monitoring.

[0193] Figure 2 Fig. shows a schematic structural diagram of an embodiment of a visualization management system 200 based on a large model and an AR real-scene map provided by the present invention. As Figure 2 shown, the system 200 includes: a first processing module 210 and a second processing module 220;

[0194] The first processing module 210 is configured to: use a pre-built large visual model to splice, recognize, and label the multi-dimensional image data of the acquired target area, obtain multiple processing results of the multi-dimensional image data, generate an initial AR live map, and display the multiple processing results through a visualization terminal;

[0195] The second processing module 220 is configured to: use a pre-built large natural language processing model to parse a preset instruction, obtain a target parsing result, and adjust at least one element in the initial AR live map according to the target parsing result to obtain a target AR live map, and update the display of the visualization terminal to achieve visualization management.

[0196] In an optional manner, it further includes: a third processing module;

[0197] The third processing module is configured to: collect multi-dimensional image data through a variety of data collection devices set in the target area;

[0198] Among them, the variety of data collection devices includes at least one of an AR panoramic eagle eye, an AR pan-tilt, and an AR dome camera.

[0199] In an optional manner, the first processing module 210 is specifically configured to:

[0200] Use the SIFT algorithm or SURF algorithm in the large visual model to extract and match the image features in the multi-dimensional image data to obtain the matched feature point pairs;

[0201] According to the feature point pairs, use perspective transformation or thin plate spline interpolation in the large visual model to align the images from different perspectives in the multi-dimensional image data to obtain an initial spliced image;

[0202] Use multi-band fusion or Poisson fusion in the large visual model to process the color difference and brightness difference at the boundary of the initial spliced image to obtain a seamless target spliced image;

[0203] Detect different target objects in the target spliced image, identify the category, current position, and current quantity of each target object to obtain target object information;

[0204] Classify each pixel in the target spliced image to identify the boundaries and attributes of different scene areas to obtain scene area information;

[0205] According to the target object information and the scene area information, perform information annotation on different target objects and different scene areas to obtain multiple processing results including category labels, attribute information, and spatial coordinates, and combine the target spliced image and the multiple processing results to generate an initial AR live map.

[0206] In an alternative manner, the second processing module 220 is specifically configured to:

[0207] Use a large natural language processing model to perform text preprocessing on a preset instruction, segment the text in the preset instruction into individual words or phrases, assign part-of-speech tags to each word or each phrase, and remove meaningless words in the text to obtain a first parsing result;

[0208] Perform syntactic analysis and semantic analysis on the preset instruction, extract the text structure and text meaning of the preset instruction to obtain a second parsing result;

[0209] According to the first parsing result and the second parsing result, perform intent recognition on the preset instruction, determine the target operation and target parameters of the preset instruction to obtain a target parsing result.

[0210] The technical solution of this embodiment realizes a comprehensive and intuitive display of the target area. It can not only quickly and real-time update and display the monitoring screen of the target area, but also flexibly adjust the monitoring content according to management requirements, improving management efficiency.

[0211] For the parameters and steps of each module in the above-mentioned visualization management system 200 based on a large model and an AR real-world map in this embodiment to implement corresponding functions, reference can be made to the parameters and steps in the embodiment of the visualization management method based on a large model and an AR real-world map in the above text, which will not be elaborated here.

[0212] As Figure 3 shown, an electronic device 300 according to an embodiment of the present invention, the electronic device 300 includes a processor 320, the processor 320 is coupled to a memory 310, and at least one computer program 330 is stored in the memory 310. The at least one computer program 330 is loaded and executed by the processor 320 so that the electronic device 300 implements any one of the above-mentioned visualization management methods based on a large model and an AR real-world map. Specifically:

[0213] The electronic device 300 can vary significantly due to different configurations or performances. It may include one or more processors 320 (Central Processing Units, CPUs) and one or more memories 310. Among them, at least one computer program 330 is stored in the one or more memories 310, and the at least one computer program 330 is loaded and executed by the one or more processors 320, so that the electronic device 300 can implement any one of the visualization management methods based on the large model and the AR real scene map provided in the above embodiments. Of course, the electronic device 300 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The electronic device 300 may also include other components for implementing the device functions, which will not be elaborated here.

[0214] In an embodiment of the present invention, a computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by a processor so that the computer can implement any one of the visualization management methods based on the large model and the AR real scene map.

[0215] Optionally, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0216] In an exemplary embodiment, a computer program product or a computer program is also provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes any one of the visualization management methods based on the large model and the AR real scene map.

[0217] It should be noted that the terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and represent a limitation on a specific order or sequence. In appropriate cases, the order of use of similar objects can be interchanged so that the embodiments of the present application described here can be implemented in an order other than the illustrated or described order.

[0218] Those skilled in the art of the present technology know that the present invention can be implemented as a system, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: it can be completely hardware, can be completely software (including firmware, resident software, microcode, etc.), or can also be in the form of a combination of hardware and software, which is generally referred to as "circuit", "module", or "system" in this article. In addition, in some embodiments, the present invention can also be implemented in the form of a computer program product in one or more computer-readable media, which contain computer-readable program code.

[0219] Any combination of one or more computer-readable media can be adopted. The computer-readable media can be computer-readable signal media or computer-readable storage media. The computer-readable storage media can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage media can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0220] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limitations on the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A visualization management method based on a large model and an AR real scene map, characterized in that, Including: Using a pre - built visual large model, splicing, recognizing, and annotating the multi - dimensional image data of the obtained target area to obtain multiple processing results of the multi - dimensional image data, generating an initial AR live map, and displaying the multiple processing results through a visualization terminal; Using a pre - built natural language processing large model to parse a preset instruction to obtain a target parsing result, and adjusting at least one element in the initial AR live map according to the target parsing result to obtain a target AR live map, and updating the display of the visualization terminal to achieve visual management.

2. The visualization management method based on a large model and an AR real - scene map according to claim 1, wherein It also includes: Collecting the multi - dimensional image data through a variety of data collection devices set in the target area; Among them, the variety of data collection devices includes at least one of an AR panoramic eagle eye, an AR pan - tilt head, and an AR dome camera.

3. The visualization management method based on a large model and an AR real - scene map according to claim 1, wherein The step of using a pre - built visual large model to splice, recognize, and annotate the multi - dimensional image data of the obtained target area to obtain multiple processing results of the multi - dimensional image data and generate an initial AR live map includes: Using the SIFT algorithm or SURF algorithm in the visual large model to extract and match the image features in the multi - dimensional image data to obtain pairs of feature points after matching; According to the pairs of feature points, using perspective transformation or thin - plate spline interpolation in the visual large model to align the images from different perspectives in the multi - dimensional image data to obtain an initial spliced image; Using multi - band fusion or Poisson fusion in the visual large model to process the color difference and brightness difference at the boundary of the initial spliced image to obtain a seamless target spliced image; Detecting different target objects in the target spliced image, identifying the category, current position, and current quantity of each target object to obtain target object information; Classifying each pixel in the target spliced image to identify the boundaries and attributes of different scene areas to obtain scene area information; According to the target object information and the scene area information, performing information annotation on different target objects and different scene areas to obtain multiple processing results including category labels, attribute information, and spatial coordinates, and combining the target spliced image and the multiple processing results to generate the initial AR live map.

4. The visualization management method based on a large model and an AR real-world map according to claim 3, characterized in that, The step of using a pre - built natural language processing large model to parse a preset instruction to obtain a target parsing result includes: Using the natural language processing large model to perform text pre - processing on the preset instruction, splitting the text in the preset instruction into individual words or phrases, assigning part - of - speech labels to each word or phrase, and removing meaningless words in the text to obtain a first parsing result; Performing syntactic analysis and semantic analysis on the preset instruction to extract the text structure and text meaning of the preset instruction to obtain a second parsing result; According to the first parsing result and the second parsing result, performing intention recognition on the preset instruction to determine the target operation and target parameters of the preset instruction to obtain the target parsing result.

5. A visualization management system based on a large model and an AR real-world map, characterized in that, Including: A first processing module and a second processing module; The first processing module is used to: utilize a pre-built visual large model to splice, recognize, and label the multi-dimensional image data of the acquired target area, obtain multiple processing results of the multi-dimensional image data, generate an initial AR real-world map, and display the multiple processing results through a visualization terminal; The second processing module is used to: utilize a pre-built natural language processing large model to parse a preset instruction, obtain a target parsing result, and adjust at least one element in the initial AR real-world map according to the target parsing result to obtain a target AR real-world map, and update the display of the visualization terminal to achieve visualization management.

6. The visualization management method based on the large model and the AR real scene map according to claim 5, wherein, It further includes: A third processing module; The third processing module is used to: collect the multi-dimensional image data through a variety of data acquisition devices arranged in the target area; Among them, the variety of data acquisition devices includes at least one of an AR panoramic eagle eye, an AR pan-tilt, and an AR dome camera.

7. The visualization management method based on a large model and an AR real-world map according to claim 5, characterized in that, Specifically, the first processing module is used to: Utilize the SIFT algorithm or SURF algorithm in the visual large model to extract and match the image features in the multi-dimensional image data to obtain the matched feature point pairs; According to the feature point pairs, utilize the perspective transformation or thin plate spline interpolation in the visual large model to align the images from different perspectives in the multi-dimensional image data to obtain an initial spliced image; Utilize the multi-band fusion or Poisson fusion in the visual large model to process the color difference and brightness difference at the boundary of the initial spliced image to obtain a seamless target spliced image; Detect different target objects in the target spliced image, identify the category, current position, and current quantity of each target object to obtain target object information; Classify each pixel in the target spliced image to identify the boundaries and attributes of different scene areas to obtain scene area information; According to the target object information and the scene area information, perform information annotation on different target objects and different scene areas to obtain multiple processing results including category labels, attribute information, and spatial coordinates, and combine the target spliced image and the multiple processing results to generate the initial AR real-world map.

8. The visualization management method based on a large model and an AR real - scene map according to claim 7, wherein, Specifically, the second processing module is used to: Utilize the natural language processing large model to perform text preprocessing on the preset instruction, segment the text in the preset instruction into individual words or phrases, assign part-of-speech tags to each word or phrase, and remove the meaningless words in the text to obtain a first parsing result; Perform syntactic analysis and semantic analysis on the preset instruction to extract the text structure and text meaning of the preset instruction to obtain a second parsing result; According to the first parsing result and the second parsing result, perform intention recognition on the preset instruction to determine the target operation and target parameters of the preset instruction to obtain the target parsing result.

9. An electronic device, characterized in that, The electronic device includes a processor, the processor is coupled to a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by the processor so that the electronic device implements the visualization management method based on a large model and an AR real-world map according to any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium. The at least one computer program is loaded and executed by a processor so that the computer-readable storage medium implements the visualization management method based on a large model and an AR real-world map according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • System and method for connecting AR (Augmented Reality) with virtual scene

    CN117793497A

  • AR virtual scene enhancement method and system, electronic equipment and storage medium

    CN118822911A

  • Semantic SLAM (Simultaneous Localization and Mapping) optimization method for fusing panoramic vision and laser radar

    CN118962716A

  • Map data processing method and device, equipment, storage medium and program product

    CN119025601A

  • Augmented reality-based display method and device, and storage medium

    US20220207811A1