Evaluation method and device for space movement track

By converting spatial movement trajectories and risk maps into image representations and utilizing a visual-language large model for multi-dimensional evaluation, the problem of poor scalability and interpretability in existing technologies is solved, achieving efficient, safe, and interpretable trajectory evaluation in complex environments.

CN121525848APending Publication Date: 2026-02-13LOW-ALTITUDE ECONOMIC BRANCH OF GUANGDONG-HONG KONG-MACAO GREATER BAY AREA DIGITAL ECONOMY RESEARCH INSTITUTE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511599367.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing trajectory processing and evaluation technologies have poor scalability and interpretability in complex and ever-changing flight environments, cannot adapt to different types of trajectories or scenarios, and are not fully functional in complex scenarios involving visual perception and language command coordination.

Method used

By converting spatial movement trajectories and risk maps into image representations and inputting them into a pre-built visual-language large model, the model's cross-modal reasoning capabilities are utilized for multi-dimensional evaluation, generating evaluation results of natural language interpretation.

Benefits of technology

It enables efficient and safe assessment of spatial movement trajectories in complex environments, provides multi-dimensional assessment results and explains the reasons for the assessment, adapts to different scenarios and trajectories, and does not require readjustment of parameters for each new scenario.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121525848A_ABST
    Figure CN121525848A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for evaluating a space moving track, and relates to the technical field of artificial intelligence. The invention discloses a spatial movement track assessment method, which comprises the following steps of: generating track-risk image data according to a spatial movement track and a corresponding risk map; inputting the trajectory-risk image data and a pre-constructed task cue word into a pre-constructed visual-language large model to obtain an evaluation result; and transmitting the evaluation result to the target end. According to the technical scheme provided by the embodiment of the invention, the space movement track and the risk map which need to be evaluated are uniformly converted into the image representation as the input of the vision-language large model, and the image representation and the task cue word are input into the vision-language large model; the general knowledge and the cross-modal reasoning ability of the vision-language large model are utilized to finish track evaluation in real time, and the expansibility and the interpretability are good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to a method and apparatus for evaluating spatial movement trajectories. Background Technology

[0002] Currently, flight planning models are a key technology for ensuring that the trajectories output by flight planning models meet expected performance and safety standards in practical applications. With the rapid development of UAV technology, their accuracy and reliability are of paramount importance. However, existing trajectory processing and evaluation technologies still have limitations in complex and ever-changing flight environments.

[0003] Most traditional trajectory processing and evaluation techniques are designed for specific types or requirements of trajectory data, and are not entirely applicable to other types of trajectory data or specific trajectory needs. Furthermore, even when using algorithms that match the type or requirements for trajectory processing and evaluation, parameters still need to be readjusted for different trajectories or scenarios, resulting in poor scalability. The evaluation results only include scalar quantities such as scores, and cannot be interpreted, leading to poor interpretability.

[0004] Meanwhile, most existing data quality assessment systems focus on measuring data quality, and their functionality for supervising flight planning models, especially in complex scenarios involving the coordination of visual perception, verbal commands, and flight maneuvers, is insufficient. They are also highly dependent on specific big data environments and technologies, and when processing trajectory data with special formats or characteristics, users may need to perform additional data preprocessing or adjust assessment parameters.

[0005] In summary, existing technologies suffer from poor scalability and poor interpretability. Summary of the Invention

[0006] Based on this, this application provides a method and apparatus for evaluating spatial movement trajectories, achieving evaluation of spatial movement trajectories with good scalability and interpretability.

[0007] According to one aspect of this application, a method for evaluating spatial movement trajectories is proposed, comprising: generating trajectory-risk image data based on the spatial movement trajectory and the corresponding risk map; inputting the trajectory-risk image data and pre-constructed task prompts into a pre-constructed visual-language large model to obtain evaluation results; and transmitting the evaluation results to the target end.

[0008] According to some embodiments, trajectory-risk image data is generated based on the spatial movement trajectory and the corresponding risk map, including: obtaining the target risk map of the area corresponding to the spatial movement trajectory; projecting the coordinates of the spatial movement trajectory onto the corresponding pixel position on the target risk map to obtain the projection curve; and superimposing the target risk map and the projection curve to obtain the trajectory-risk image data.

[0009] According to some embodiments, obtaining a target risk map of the area corresponding to the spatial movement trajectory includes: obtaining risk information of the area corresponding to the spatial movement trajectory; converting the risk information of the area corresponding to the spatial movement trajectory into a two-dimensional image to generate a target risk map, wherein the two-dimensional image includes a binarized image and a grayscale image.

[0010] According to some embodiments, trajectory-risk image data and pre-built task prompts are input into a pre-built visual-language large model to obtain evaluation results, including: preprocessing the trajectory-risk image data and task prompts to obtain model input data; inputting the model input data into the visual-language large model to obtain evaluation results, wherein the evaluation results include evaluation data in natural language format.

[0011] According to some embodiments, inputting model input data into a visual-language large model and obtaining evaluation results includes: inputting model input data into a visual-language large model, reasoning on the model input data based on the prior knowledge embedded in the visual-language large model, and obtaining streaming data output in real time by the visual-language large model; and using the streaming data as the evaluation result.

[0012] According to some embodiments, the trajectory-risk image data and task prompt words are preprocessed to obtain model input data, including: converting the format of the trajectory-risk image data to obtain visual input data; encoding the task prompt words to obtain text input data; and obtaining model input data based on the visual input data and text input data.

[0013] According to some embodiments, transmitting the evaluation results to the target end includes: transmitting the evaluation results to the target end in real time based on a preset streaming protocol.

[0014] According to some embodiments, the method further includes: visualizing the evaluation results and displaying the results of the visualization on the target device.

[0015] According to some embodiments, task prompts include evaluation dimensions.

[0016] According to one aspect of this application, a spatial movement trajectory evaluation device includes: a data processing module for generating trajectory-risk image data based on the spatial movement trajectory and the corresponding risk map; a model evaluation module for inputting the trajectory-risk image data and pre-built task prompts into a pre-built visual-language large model to obtain an evaluation result; and a result transmission module for transmitting the evaluation result to a target end.

[0017] According to one aspect of this application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs; and, when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the method as described above.

[0018] According to one aspect of this application, a computer-readable medium is provided that stores a computer program or instructions thereon, which, when executed by a processor, implement the method as described above.

[0019] Through the above embodiments provided in this application, the spatial movement trajectory and risk map that need to be evaluated are uniformly converted into image representations, which are used as inputs to the visual-language big model. The image representations and task prompts are input into the visual-language big model, and the general knowledge and cross-modal reasoning capabilities of the visual-language big model are used to complete the evaluation of the trajectory in real time. It has good scalability and interpretability. Attached Figure Description

[0020] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this application.

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings, without exceeding the scope of protection claimed by this application.

[0022] Figure 1 A flowchart illustrating the method for evaluating spatial movement trajectories provided in this application embodiment; Figure 2 A flowchart for generating trajectory-risk image data based on spatial movement trajectory and corresponding risk map provided in this application embodiment; Figure 3 This is a schematic diagram of trajectory-risk image data provided in an embodiment of this application; Figure 4 A flowchart illustrating how model input data is input into a large visual-language model to obtain evaluation results, provided in an embodiment of this application; Figure 5 A flowchart illustrating the process of preprocessing trajectory-risk image data and task prompts to obtain model input data, as provided in this embodiment of the application. Figure 6 A block diagram of a spatial movement trajectory evaluation device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0025] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0026] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0027] It should be understood that although the terms first, second, third, etc., may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Therefore, the first component discussed below may be referred to as the second component without departing from the teachings of this application. As used herein, the term "and / or" includes all combinations of any one and more of the associated listed items.

[0028] For specific implementation details, please refer to the following examples.

[0029] Figure 1 A flowchart illustrating the method for evaluating spatial movement trajectories provided in embodiments of this application. Figure 1 As shown, the method includes steps S110-S130.

[0030] In step S110, trajectory-risk image data is generated based on the spatial movement trajectory and the corresponding risk map.

[0031] A spatial movement trajectory is a movement trajectory designed for a space-based mobile device. The space-based mobile device executes the spatial movement trajectory within the spatial domain to achieve spatial movement. The spatial domain includes areas such as low altitude, high altitude, outer space, and ocean, which are not limited in this application.

[0032] Based on this, a spatial movement trajectory is a collection of trajectory points arranged in chronological order, which can be represented as a curve or a sequence. The coordinate values ​​of the trajectory points can be (longitude, latitude).

[0033] Furthermore, the spatial movement trajectory can be an obstacle avoidance trajectory generated by a spatial movement planning model (such as a path planning algorithm like A* algorithm) or a spatial movement trajectory input by the user; this application does not impose any restrictions on this.

[0034] In the specific implementation process, the spatial movement trajectory and the risk map of the area involved in the spatial movement trajectory are visually integrated, and the integration result is used as trajectory-risk image data. The generated trajectory-risk image data is a unified image containing trajectory and risk information, which is used as visual input for the visual-language large model.

[0035] The risk map includes risk information such as obstacles, weather, and electromagnetic interference.

[0036] In step S120, the trajectory-risk image data and pre-built task prompts are input into a pre-built visual-language large model to obtain the evaluation results.

[0037] To meet the need for efficient, safe, and autonomous movement of space mobile devices in complex environments, this application uses a visual-language large model to supervise and evaluate the trajectory of space movement.

[0038] Because the Visual-Language Model (VLM) embeds image understanding and world knowledge, a single prompt can be used across cities, seas, and scenes, and can usually provide a direct natural language explanation (why it is abnormal and how to improve it). This application utilizes the Visual-Language Model to deeply integrate language instruction parsing, risk text interpretation, and spatial movement action assessment functions to comprehensively evaluate the smoothness, collision risk, and rationality of a trajectory, and output the evaluation results.

[0039] This application is particularly applicable to the evaluation of spatial movement trajectories output by spatial mobility planning models (such as low-altitude flight planning models), thereby utilizing large language models for zero-shot inference and feedback correction of the spatial mobility planning model. The multimodal supervision mechanism proposed in this application not only enables the spatial mobility planning model to more effectively avoid collisions and optimize spatial movement paths, but also comprehensively ensures the safety, effectiveness, and feasibility of the spatial mobility planning model, providing solid technical support for the autonomous movement of space-based mobile devices. It solves the problems of poor adaptability and delayed verification of traditional planning models in open scenarios, providing real-time supervision and optimization support for traffic management and autonomous decision-making of unmanned systems in various spatial domains.

[0040] In the specific implementation process, the generated trajectory-risk image data and task prompts are simultaneously input into VLM. The model performs semantic reasoning based on the instructions and uses the powerful functions of VLM to conduct multi-dimensional performance evaluation of various key indicators of the trajectory, and finally obtains detailed evaluation results.

[0041] It should be explained that the pre-trained Visual-Language Large Model (VLM) is a multimodal large model, which includes a neural network structure and corresponding pre-trained parameters. The pre-trained parameters are trained on a large-scale, high-quality multimodal dataset containing detailed, manually annotated image-text pairs, thus ensuring that the model has excellent cross-modal understanding and inference performance.

[0042] The neural network architecture employs a concise and efficient combined architecture, where a visual encoder processes image input and a language model processes text input, linked by a connection module. This vision-language combined architecture enables the model to handle multimodal inputs and perform inference and generation tasks based on these inputs.

[0043] This application requires no additional training or fine-tuning. It can leverage the general knowledge and cross-modal reasoning capabilities of VLM to integrate traditional geometric metrics and task-level metrics into a semantic evaluation framework, significantly improving the adaptability, safety, and reliability of spatial mobility trajectory planning models in complex traffic environments.

[0044] Furthermore, the pre-built task prompts are task prompts and evaluation metrics designed for VLM based on the features of trajectory-risk image data, aiming to guide VLM to analyze and evaluate specific attributes of the input image.

[0045] According to the example embodiment, the task prompt can be constructed as an instruction containing multi-dimensional evaluation requirements, such as "Analyze the quality (smoothness, obstacle avoidance, precision, path length, adaptability, robustness) of the UAV trajectory (start from red circle point and end at red X mark) in the given scenario where black represents obstacles."

[0046] The setting of task prompts enables the model to perform multi-dimensional performance evaluation of spatial movement trajectories in trajectory-risk images and generate corresponding evaluation results to support risk analysis.

[0047] In step S130, the evaluation results are transmitted to the target device.

[0048] The evaluation results text generated by VLM is transmitted to the target end (such as downstream systems) in real time to ensure the efficiency and low latency of the evaluation process.

[0049] The target end (e.g., a remote terminal or downstream application) receives and parses the streaming data from the evaluation results. The parsed data can be fed back in various forms, constructing a "command-reasoning-feedback" closed loop at the visual and semantic levels, enabling real-time monitoring and response to trajectory risks.

[0050] This application leverages the prior knowledge embedded in the Visual-Language Model (VLM) to achieve universal evaluation of different scenarios, spatial domains, and spatial movement trajectories. Compared to existing technologies (e.g., traditional evaluation methods that are only applicable to specific trajectory types or require repeated parameter adjustments for new datasets), this application eliminates the need to reselect thresholds, distance metrics, or perform physical verification for each new scenario. Users only need to provide a single instruction to achieve efficient application across cities, seas, and scenarios, significantly reducing deployment and maintenance costs. In practical implementation, the spatial movement trajectories and risk maps to be evaluated are uniformly converted into image representations, which are then used as input to the Visual-Language Model. The image representations and task prompts are input into the Visual-Language Model, and the general knowledge and cross-modal reasoning capabilities of the Visual-Language Model are utilized to complete trajectory evaluation in real time, demonstrating good scalability and interpretability.

[0051] According to some embodiments, refer to Figure 2 In step S110, trajectory-risk image data is generated based on the spatial movement trajectory and the corresponding risk map, which can be achieved through steps S210-S230.

[0052] In step S210, a target risk map of the area corresponding to the spatial movement trajectory is obtained.

[0053] Obtain a risk map, including risk information, for the area corresponding to the spatial movement trajectory. Risk information includes information on obstacles, weather, and electromagnetic interference.

[0054] In step S220, the coordinates of the spatial movement trajectory are projected onto the corresponding pixel positions on the target risk map to obtain the projection curve.

[0055] The latitude and longitude coordinates of the spatial movement trajectory are projected onto the corresponding pixel positions on the target risk map. The trajectory sequence is visualized as a curve on the image. Simultaneously, the starting and ending points of the trajectory are marked as "o" and "x" respectively to clearly indicate the direction of movement, such as... Figure 3 As shown.

[0056] In step S230, the target risk map and the projection curve are superimposed to obtain trajectory-risk image data.

[0057] The target risk map from step S210 is overlaid with the visualized trajectory sequence (i.e., projection curve) generated in step S220 to produce a complete visual image containing trajectory and risk information, denoted as trajectory-risk image data. This trajectory-risk image data is used as input to the subsequent VLM model for risk analysis and assessment. Figure 3 As shown, Figure 3 In the diagram, black pixels represent impassable areas, white pixels represent passable areas, curves represent spatial movement trajectories, "o" represents the starting point of the trajectory, and "x" represents the ending point of the trajectory.

[0058] This application summarizes and intuitively represents various information in the space mobility environment, such as buildings, terrain, airflow, wind direction, and electromagnetic interference, on a risk integration map (i.e., trajectory-risk image data), thereby more accurately assessing the distance to various risk factors, such as how many meters away from airflow and how many meters away from obstacles.

[0059] According to some embodiments, in step S210, a target risk map of the area corresponding to the spatial movement trajectory is obtained, which can be specifically implemented through steps S211-S212.

[0060] In step S211, risk information of the area corresponding to the spatial movement trajectory is obtained.

[0061] The risk information includes information such as obstacles, population density, and weather that may pose risks to spatial movement trajectories, which will not be exhaustively listed here.

[0062] The risk information referred to in this application can be obtained directly through public channels or based on a 2D risk map corresponding to the area of ​​the spatial movement trajectory. This application does not impose any restrictions on this.

[0063] In step S212, the risk information of the area corresponding to the spatial movement trajectory is converted into a two-dimensional image to generate a target risk map. The two-dimensional image includes a binarized image and a grayscale image.

[0064] Based on a 2D map or 2D risk map of the area corresponding to the spatial movement trajectory, the risk information is abstracted into a two-dimensional image.

[0065] Two-dimensional images can be binarized images or grayscale images. According to an example embodiment, in a binarized image, a pixel value of 0 represents an impassable area, and 1 represents a passable area; in a grayscale image, the grayscale value of a pixel is positively correlated with the risk level of the location.

[0066] According to some embodiments, in step S120, the trajectory-risk image data and the pre-built task prompt words are input into the pre-built visual-language large model to obtain the evaluation result, which can be specifically implemented through steps S121-S122.

[0067] In step S121, the trajectory-risk image data and task prompts are preprocessed to obtain model input data.

[0068] The trajectory-risk image data and task prompts are preprocessed to convert the multimodal data into a unified format that the model can understand, so as to ensure that the model can correctly receive and understand the image and text input. The processed data is recorded as the model input data.

[0069] In practice, according to the example implementation, a multimodal processor is configured in the visual-language large model to convert multimodal data into a unified format that the model can understand. The input trajectory-risk image data and task prompts are converted into tensor formats required by the model, ensuring consistency with the training data format of the VLM model.

[0070] In step S122, the model input data is input into the visual-language large model to obtain the evaluation results, which include evaluation data in natural language format.

[0071] The multimodal images and prompts included in the model input data are simultaneously input into the multimodal large model. Based on the instructions, the model outputs the text assessment results of the trajectory risk in the form of natural language text.

[0072] Based on the above embodiments, the prompt includes: "Analyze the quality (smoothness, obstacle avoidance, precision, path length, adaptability, robustness) of the UAV trajectory (start from red circle point and end at red X mark) in the given scenario where black represents obstacles." Based on this, the evaluation results of this embodiment include: The UAV trajectory in this scenario exhibits the following characteristics: Smoothness: Low. The path is not continuous, with sudden changes in direction. Obstacle avoidance: Poor. The UAV gets very close to obstacles, especially in the final segment. Precision: Low. The trajectory lacks accuracy and deviates significantly from a straight line. Path length: Moderate. The total distance traveled is relatively long compared to a direct path. Adaptability: Low. The path doesn't seem to adapt well to the changing obstacle landscape. Robustness: Low. The UAV struggles to maintain a consistent course and fails to handle unexpected obstacles. Overall, this trajectory demonstrates poor performance across most quality metrics. It's highly sensitive to obstacles, lacks precision, and doesn't effectively balance path length with adaptability and robustness.

[0073] The visual-language large-scale model for assessment proposed in this application achieves multi-dimensional and comprehensive trajectory evaluation by deeply integrating language instruction parsing, risk image interpretation, and spatial movement action assessment. This model can not only evaluate the smoothness, collision risk, and rationality of the trajectory, but also visualize various information in the spatial movement environment (such as terrain, airflow, and electromagnetic interference) as risk maps, thereby intuitively assessing the relative geometric position of the trajectory and risk factors. This contrasts sharply with traditional methods that only provide a single scalar or focus on data quality.

[0074] Furthermore, by leveraging VLM's natural language generation capabilities, this application can directly provide interpretable evaluation chains, explaining the reasons behind the evaluation results, such as: "Because the flight maintained a high descent rate despite thunderstorms, it violated the company's standard operating procedures (SOP)." This interpretability fills the gap in traditional methods' inability to explain evaluation results, enabling users to intuitively understand the problem rather than simply receiving an abstract numerical value. This solves the technical problem of traditional evaluation methods only providing numerical results without explaining their causes, enabling the generation of human-readable chains of cause for anomalies, and automatically or semi-automatically optimizing and adjusting the spatial movement trajectory planning model based on the explanation.

[0075] According to some embodiments, refer to Figure 4 In step S122, the model input data is input into the visual-language large model to obtain the evaluation result, which can be achieved through steps S410-S430.

[0076] In step S410, the model input data is input into the visual-language big model, and reasoning is performed on the model input data based on the prior knowledge embedded in the visual-language big model to obtain the streaming data output by the visual-language big model in real time.

[0077] The Visual-Language Large Model (VLM) analyzes the risk map and structured track / action descriptions in the model input data. It leverages the visual common sense and logical reasoning capabilities of the VLM to conduct real-time assessments of the safety, effectiveness, and feasibility of the track and generate textual parameter adjustment feedback suggestions, which are then output as streaming data.

[0078] In the process of generating text evaluation results using a large visual-language model, the model output is generated as streaming data, either token-by-token or line-by-line. This application achieves real-time acquisition of the model output by capturing the streaming data.

[0079] In step S420, streaming data is used as the evaluation result.

[0080] In step S430, the evaluation result text generated by VLM is transmitted to the downstream system in real time in the form of a data stream to ensure the efficiency and low latency of the evaluation process.

[0081] The evaluation results are returned as a real-time data stream in the form of streaming (e.g., by sending Event SSE via a server), providing a real-time and reliable basis for the optimization and improvement of spatial movement trajectories.

[0082] According to some embodiments, refer to Figure 5 In step S121, the trajectory-risk image data and task prompt words are preprocessed to obtain model input data, which can be specifically implemented through steps S510-S530.

[0083] In step S510, the trajectory-risk image data is converted to a new format to obtain visual input data.

[0084] The trajectory-risk image data is preprocessed to convert the raw image data into a tensor format required by the model and make it consistent with the training data format of the VLM model to ensure that the model can correctly receive and understand the image input. The processed image data is recorded as visual input data.

[0085] In step S520, the task prompt words are encoded to obtain text input data.

[0086] The task prompts are preprocessed, including converting the text data into a numerical sequence that the model can recognize (i.e., encoding). The processed text data is recorded as the text input data.

[0087] In step S530, model input data is obtained based on visual input data and text input data.

[0088] Visual input data and text input data are used as input data for the model.

[0089] According to an example embodiment, a multimodal processor is configured in a large vision-language model. The multimodal processor includes a feature extractor and a word segmenter. The feature extractor preprocesses the trajectory-risk image data, and the word segmenter preprocesses the text data.

[0090] This application provides an assessment mechanism that deeply integrates multimodal information (including but not limited to verbal instructions, risk text, and visual perception). This mechanism integrates and maps multi-source heterogeneous information (e.g., terrain data, airflow information, electromagnetic interference data, written instructions, etc.) into a unified risk assessment framework, and uses this framework to comprehensively assess spatial movement trajectories. The spatial movement trajectory assessment method provided in this application transcends the limitations of traditional single-data-dimensional assessments, and utilizes this framework to comprehensively assess the smoothness, collision risk, and reasonableness of the trajectory.

[0091] According to some embodiments, in step S130, the evaluation result is transmitted to the target end, which can be specifically implemented through step S131.

[0092] In step S131, the evaluation results are transmitted to the target terminal in real time based on a preset streaming protocol.

[0093] Specifically, streaming protocols, such as WebSocket or HTTP / 2 Server-Sent Events (SSE), are used to transmit the captured evaluation results data to the target end (such as a remote terminal or downstream application) in real time.

[0094] This transmission method avoids the delay of waiting for the entire result to be generated before sending, thereby improving the system's response speed.

[0095] According to some embodiments, the method further includes step S140.

[0096] In step S140, the evaluation results are visualized and the results of the visualization are displayed on the target device.

[0097] The target device (e.g., a remote terminal or downstream application) receives and parses the streaming data from the evaluation results. The parsed data is then visualized and fed back to the operators at the target device.

[0098] According to the example embodiment, the evaluation progress is obtained by processing the evaluation results, and the evaluation progress is displayed in real time in the graphical user interface (GUI), the evaluation result text is presented to the operator, or it is used as input for further automated decision-making.

[0099] This application enables real-time monitoring and response to trajectory risks through feedback.

[0100] According to some embodiments, task prompts include evaluation dimensions.

[0101] Evaluation dimensions, such as smoothness, obstacle avoidance, accuracy, path length, adaptability, and robustness, can be set according to the actual evaluation situation. This application will not exhaustively list them here.

[0102] This application enables trajectory evaluation and correction across domains and scenarios through a one-time prompt, without the need to readjust parameters or model structure for each new scenario.

[0103] The following describes an apparatus embodiment of this application, which can be used to perform the method embodiment of this application. For details not disclosed in the apparatus embodiment of this application, please refer to the method embodiment of this application.

[0104] Figure 6 A block diagram of an apparatus for evaluating spatial movement trajectories according to an exemplary embodiment is shown.

[0105] Figure 6 The apparatus shown can perform the aforementioned method for evaluating spatial movement trajectories according to embodiments of this application.

[0106] like Figure 6 As shown, the device for evaluating spatial movement trajectories may include: See Figure 6 Referring to the preceding description, the data processing module 610 is used to generate trajectory-risk image data based on the spatial movement trajectory and the corresponding risk map.

[0107] The model evaluation module 620 inputs the trajectory-risk image data and pre-built task prompts into a pre-built visual-language large model to obtain the evaluation results.

[0108] The result transmission module 630 is used to transmit the evaluation results to the target end.

[0109] The device performs functions similar to those described above; other functions are described in the preceding descriptions and will not be repeated here.

[0110] This application discloses an electronic device, including: a processor; and a memory storing a computer program, which, when executed by the processor, causes the processor to execute the above-described instruction generation method.

[0111] For example, refer to Figure 7 , Figure 7 The illustrated electronic device 700 includes a processor 701 and a memory 703. The processor 701 and the memory 703 are connected, for example, via a bus 702. Optionally, the electronic device 700 may also include a transceiver 704. It should be noted that in practical applications, the transceiver 704 is not limited to one type, and the structure of this electronic device 700 does not constitute a limitation on the embodiments of the present invention.

[0112] Processor 701 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in this disclosure. Processor 701 may also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0113] Bus 702 may include a pathway for transmitting information between the aforementioned components. Bus 702 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 702 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 7The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0114] The memory 703 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other storage medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0115] The memory 703 stores application code that executes the present invention, and its execution is controlled by the processor 701. The processor 701 executes the application code stored in the memory 703 to implement the content shown in the foregoing method embodiments.

[0116] Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0117] This application discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, causes the processor to execute an instruction generation method.

[0118] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0119] The above are only some embodiments of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for evaluating spatial movement trajectories, characterized in that The method comprises the following steps: generating trajectory-risk image data according to a spatial movement trajectory and a corresponding risk map; inputting the trajectory-risk image data and a pre-constructed task prompt word into a pre-constructed visual-language large model to obtain an evaluation result; transmitting the evaluation result to a target end.

2. The method of claim 1, wherein, The method for generating trajectory-risk image data according to a spatial movement trajectory and a corresponding risk map comprises the following steps: obtaining a target risk map of a region corresponding to the spatial movement trajectory; projecting coordinates of the spatial movement trajectory to corresponding pixel positions on the target risk map to obtain a projection curve; superimposing the target risk map and the projection curve to obtain trajectory-risk image data.

3. The method of claim 2, wherein, The method for obtaining a target risk map of a region corresponding to the spatial movement trajectory comprises the following steps: obtaining risk information of a region corresponding to the spatial movement trajectory; converting the risk information of the region corresponding to the spatial movement trajectory into a two-dimensional image to generate a target risk map, wherein the two-dimensional image comprises a binary image and a grayscale image.

4. The method of claim 1, wherein, The method for inputting the trajectory-risk image data and a pre-constructed task prompt word into a pre-constructed visual-language large model to obtain an evaluation result comprises the following steps: preprocessing the trajectory-risk image data and the task prompt word to obtain model input data; inputting the model input data into the visual-language large model to obtain an evaluation result, wherein the evaluation result comprises evaluation data in a natural language format.

5. The method of claim 4, wherein, The method for inputting the model input data into the visual-language large model to obtain an evaluation result comprises the following steps: inputting the model input data into the visual-language large model to perform reasoning on the model input data based on prior knowledge embedded in the visual-language large model, to obtain streaming data output by the visual-language large model in real time; taking the streaming data as the evaluation result.

6. The method of claim 4, wherein, The method for preprocessing the trajectory-risk image data and the task prompt word to obtain model input data comprises the following steps: performing format conversion on the trajectory-risk image data to obtain visual input data; encoding the task prompt word to obtain text input data; obtaining model input data according to the visual input data and the text input data.

7. The method of claim 1, wherein, The method for transmitting the evaluation result to a target end comprises the following steps: transmitting the evaluation result to the target end in real time based on a pre-set streaming transmission protocol.

8. An apparatus for evaluating spatial movement trajectories, characterized by The method comprises the following steps: a data processing module is configured to generate trajectory-risk image data according to a spatial movement trajectory and a corresponding risk map; a model evaluation module is configured to input the trajectory-risk image data and a pre-constructed task prompt word into a pre-constructed visual-language large model to obtain an evaluation result; a result transmission module is configured to transmit the evaluation result to a target end.

9. An electronic device, comprising: The method comprises the following steps: one or more processors; a storage device configured to store one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.

10. A computer readable storage medium having stored thereon a computer program or instructions, characterized in that, The computer program or instructions implement the method according to any one of claims 1-7 when executed by a processor.

Citation Information

Cited By

  • Multi-autonomous underwater vehicle formation and obstacle avoidance strategy generation method

    CN122131777A