Instruction processing method and device, electronic equipment and storage medium

By identifying target image data from image data from different perspectives and combining it with 3D representation for in-depth analysis, this method solves the problem of traditional methods failing to fully utilize 3D and semantic information, achieving higher accuracy and efficiency in instruction processing. It is applicable to scenarios such as intelligent driving, mobile robot navigation, and autonomous flight of drones.

CN121860040APending Publication Date: 2026-04-14XIAOMI EV TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Traditional image analysis methods fail to fully utilize the three-dimensional information of images and the semantic information related to instructions, resulting in inference results that are difficult to meet the needs of practical applications in terms of accuracy and reliability.

Method used

By identifying target image data from candidate image data from different perspectives and combining it with 3D representation for in-depth analysis, a pre-trained model is used for feature extraction and fusion, including a 3D feature extraction module, a visual encoding module, a text encoding module, and a feature analysis and processing module, to achieve deep fusion and inference of multiple features.

Benefits of technology

It improves the accuracy and efficiency of command processing, enabling rapid and accurate command processing in scenarios such as intelligent driving, mobile robot navigation, and autonomous drone flight, providing reliable support for decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860040A_ABST
    Figure CN121860040A_ABST
Patent Text Reader

Abstract

The invention provides an instruction processing method and device, electronic equipment and a storage medium, and the method comprises the steps: responding to received input instruction information, and determining target image data from at least one piece of candidate image data according to the instruction information; wherein the at least one piece of candidate image data is image data shot at different visual angles; and based on the instruction information, performing analysis processing on the three-dimensional representation corresponding to the candidate image data and the target image data to obtain a reasoning result corresponding to the instruction information. Analysis is carried out by combining the three-dimensional representation of the candidate image data and the target image data, the two-dimensional and three-dimensional information of the image is fully utilized, and compared with a traditional processing mode which only depends on the two-dimensional image information, the image content can be understood more comprehensively and deeply, and the accuracy and reliability of a reasoning result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the application of image processing technology in the vehicle field, and more particularly to an instruction processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] In the field of image analysis and processing, traditional methods often have many limitations when faced with complex and ever-changing user needs. When analyzing image data to obtain inference results, traditional methods usually process the two-dimensional information of the image in isolation, failing to fully utilize the three-dimensional information of the image and the semantic information related to the instructions. They also fail to fully explore the potential relationships between different types of features, making it difficult for the accuracy and reliability of the inference results to meet the needs of practical applications. Summary of the Invention

[0003] This application aims to at least partially address one of the technical problems in the related art.

[0004] Therefore, this application proposes a method, apparatus, electronic device, and storage medium.

[0005] One embodiment of this application proposes an instruction processing method, including: In response to receiving input instruction information, a target image data is determined from at least one candidate image data according to the instruction information; wherein, the at least one candidate image data is image data taken from different perspectives; Based on the instruction information, the three-dimensional representation corresponding to the candidate image data and the target image data are analyzed and processed to obtain the inference result corresponding to the instruction information.

[0006] It can accurately determine target image data from multi-view candidate image data according to instructions, and then perform in-depth analysis based on 3D representation, improving the accuracy and efficiency of instruction processing. In scenarios with extremely high requirements for real-time performance and accuracy, such as intelligent monitoring and autonomous driving, it can process instructions quickly and accurately, providing strong support for decision-making.

[0007] Optionally, determining the target image data from at least one candidate image data according to the instruction information includes: Extract the azimuth information contained in the instruction information to obtain the extraction result; Based on the extraction results, target image data is determined from at least one candidate image data.

[0008] By accurately extracting the azimuth information from the command information and filtering the target image data accordingly, image data that meets the command azimuth requirements can be precisely located, improving the accuracy of filtering and avoiding erroneous filtering caused by inaccurate azimuth information extraction or improper filtering methods, thus providing a correct data foundation for subsequent analysis and processing.

[0009] Optionally, determining the target image data from at least one candidate image data based on the extraction result includes any one of the following: In response to the instruction information containing the azimuth information, a target viewpoint matching the target azimuth is selected according to the target azimuth corresponding to the azimuth information, and the candidate image data corresponding to the target viewpoint is used as the target image data; In response to the fact that the instruction information does not contain the orientation information, each of the candidate image data is used as the target image data.

[0010] The system can employ different strategies to determine the target image data based on the presence or absence of azimuth information in the command, making it more adaptable and flexible. Whether the command has explicit or insufficient azimuth information, it can rationally select the target image data, ensuring the smooth execution of the command processing flow. When azimuth information is missing, all candidate image data are used as the target image data, guaranteeing comprehensive consideration of all potentially relevant image information. This avoids inaccurate analysis results due to information omissions and improves the system's ability to handle complex commands and incomplete information.

[0011] Optionally, the step of analyzing and processing the 3D representation corresponding to the candidate image data and the target image data based on the instruction information to obtain the inference result corresponding to the instruction information includes: The instruction information, the candidate image data, and the target image data are input into a pre-trained first model to extract the three-dimensional representation corresponding to the candidate image data, and inference is performed based on the instruction information, the three-dimensional representation corresponding to the candidate image data, and the target image data to obtain the inference result corresponding to the instruction information; wherein, the first model includes: a three-dimensional feature extraction module, a visual encoding module, a text encoding module, and a feature analysis and processing module.

[0012] The first model integrates multiple functional modules to achieve deep fusion and system processing of command information, image data, and their 3D representations. These modules work collaboratively, fully exploring their potential connections, and can provide high-precision inference results for intelligent driving, mobile robot navigation, and autonomous drone flight, thereby supporting accurate decision-making.

[0013] Optionally, the step of extracting the three-dimensional representation corresponding to the candidate image data and performing inference based on the instruction information, the three-dimensional representation corresponding to the candidate image data, and the target image data to obtain the inference result corresponding to the instruction information includes: The instruction information is input into the text encoding module for feature extraction to obtain semantic feature data corresponding to the instruction information; The target image data is input into the visual encoding module for feature extraction to obtain the first image feature data corresponding to the target image data; The three-dimensional representation corresponding to the candidate image data is input into the three-dimensional feature extraction module for feature extraction to obtain reference three-dimensional feature data; The reference 3D feature data, the semantic feature data, and the first image feature data are input into the feature analysis and processing module for fusion and decoding to obtain the inference result.

[0014] By employing advanced natural language processing, computer vision, and 3D analysis techniques, key features are meticulously extracted from the 3D representations of command information, target image data, and candidate image data, maximizing the extraction of useful information and minimizing information omissions and redundancy. This enables the system to more accurately understand command intent and perceive the surrounding environment, providing a solid foundation for subsequent reasoning and thus improving the accuracy and reliability of the reasoning results. Through innovative feature fusion and reasoning mechanisms, the inherent connections between different features are fully utilized for rigorous reasoning. The application of attention mechanisms allows the model to adaptively focus on important features, improving the relevance and effectiveness of reasoning. Decoding structures such as multilayer perceptrons accurately map comprehensive features to the reasoning result space, providing reasonable decision-making basis for intelligent driving, mobile robot navigation, and autonomous flight of unmanned aerial vehicles, meeting the stringent requirements for high-precision decision-making in practical applications.

[0015] Optionally, the feature analysis and processing module includes: a decoding submodule, a cross-view 3D geometry enhancement submodule, and a prediction submodule; the step of inputting the reference 3D feature data, the semantic feature data, and the first image feature data into the feature analysis and processing module for fusion and decoding to obtain the inference result includes: The semantic feature data and the first image feature data are concatenated and input into the decoding layer in the decoding submodule for decoding. The feature data output from each decoding layer in the decoding submodule and the reference 3D feature data are input into the cross-view 3D geometry enhancement submodule for fusion to obtain fused 3D feature data. The fused 3D feature data is input into the decoding submodule or the prediction module for processing to obtain the inference result.

[0016] Through the collaborative work of the decoding submodule, the cross-view 3D geometry enhancement submodule, and the prediction submodule, deep fusion of semantic features, image features, and 3D geometric features is achieved. This fusion approach fully leverages the complementary information and complex relationships between different features, enabling the model to more comprehensively and accurately understand scene information, thereby improving the accuracy and reliability of inference results. In scenarios with extremely high requirements for decision-making accuracy, such as intelligent driving, mobile robot navigation, and autonomous drone flight, it can make decisions that are more consistent with the actual situation.

[0017] Optionally, concatenating the semantic feature data and the first image feature data and inputting them into the decoding layer of the decoding submodule for decoding includes: The semantic feature data and the first image feature data are concatenated and input into the starting decoding layer of the decoding module for decoding to obtain the first two-dimensional feature data; wherein, the decoding submodule includes multiple decoded layers in series, and the decoded layers include a starting decoding layer, an ending decoding layer and at least one intermediate decoding layer.

[0018] By concatenating semantic feature data and first image feature data and then decoding them at the initial decoding layer, the complementary information of both can be fully utilized to extract task-related features with precision. This approach avoids the information omission problem that may occur when processing feature data separately, improving the accuracy and completeness of feature extraction. In scenarios such as intelligent driving, mobile robot navigation, and autonomous drone flight, it can more accurately identify target objects, plan paths, or judge environmental states, providing a more reliable basis for subsequent decision-making.

[0019] Optionally, the step of fusing the feature data output from each decoding layer in the decoding submodule with the reference 3D feature data into the cross-view 3D geometry enhancement submodule to obtain fused 3D feature data includes: For each of the decoding layers, the first two-dimensional feature data output by the decoding layer is input into the fusion with the reference three-dimensional feature data to obtain the fused three-dimensional feature data corresponding to the decoding layer.

[0020] Through a series of operations including unfolding, dimensionality compression, attention calculation, and weighted fusion, deep fusion of the first-dimensional feature data and the reference three-dimensional feature data is achieved. This fusion method can fully explore the inherent connections and complementarities between different feature data, making the fused feature data more comprehensive and accurate in reflecting scene information, thereby improving the accuracy and reliability of inference results. In scenarios with extremely high requirements for decision-making accuracy, such as intelligent driving, mobile robot navigation, and autonomous drone flight, it can make decisions that are more in line with the actual situation.

[0021] Optionally, the step of inputting the fused 3D feature data into the decoding submodule or the prediction module for processing to obtain the inference result includes: The fused three-dimensional feature data corresponding to the decoding layer is fused with the first two-dimensional feature data output by the decoding layer, and then input into the next adjacent decoding layer for decoding. The first two-dimensional feature data output by the last decoding layer is fused with the fused three-dimensional feature data corresponding to the last decoding layer, and then input into the prediction submodule for classification to obtain the inference result.

[0022] By fusing feature data from multiple decoding layers, complementary information between features of different levels and types can be fully explored, continuously optimizing feature representation and making the feature data more comprehensive and accurate in reflecting the actual scene. This helps improve the accuracy and reliability of inference results, providing stronger support for agents to make decisions in complex environments. The feature data from the final decoding layer is then fused again with the fused 3D feature data before being input into the prediction submodule, ensuring that the prediction submodule can obtain the most comprehensive and representative feature information.

[0023] Optionally, the step of inputting the first two-dimensional feature data output by the decoding layer into the fused three-dimensional feature data corresponding to the decoding layer by fusing it with the reference three-dimensional feature data includes: The reference three-dimensional feature data is unfolded to obtain the second two-dimensional feature data; The first two-dimensional feature data and the second two-dimensional feature data are dimensionally compressed, and attention is calculated based on the compressed first two-dimensional feature data and the second two-dimensional feature data to obtain an attention matrix; The compressed first two-dimensional feature data and the second two-dimensional feature data are weighted and fused according to the attention matrix, and the dimensions are restored to obtain the fused three-dimensional feature data.

[0024] Through a series of operations including unfolding, dimensionality compression, attention computation, and weighted fusion, deep integration of the first-dimensional feature data and the reference three-dimensional feature data is achieved. This fusion method can fully explore the complementarity and potential connections between the two types of feature data, enabling the fused feature data to more comprehensively and accurately reflect the actual situation of the scene, thereby improving the accuracy and reliability of the inference results. In scenarios with extremely high requirements for decision-making accuracy, such as intelligent driving, mobile robot navigation, and autonomous drone flight, it can provide more reliable decision-making basis for intelligent agents.

[0025] Optionally, the step of obtaining the three-dimensional representation corresponding to the candidate image data includes: Perform 3D modeling based on the candidate image data to obtain an initial 3D representation; Feature extraction is performed on the camera parameters corresponding to the candidate image data to obtain parameter features; The parameter features are fused with the initial three-dimensional representation to obtain the fused three-dimensional representation.

[0026] By optimizing the initial 3D representation by incorporating camera parameter features, the true position, shape, and spatial relationships of objects in a scene can be more accurately reflected. In autonomous driving, this allows for precise determination of the distance and position of obstacles, providing more reliable information for safe driving; in mobile robot navigation, it helps generate more accurate maps and optimize path planning; and in autonomous drone flight, it enables drones to perceive their surroundings more accurately, avoid collisions, and complete tasks efficiently.

[0027] Another embodiment of this application proposes an instruction processing apparatus, including: An image selection module is configured to, in response to receiving input instruction information, determine target image data from at least one candidate image data according to the instruction information; wherein, the at least one candidate image data is image data taken from different perspectives; The instruction processing module is used to analyze and process the three-dimensional representation corresponding to the candidate image data and the target image data based on the instruction information, and obtain the inference result corresponding to the instruction information.

[0028] Optionally, the image selection module includes: The extraction submodule is used to extract the azimuth information contained in the instruction information and obtain the extraction result; An image selection submodule is used to determine target image data from at least one candidate image data based on the extraction results.

[0029] Optionally, the image selection submodule includes: The first selection unit is configured to respond to the instruction information containing the azimuth information, select a target viewpoint matching the target azimuth according to the target azimuth corresponding to the azimuth information, and use the candidate image data corresponding to the target viewpoint as the target image data; The second selection unit is configured to, in response to the instruction information not containing the orientation information, select each of the candidate image data as the target image data.

[0030] Another embodiment of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method described in the foregoing aspect.

[0031] Another embodiment of this application proposes a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the foregoing aspect.

[0032] Another embodiment of this application proposes a chip including processing circuitry configured to perform the method described in one aspect above.

[0033] Another embodiment of this application proposes a computer program product that, when executed by a processor, implements the method described in the foregoing aspect.

[0034] The instruction processing method, apparatus, electronic device, chip, and storage medium proposed in this application can achieve the following beneficial effects: Precise identification of target image data: By responding to input command information and based on the orientation information contained in the command, the target image data can be accurately identified from candidate image data taken from at least one different perspective. This greatly improves the accuracy and efficiency of image screening, meeting users' needs for images from specific perspectives.

[0035] This approach leverages multi-source information for analysis: based on instruction information, it analyzes and processes the 3D representations of candidate image data and target image data. By extracting semantic feature data, image feature data, and 3D feature data, and then fusing and decoding them, it fully utilizes various types of information. Compared to traditional analysis methods that rely solely on 2D image information, this approach provides a more comprehensive understanding of image content and improves the accuracy of inference results.

[0036] Optimized feature fusion and decoding process: Detailed feature fusion and decoding steps, such as concatenating semantic feature data and first image feature data and decoding them in the decoding layer of the inference model, then fusing them with 3D feature data and further decoding, ensure that different feature information can be effectively fused and processed. Especially when fusing 3D feature data, operations such as unfolding, dimensionality compression, and attention calculation are used to fully explore the relationship between 3D and 2D features, improving the inference model's ability to process complex information, thereby obtaining more reliable inference results.

[0037] Improved adaptability to various application scenarios: This method is applicable to a variety of application scenarios with high requirements for image analysis accuracy and real-time performance, such as autonomous driving and robot control. It can quickly and accurately process image data from different perspectives in complex environments, providing a reliable basis for decision-making and enhancing the stability and reliability of the system in practical applications.

[0038] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0039] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating an instruction processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a first model provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an instruction processing device provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a chip proposed in an embodiment of this application. Detailed Implementation

[0040] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0041] The instruction processing method, apparatus, electronic device, chip, and storage medium of this application are described below with reference to the accompanying drawings.

[0042] Figure 1 This is a flowchart illustrating an instruction processing method provided in an embodiment of this application.

[0043] As one implementation, the instruction processing method of this application embodiment can be configured in an instruction processing device, which can be applied to any electronic device so that the electronic device can perform instruction processing functions.

[0044] Among them, electronic devices can be any device with computing capabilities, such as mobile terminals. Mobile terminals can be hardware devices with various operating systems, touch screens and / or displays, such as vehicles, mobile phones, tablets, personal digital assistants, wearable devices, etc.

[0045] As another implementation, the instruction processing method of this application embodiment can also be executed by a chip with processing capabilities. Chips include image signal processing chips (ISP), central processing units (CPU), application-specific integrated circuits (ASIC), digital signal processors (DSP), field-programmable gate arrays (FPGA), systems on a chip (SOC), reduced instruction set computers (RISC), etc., which will not be listed here.

[0046] It should be noted that all data collection operations related to users in this application are conducted with the user's authorization and in strict compliance with relevant laws and regulations such as privacy and security.

[0047] like Figure 1 As shown, the method may include the following steps: Step 101: In response to receiving input instruction information, determine target image data from at least one candidate image data according to the instruction information; wherein, the at least one candidate image data is image data taken from different perspectives; Step 102: Based on the instruction information, analyze and process the three-dimensional representation corresponding to the candidate image data and the target image data to obtain the inference result corresponding to the instruction information.

[0048] In this embodiment, in cutting-edge technology applications such as intelligent driving, mobile robot navigation, and autonomous drone flight, the accuracy and efficiency of image processing are crucial to ensuring the safe and stable operation of the system. With the development of sensor technology, intelligent driving vehicles, mobile robots, and drones are equipped with multiple cameras, capable of acquiring massive amounts of image data from different perspectives. However, how to quickly and accurately extract useful information from this complex image data and perform in-depth analysis according to specific needs has become a bottleneck restricting the further development of these technologies.

[0049] For example, in intelligent driving scenarios, vehicles need to process images from cameras at different angles in real time to identify targets such as roads, traffic signs, other vehicles, and pedestrians. Traditional methods struggle to quickly and accurately filter out key image data when faced with complex and ever-changing road conditions and diverse driver commands, leading to decision delays or errors. For instance, when a driver instructs to check for oncoming traffic in a specific direction, traditional systems may not be able to quickly and accurately locate the target image from numerous candidate images, affecting driving safety. When mobile robots navigate in complex environments, they need to make real-time decisions based on environmental information. Traditional image processing methods cannot effectively integrate multi-view images received by the robot with user commands, leading to deviations in the robot's understanding of the environment, potentially causing collisions with obstacles or failure to accurately reach the target location during navigation. Similarly, drones face the challenge of processing large amounts of image data collected from different perspectives during autonomous flight. When performing tasks such as search and rescue and geographic mapping, traditional methods cannot effectively combine user commands to determine key images from multi-view images, affecting task execution efficiency and accuracy.

[0050] Command information: In scenarios such as intelligent driving, mobile robot navigation, or autonomous drone flight, this information is input by the user (such as a driver or operator) or generated by the system based on the mission objective, used to indicate the direction and specific requirements of image processing. It can be a requirement for identifying a specific target, a viewing requirement from a specific perspective, or a monitoring instruction for a specific area, such as "identify the color of the traffic light 50 meters ahead," "check if there are any obstacles on the left side of the robot," or "the drone searches for signs of life in a certain area."

[0051] Candidate image data: Image data captured from different perspectives by intelligent driving vehicles, mobile robots, or drones using multiple cameras onboard. This data encompasses a wealth of information about the environment surrounding the device, but before filtering, it contains a significant amount of redundant and non-critical information.

[0052] Target image data: Based on the instruction information, the image data selected from the candidate image data that best meets the requirements of the current task. It is the core object of subsequent image processing and analysis, and can provide key information for system decision-making.

[0053] Reasoning Results: By analyzing and processing the 3D representations corresponding to candidate image data and target image data, conclusions corresponding to the instruction information are derived. In intelligent driving scenarios, this might involve judging road conditions and predicting the behavior of other traffic participants; in mobile robot navigation, it might involve determining its own position and planning its direction of travel; in autonomous drone flight, it might involve identifying the mission target and suggesting adjustments to the flight path.

[0054] Once the system receives the instruction, it sets a clear target direction for the image processing flow. At this point, the system filters candidate images captured from numerous different perspectives based on the instruction to determine the target image. This process is similar to finding the most matching image in a vast image database based on a specific index (i.e., the instruction). After determining the target image, the system analyzes the 3D representations of the candidate images and the target image itself based on the instruction. By combining the information from both, it mines the features, relationships, and potential information within the images, ultimately deriving a reasoning result corresponding to the instruction. For example, in an autonomous driving scenario, if the instruction is to determine whether it is safe to cross an intersection ahead, the system first filters the target image data from the perspective of the intersection, while simultaneously constructing a 3D scene representation of the intersection using images from other perspectives. Then, it comprehensively analyzes the 2D features of the target image (such as traffic light status and vehicle positions) and the 3D scene features (such as the distance and relative positional relationship between vehicles and the intersection) to arrive at a reasoning result regarding whether it is safe to cross the intersection.

[0055] Beneficial effects: Precise decision support: This feature enables precise identification of target image data from multi-view image data based on command information. It provides accurate decision-making support for intelligent driving, mobile robot navigation, and autonomous drone flight systems, significantly improving the accuracy and timeliness of decision-making and ensuring the safe and stable operation of the system. For example, intelligent driving vehicles can quickly acquire key images, accurately assess road conditions, and avoid traffic accidents.

[0056] Comprehensive information utilization: By comprehensively analyzing the 3D representations corresponding to candidate image data and target image data, the two-dimensional and three-dimensional information of the images is fully explored. Compared with traditional processing methods that only rely on 2D image information, this approach enables a more comprehensive and in-depth understanding of the scene, improving the system's adaptability to complex environments. For example, mobile robots can better perceive the surrounding 3D spatial environment and plan the optimal path.

[0057] In an intelligent driving scenario, suppose a driver issues the command "Check if there is a vehicle overtaking on the right ahead" while the vehicle is in motion. Upon receiving this command, the system quickly selects the image from the right-hand view of the vehicle as the target image from candidate image data collected by multiple cameras. Simultaneously, it constructs a 3D representation of the vehicle's surrounding environment by combining image data from other perspectives. Feature extraction is performed on the target image data to obtain 2D feature information such as vehicle shape and color, and 3D feature information such as relative positions and speeds between vehicles is extracted from the 3D representation. By analyzing and processing this information, the system infers whether there is a vehicle overtaking on the right ahead and feeds the result back to the driver or the intelligent driving decision-making system to make appropriate driving decisions.

[0058] In a mobile robot navigation scenario, the mobile robot performs goods handling tasks in a warehouse. The operator issues the command, "Detect whether there are obstacles in front of the robot; if so, plan a path to avoid them." The robot's camera collects candidate image data from multiple perspectives. Based on the command, the system extracts the target image data from the front perspective and constructs a 3D model of the warehouse environment. By analyzing the 2D features of the target image (such as the shape and size of the obstacle) and the features of the 3D model (such as the location of the obstacle and its distance from the robot), the system infers whether an obstacle exists in front. If an obstacle exists, it further plans a path to avoid it based on the 3D model, guiding the robot to complete the task safely and efficiently.

[0059] For autonomous drone flight, assuming the drone is performing a forest fire monitoring mission, the operator issues the command "Search for signs of fire in a certain forest area." Multiple cameras on the drone collect images of the forest area from different perspectives as candidate image data. The system filters the target image data from specific areas of the forest area according to the command and constructs a 3D scene of the forest area using images from other perspectives. The target image data and the 3D scene are analyzed to extract 2D features such as smoke color and flame shape, as well as 3D features such as the location of the fire source in 3D space. By comprehensively analyzing these features, an inference result is derived regarding the presence of signs of fire in the forest area, providing crucial information for fire monitoring and rescue.

[0060] Optionally, determining the target image data from at least one candidate image data according to the instruction information includes: Extract the azimuth information contained in the instruction information to obtain the extraction result; Based on the extraction results, target image data is determined from at least one candidate image data.

[0061] In this embodiment, accurately parsing the azimuth information in command information is crucial for determining target image data in intelligent driving, mobile robot navigation, and autonomous drone flight. However, traditional methods have many shortcomings in processing azimuth information. In intelligent driving scenarios, road conditions are complex and changeable, and the driver's commands may contain precise azimuth descriptions, such as "check the traffic sign 45 degrees to the left front." Traditional systems may not be able to accurately extract this azimuth information, resulting in the inability to acquire key images and affecting driving decisions. When mobile robots navigate in complex indoor and outdoor environments, the operator's commands may involve specific locations, such as "check if there are steps on the left side behind the robot." Traditional methods may fail to effectively extract azimuth information, causing the robot to have a distorted perception of the environment, leading to navigation errors. During drone missions, such as power line inspections, if the operator's command is "check the status of the insulator on the right side of the tower ahead," traditional methods may fail to accurately extract azimuth information, preventing the drone from acquiring images from the correct perspective, affecting the accuracy and efficiency of the inspection task.

[0062] Orientation information: In the instruction information of intelligent driving, mobile robot navigation and drone autonomous flight scenarios, it is used to clearly indicate the specific description of the required image view direction, such as "front", "back", "30 degrees to the left", "upper right", etc., which accurately describes the orientation of the image that the user expects to obtain.

[0063] Extraction Result: The location information obtained after parsing and processing the instruction information using specific algorithms and techniques. It may be an explicit location description, or it may indicate that the instruction does not contain location information.

[0064] The system first meticulously analyzes the input command information, using natural language processing techniques or specific information extraction algorithms to extract location information, thus obtaining the extraction result. Subsequently, it determines the target image data based on this extraction result. For example, when the extraction result contains explicit location information, such as "20 degrees to the left of the rear," the system will search for a target viewpoint matching that location in the candidate image data, based on a pre-established correspondence between image viewpoints and locations. Each camera corresponds to a specific viewpoint range when acquiring images; the system finds candidate image data captured from the corresponding viewpoint through matching and identifies it as the target image data. If the extraction result shows that the command does not contain location information, considering that the user may need to perform a comprehensive analysis of the entire scene, the system will use all candidate image data as the target image data, providing a comprehensive data foundation for subsequent integrated analysis.

[0065] Beneficial effects: Precise target image positioning: By accurately extracting the directional information from the command message, the system can accurately filter out target image data that meets the user's directional requirements from candidate image data, improving the targeting and efficiency of image processing and providing accurate image information support for the system's decision-making in complex environments. For example, autonomous vehicles can quickly acquire road condition images from specific locations and react promptly.

[0066] Flexible handling of ambiguous commands: When the command does not contain orientation information, the strategy of using all candidate image data as the target image data enables the system to flexibly respond to ambiguous command situations, ensuring that no potentially useful information is missed, thus enhancing the system's adaptability and reliability under different command conditions. For example, when a mobile robot faces ambiguous commands, it can comprehensively analyze the environmental image to avoid navigation failure due to missing information.

[0067] In a smart driving scenario, when a self-driving car is traveling on a highway, the driver issues the command "Check if there are any vehicles approaching at high speed from the right rear." The system parses the command and extracts the location information of "right rear." Then, based on the correspondence between the vehicle's camera viewpoint and location, it selects the image of the right rear viewpoint from the candidate image data as the target image data. By analyzing this target image data, the system can determine whether there are any vehicles approaching at high speed from the right rear, providing the driver with timely safety alerts.

[0068] In a mobile robot navigation scenario, the robot moves through a hospital corridor. The operator issues the command "Check the surrounding environment," without specifying location information. The system uses all candidate image data captured by the robot's camera as target image data. Subsequently, by analyzing this image data, the robot can gain a comprehensive understanding of its surroundings, including whether there are pedestrians ahead, obstacles on both sides, and oncoming vehicles behind, thus enabling it to navigate safely through the corridor.

[0069] For autonomous drone flight, assuming the drone is conducting an urban topographic mapping task, the operator issues the command "collect images of a certain street," without specifying the location. The drone's camera collects candidate image data from multiple perspectives, and the system uses all these image data as the target image data. Then, through processing and stitching these image data, the drone can generate a comprehensive topographic mapping result for the street, meeting the mission requirements.

[0070] Optionally, determining the target image data from at least one candidate image data based on the extraction result includes any one of the following: In response to the instruction information containing the azimuth information, a target viewpoint matching the target azimuth is selected according to the target azimuth corresponding to the azimuth information, and the candidate image data corresponding to the target viewpoint is used as the target image data; In response to the fact that the instruction information does not contain the orientation information, each of the candidate image data is used as the target image data.

[0071] In this embodiment, the diversity of command information in practical applications of intelligent driving, mobile robot navigation, and autonomous drone flight requires the system to possess high flexibility and accuracy in determining target image data based on location information. Traditional methods often use fixed patterns to process location information, which cannot adapt to different command scenarios. In intelligent driving, the driver's commands may vary depending on road conditions; sometimes it is necessary to precisely view details in a specific location, while at other times a quick overview of the general scene is required. Traditional systems struggle to accurately determine target image data based on these variations, potentially increasing driving risks. Mobile robots perform tasks in different environments, and the operator's command location information is expressed in diverse forms. Traditional methods may fail to effectively understand and respond correctly, affecting the robot's efficiency and safety. When drones perform various complex tasks, such as agricultural plant protection and logistics delivery, the operator's commands may contain ambiguous or precise location requirements. Traditional methods cannot flexibly handle these requirements, potentially leading to task failure.

[0072] Target location: The specific directional position determined by the location information in the instruction information. In scenarios such as intelligent driving, mobile robot navigation and drone autonomous flight, it clarifies the actual spatial location corresponding to the image that the user expects to obtain, such as "30 meters in the northeast" or "5 meters below vertically", and is an important basis for screening target image data.

[0073] Target viewpoint: The image acquisition viewpoint corresponding to the target location. In the camera layout of intelligent driving vehicles, mobile robots and drones, each target location has a specific camera viewpoint to match it. The image captured through this viewpoint can accurately reflect the scene information of the target location.

[0074] When the extracted results contain azimuth information, the system first identifies the target azimuth corresponding to that information. For example, if the command is "view the exterior of a building 20 degrees to the left of the drone," the system determines "20 degrees to the left of the drone" as the target azimuth. Next, based on a pre-built mapping between camera views and azimuth, the system finds the target viewpoint that matches this azimuth. Assuming that camera A of the drone corresponds to this target viewpoint, the candidate image data captured by camera A is the target image data. This allows for the accurate acquisition of images that meet the user's needs, providing accurate data for subsequent analysis.

[0075] When the extracted results do not contain location information, considering that users may need to perform a comprehensive analysis of the entire scene, or that subsequent processing can extract useful information from all image data, the system uses all candidate image data as target image data. For example, when a mobile robot is performing a warehouse inventory task, the operator issues the instruction "Check the goods in the warehouse" without specifying the location. The system uses all candidate image data collected by the robot's various cameras as target image data to comprehensively understand the distribution and quantity of goods in the warehouse.

[0076] Beneficial effects: Precisely meeting specific needs: When the command includes location information, by accurately matching the target location and target viewpoint, it can precisely filter target image data that meets the user's specific location requirements, improving the accuracy and efficiency of image processing. This satisfies the precise location information needs of intelligent driving, mobile robot navigation, and autonomous drone flight in complex scenarios. For example, intelligent driving vehicles can quickly acquire traffic condition images of the required location and make accurate driving decisions.

[0077] Fully adaptable to fuzzy commands: When commands do not contain location information, all candidate image data are included in the processing scope, ensuring that the system can fully adapt to fuzzy command situations and will not miss important information due to insufficient information. This enhances the system's robustness and adaptability in different task scenarios. For example, when a UAV executes a fuzzy command task, it can flexibly adjust the task execution method by analyzing all image data.

[0078] In intelligent driving scenarios, when a self-driving car enters a complex intersection, the driver issues the command "Check the traffic light status 30 degrees to the left front." The system extracts the directional information of "30 degrees to the left front" to determine the target location. Based on the correspondence between the vehicle's camera viewpoint and the location, it finds the image captured by the camera at the corresponding 30-degree left front viewpoint as the target image data. By analyzing this target image data, the system accurately obtains the traffic light status, providing a basis for intelligent driving decisions and ensuring the vehicle safely passes through the intersection.

[0079] In a mobile robot navigation scenario, a mobile robot performs cleaning tasks in a large shopping mall. The operator issues the instruction "Check the surroundings for litter," without specifying the location. The system uses candidate image data collected by the robot's various cameras as target image data. By analyzing this image data, the robot can comprehensively scan the mall environment, detect litter in every corner, and efficiently complete the cleaning task.

[0080] For autonomous drone flight, assuming the drone is conducting a power line inspection mission, the operator issues the command "Check the situation near a certain section of the line," without specifying the location. The drone's camera collects candidate image data from multiple perspectives, and the system uses all of these image data as the target image data. By analyzing this image data, the drone can comprehensively inspect the growth of trees, the distance to buildings, and other conditions around the power line, promptly identifying potential safety hazards.

[0081] Optionally, the step of analyzing and processing the 3D representation corresponding to the candidate image data and the target image data based on the instruction information to obtain the inference result corresponding to the instruction information includes: The instruction information, the candidate image data, and the target image data are input into a pre-trained first model to extract the three-dimensional representation corresponding to the candidate image data, and inference is performed based on the instruction information, the three-dimensional representation corresponding to the candidate image data, and the target image data to obtain the inference result corresponding to the instruction information; wherein, the first model includes: a three-dimensional feature extraction module, a visual encoding module, a text encoding module, and a feature analysis and processing module.

[0082] In this embodiment, in the fields of intelligent driving, mobile robot navigation, and autonomous drone flight, in-depth analysis and processing of the 3D representations corresponding to candidate image data and target image data based on command information is a key step in achieving accurate inference results and realizing intelligent decision-making. However, traditional methods often have many shortcomings in this process. On the one hand, the lack of an integrated and efficient model to systematically process command information, image data, and their 3D representations leads to relatively independent processing steps, failing to fully explore the complex relationships between them, thus significantly reducing the accuracy and reliability of the inference results. On the other hand, the algorithms used in traditional methods for feature extraction and inference are relatively simple, making it difficult to fully utilize the rich information in images and commands, and failing to meet the stringent requirements of high-precision decision-making in these fields. Therefore, there is an urgent need for a method based on advanced models that can comprehensively and deeply perform feature extraction and inference.

[0083] The first model is a comprehensive intelligent model pre-trained on a large amount of data, specifically designed for scenarios such as intelligent driving, mobile robot navigation, and autonomous flight of unmanned aerial vehicles. It integrates multiple powerful modules, including a 3D feature extraction module, a visual encoding module, a text encoding module, and a feature analysis and processing module, which are used to systematically analyze and process command information, candidate image data, and their 3D representations to obtain accurate and reliable inference results.

[0084] The 3D feature extraction module is a crucial component of the first model. Utilizing advanced 3D modeling and analysis techniques, it extracts key 3D feature information from the 3D representations corresponding to candidate image data. These features accurately reflect important information such as the spatial structure, positional relationships, and motion trends of objects in the scene, providing rich 3D spatial evidence for subsequent reasoning and helping the system better understand the spatial layout of the environment.

[0085] The visual encoding module focuses on feature extraction from target image data. It utilizes deep learning and computer vision techniques to perform multi-level and multi-angle analysis of the target image data, transforming it into a feature representation that is easily understood and processed by computers. The extracted first image feature data contains rich visual information from the image, such as the shape, color, texture, and edges of objects, providing intuitive image evidence for subsequent reasoning based on other features.

[0086] The text encoding module is primarily responsible for feature extraction from instruction information. Through natural language processing (NLP) technology, it performs in-depth analysis of the instruction information, transforming the textual instruction information into semantic feature data. This semantic feature data accurately reflects the intent, objectives, and various constraints of the instruction, providing clear guidance for the entire analysis and processing process and ensuring that the system accurately processes the instructions according to their requirements.

[0087] Feature Analysis and Processing Module: As the core of the first model, it is responsible for fusing the feature data extracted by the 3D feature extraction module, visual encoding module, and text encoding module, and deriving the inference result through complex decoding and reasoning operations. It can uncover the inherent relationships between different feature data, comprehensively consider instruction semantics, image visual information, and scene 3D spatial information, thereby making accurate decisions and determining the quality and reliability of the final inference result.

[0088] The system inputs instruction information, candidate image data, and target image data into the pre-trained first model. First, the text encoding module begins its work, analyzing the instruction information word by word and sentence by sentence. Using natural language processing techniques such as word vector representation and syntactic analysis, it extracts key semantic elements from the instruction, such as actions, objects, and conditions, and encodes these semantic elements into semantic feature data. For example, in an autonomous driving scenario, for the instruction "identify the red traffic light on the road ahead and determine its status," the text encoding module extracts key semantic information such as "road ahead," "red," "traffic light," and "determine status," and converts them into corresponding semantic feature vectors. These vectors accurately represent the semantic meaning of the instruction.

[0089] Next, the visual encoding module processes the target image data. It employs deep learning architectures such as convolutional neural networks, starting with pixel-level information of the target image and gradually extracting low-level features (such as edges and textures) and high-level features (such as object categories and shapes) to form the first image feature data. Taking the road image in autonomous driving as an example, the visual encoding module can extract visual features such as the shape of the road, the position and color of traffic lights, and the outlines of vehicles, providing intuitive image information for subsequent reasoning.

[0090] Simultaneously, the 3D feature extraction module extracts features from the 3D representation corresponding to the candidate image data. Based on techniques such as 3D geometric analysis and spatial transformation, it extracts information such as the spatial coordinates, volume, shape parameters, and relative spatial relationships between objects from the 3D representation, obtaining reference 3D feature data. For example, in a mobile robot navigation scenario, the 3D feature extraction module can extract the position and orientation information of shelves, goods, and the robot itself in 3D space from the 3D representation of the warehouse environment, providing spatial reference for path planning and task execution.

[0091] Finally, the feature analysis and processing module fuses the reference 3D feature data, semantic feature data, and first image feature data. It uses specific fusion algorithms, such as feature stitching and attention mechanisms, to uncover complementary information and intrinsic relationships between different feature data. Then, after decoding, the fused feature data is transformed into the final inference result. For example, in an autonomous drone flight scenario, by combining command semantics (such as finding a specific landmark and planning a route to it), visual information in the target image (landmark appearance features), and scene 3D spatial information (the spatial relationship between the landmark and the drone), the feature analysis and processing module can plan an optimal flight route to avoid obstacles as the inference result.

[0092] Beneficial effects Precise Decision Support: The first model integrates multiple functional modules to achieve deep fusion and system processing of command information, image data, and their 3D representations. These modules work collaboratively, fully exploring their potential connections to provide high-precision inference results for intelligent driving, mobile robot navigation, and autonomous drone flight, thus supporting precise decision-making. In intelligent driving, it accurately identifies traffic light status and makes corresponding decisions, ensuring driving safety and efficiency; in mobile robot navigation, it precisely plans paths to complete complex tasks; and in autonomous drone flight, it rationally plans flight routes to ensure the smooth execution of flight missions.

[0093] Efficient Information Utilization: The model can simultaneously process multiple types of information from instructions, images, and 3D representations, fully leveraging the advantages of different information sources. Whether it's semantic information in instructions, visual information from images, or 3D spatial information from scenes, all can be effectively utilized in the model, avoiding information omissions and waste. This not only improves information processing efficiency but also enhances the system's adaptability to complex scenes and diverse instructions, maintaining good performance under different environments and task requirements.

[0094] In an optional embodiment, in an intelligent driving scenario, the vehicle is driving on a city street, and the driver issues the instruction "plan a safe acceleration path when encountering a green light at the intersection ahead." The system inputs this instruction information, candidate image data collected by each camera of the vehicle, and target image data of the intersection ahead into the first model. The text encoding module extracts semantic feature data such as "intersection ahead," "green light," and "plan acceleration path" from the instruction. The visual encoding module performs feature extraction on the target image data to obtain first image feature data containing information such as the color of the traffic lights at the intersection, vehicle distribution, and road conditions. The three-dimensional feature extraction module extracts features from the three-dimensional representation corresponding to the candidate image data (including three-dimensional spatial information of buildings, vehicles, etc. around the intersection) to obtain reference three-dimensional feature data reflecting the three-dimensional spatial layout of the intersection.

[0095] The feature analysis and processing module integrates these three types of feature data. By analyzing information such as traffic light status, vehicle position and speed in three-dimensional space, and road geometry, and combining it with instruction requirements, it plans a safe acceleration path as the reasoning result. For example, "Currently green light, it is recommended to accelerate evenly to 30 km / h within three seconds and keep driving in the right lane," providing drivers with accurate driving guidance.

[0096] Optionally, the step of extracting the three-dimensional representation corresponding to the candidate image data and performing inference based on the instruction information, the three-dimensional representation corresponding to the candidate image data, and the target image data to obtain the inference result corresponding to the instruction information includes: The instruction information is input into the text encoding module for feature extraction to obtain semantic feature data corresponding to the instruction information; The target image data is input into the visual encoding module for feature extraction to obtain the first image feature data corresponding to the target image data; The three-dimensional representation corresponding to the candidate image data is input into the three-dimensional feature extraction module for feature extraction to obtain reference three-dimensional feature data; The reference 3D feature data, the semantic feature data, and the first image feature data are input into the feature analysis and processing module for fusion and decoding to obtain the inference result.

[0097] In this embodiment, extracting key features from candidate image data, their 3D representations, and command information, and then using these features to make accurate inferences to arrive at valid results, is a crucial step in application scenarios such as intelligent driving, mobile robot navigation, and autonomous drone flight. However, traditional methods have many shortcomings in this process. In terms of feature extraction, they often fail to precisely extract the most representative features from different data sources, leading to information omissions or redundancy. For example, in intelligent driving scenarios, it may be impossible to accurately extract the spatial position and velocity features of target objects that are closely related to vehicle driving safety; in mobile robot navigation, the influence of the shape and texture features of target objects on their recognition and grasping may be overlooked; and in autonomous drone flight, it may be impossible to fully extract the 3D features of the terrain to plan a safe flight route.

[0098] In the reasoning process, traditional methods may not adequately integrate these features, resulting in less rigorous reasoning logic and an inability to fully utilize the inherent relationships between features. This leads to inaccurate inference results, failing to meet the high-precision decision-making requirements of practical applications. Furthermore, traditional methods lack standardized and modular processing flows, making the entire process complex and difficult to optimize, increasing development and maintenance costs. Therefore, a method with refined feature extraction and a reasonable reasoning process is needed to improve the quality of inference results and the operability of the processing.

[0099] The system first inputs the instruction information into the text encoding module for feature extraction. The text encoding module employs advanced natural language processing technology to perform in-depth semantic understanding of the instruction information. It represents each word in the instruction as a word vector, captures the semantic relationships between words, and identifies key semantic elements in the instruction, such as actions, objects, and conditions, through syntactic analysis and semantic role labeling. Then, these semantic elements are converted into corresponding semantic feature data, represented in vector form. These vectors accurately represent the semantic connotation of the instruction. For example, for the instruction "When encountering a red obstacle ahead, the drone lowers its altitude and changes its flight direction," the text encoding module extracts semantic information such as "ahead," "red," "obstacle," "lower altitude," and "change flight direction," encoding them into corresponding semantic feature vectors. These vectors not only contain the key information of the instruction but also reflect the logical relationships between them.

[0100] Next, the target image data is input into the visual encoding module. The visual encoding module utilizes the multi-layered structure of a convolutional neural network to analyze the target image data starting from the pixel-level information. The lower convolutional layers are responsible for extracting low-level features of the image, such as edges and textures; as the network deepens, higher convolutional layers gradually extract higher-level features, such as the shape and category of objects. Through this multi-layered feature extraction method, the visual encoding module can comprehensively capture the visual information in the target image, forming the first image feature data. For example, in an intelligent driving scenario, for a target image of the road ahead, the visual encoding module can extract features such as the shape, color, and position of objects like vehicles, pedestrians, and traffic signs, representing them as tensors as the first image feature data, providing rich visual evidence for subsequent reasoning.

[0101] Simultaneously, the 3D representations corresponding to the candidate image data are input into the 3D feature extraction module. Based on 3D geometric analysis, spatial transformation, and deep learning techniques, the 3D feature extraction module extracts key information such as the spatial location, shape, size, and spatial relationships of objects from the 3D representation, obtaining reference 3D feature data. For example, in a mobile robot navigation scenario, for a 3D representation of a warehouse environment, the 3D feature extraction module can extract information such as the 3D coordinates, volume, shape parameters, and relative positional relationships of shelves and goods, representing them in matrix form as reference 3D feature data, providing important spatial references for robot path planning and operation.

[0102] Finally, the reference 3D feature data, semantic feature data, and first image feature data are input into the feature analysis and processing module. This module first merges the three feature data along a specific dimension using feature concatenation to form a comprehensive feature vector. Then, an attention mechanism is used to process the comprehensive feature vector, calculating the correlation weights between different features, enabling the model to focus more on feature information relevant to the current task. Next, a neural network structure such as a multilayer perceptron (MLP) is used to decode the weighted feature vector, mapping it to the final inference result space. For example, in an intelligent driving scenario, combining the semantic features of the instruction (such as the operational requirements when encountering an obstacle), the visual features of the target image (the appearance and location of the obstacle), and the 3D spatial features of the scene (the spatial relationship between the obstacle and the vehicle), the feature analysis and processing module fuses and decodes the data to derive the specific actions the vehicle should take, such as "avoid to the left and decelerate to 20 km / h" as the inference result.

[0103] Beneficial effects: High-precision feature extraction: By employing advanced natural language processing, computer vision, and 3D analysis techniques, key features are meticulously extracted from the 3D representations of instruction information, target image data, and candidate image data, maximizing the extraction of useful information from the data and reducing information omissions and redundancy. This enables the system to more accurately understand instruction intent and perceive the surrounding environment, providing a solid foundation for subsequent reasoning and thus improving the accuracy and reliability of the reasoning results.

[0104] Rational Reasoning and Decision Making: The feature analysis and processing module, through innovative feature fusion and reasoning mechanisms, fully leverages the inherent connections between different features to perform rigorous reasoning. The application of attention mechanisms enables the model to adaptively focus on important features, improving the relevance and effectiveness of reasoning. Decoding structures such as multilayer perceptrons accurately map comprehensive features to the reasoning result space, providing rational decision-making basis for intelligent driving, mobile robot navigation, and autonomous drone flight, meeting the stringent requirements for high-precision decision-making in practical applications.

[0105] Standardization and Operability: The clear modular processing flow ensures good standardization and operability throughout the feature extraction and inference process. Each module is responsible for specific types of data processing and feature extraction, with clearly defined responsibilities, facilitating development, maintenance, and optimization. This modular design also facilitates model expansion and upgrades, enabling easy integration of new technologies and algorithms to adapt to ever-changing application needs and data characteristics.

[0106] In an optional embodiment, in an intelligent driving scenario, the vehicle is traveling on a highway, and the driver issues the instruction "When a slow-moving large vehicle is detected ahead, maintain a safe distance and prepare to overtake." The system inputs this instruction information into a text encoding module, which extracts semantic feature data such as "ahead," "slow-moving," "large vehicle," "maintain a safe distance," and "prepare to overtake." The target image data of the road ahead is input into a visual encoding module, which extracts first image feature data such as the vehicle's shape, speed, and distance. A three-dimensional representation of the vehicle's surrounding environment is input into a three-dimensional feature extraction module, which extracts reference three-dimensional feature data such as the spatial position and size of the vehicle, road, and surrounding objects.

[0107] The feature analysis and processing module concatenates these three types of feature data, and then uses an attention mechanism to calculate the weights of different features, paying more attention to feature information related to large vehicles. After decoding by a multilayer perceptron, the inference result is "There is a large truck traveling at 60 km / h 200 meters ahead. Maintain a safe distance of 50 meters and wait for a suitable opportunity to overtake." This result is fed back to the driver to provide accurate support for driving decisions.

[0108] In a mobile robot navigation scenario, the robot performs a part-grabbing task in a factory workshop. Upon receiving the instruction "Find the yellow cylindrical part and move directly above it to prepare for grabbing," the system inputs the instruction information into a text encoding module, extracting semantic feature data such as "yellow," "cylindrical," "part," and "move directly above to prepare for grabbing." The system then inputs target image data that may contain the part into a visual encoding module, which extracts primary image feature data such as the part's color, shape, and position. Finally, a 3D representation of the workshop environment is input into a 3D feature extraction module, which extracts reference 3D feature data such as the spatial position and orientation of the part, machinery, and the robot itself.

[0109] The feature analysis and processing module fuses three types of feature data and highlights features related to the yellow cylindrical part through an attention mechanism. After decoding by a multilayer perceptron, it plans the path for the robot to move directly above the part as the inference result, guiding the robot to execute the task accurately.

[0110] In an autonomous drone flight scenario, the drone performs a patrol mission in an urban area and receives an instruction to "find and photograph landmark buildings in the city, avoiding tall buildings and power lines." The system inputs the instruction information into a text encoding module to extract semantic feature data such as "landmark buildings," "find and photograph," and "avoid tall buildings and power lines." The target image data of the urban area is input into a visual encoding module, which extracts first-order image feature data such as the shape and height of buildings and the location of power lines. Finally, a 3D representation of the urban area is input into a 3D feature extraction module, which extracts reference 3D feature data such as the position and height of buildings and power lines in 3D space.

[0111] The feature analysis and processing module integrates these feature data and uses an attention mechanism to focus on features related to landmark buildings and obstacles. After decoding by a multilayer perceptron, a flight path that avoids tall buildings and power lines and can capture images of landmark buildings is planned as the inference result, enabling the drone to successfully complete the patrol mission.

[0112] Optionally, the feature analysis and processing module includes: a decoding submodule, a cross-view 3D geometry enhancement submodule, and a prediction submodule; the step of inputting the reference 3D feature data, the semantic feature data, and the first image feature data into the feature analysis and processing module for fusion and decoding to obtain the inference result includes: The semantic feature data and the first image feature data are concatenated and input into the decoding layer in the decoding submodule for decoding. The feature data output from each decoding layer in the decoding submodule and the reference 3D feature data are input into the cross-view 3D geometry enhancement submodule for fusion to obtain fused 3D feature data. The fused 3D feature data is input into the decoding submodule or the prediction module for processing to obtain the inference result.

[0113] In this embodiment, in complex application scenarios such as intelligent driving, mobile robot navigation, and autonomous drone flight, the feature analysis and processing module, as the core component for obtaining accurate inference results, is crucial in terms of the rationality and efficiency of its internal structure and processing flow. Traditional feature analysis and processing methods often lack refined structural design and effective information fusion strategies. For example, when processing multiple feature data, they may simply splice or linearly combine them, failing to fully explore the complex relationships and complementary information between different features, thus limiting the accuracy of the inference results. Furthermore, traditional methods may employ a relatively singular approach in feature decoding and prediction, unable to flexibly adapt to the needs of different scenarios and tasks, resulting in weak system generalization ability. Therefore, a feature analysis and processing module with optimized internal structure and efficient processing flow is needed to improve the system's performance in different application scenarios.

[0114] The decoding submodule is a crucial component of the feature analysis and processing module. It is responsible for decoding the input feature data, transforming it into a more understandable and processable form. It comprises multiple cascaded decoding layers, each with a specific function. Through progressive decoding, it extracts task-relevant information from the feature data, providing a foundation for subsequent fusion and prediction.

[0115] Cross-view 3D geometry enhancement submodule: This submodule focuses on enhancing the fusion effect of 3D geometric features from different viewpoints. Through specific algorithms and mechanisms, it deeply fuses the feature data output by the decoding submodule with reference 3D feature data, fully exploring the intrinsic connections between 3D spatial information and other features. This results in more representative fused 3D feature data, improving the accuracy and reliability of inference results.

[0116] The prediction submodule performs the final classification or prediction operation on the fused feature data. Based on the nature of the task (such as target recognition, path planning, state determination, etc.), it maps the fused 3D feature data to the corresponding result space, deriving the final inference result and providing direct basis for the agent's decision-making in different application scenarios.

[0117] The system concatenates semantic feature data and first image feature data. This step is like integrating information from different domains (instruction semantics and image vision) to form a more comprehensive feature set. For example, in an intelligent driving scenario, semantic feature data contains the description of traffic conditions in the instruction, while the first image feature data contains visual information about vehicles and traffic signs in the image of the road ahead. Concatenating them together provides a richer source of information for subsequent decoding.

[0118] Next, the concatenated feature data is input into the initial decoding layer of the decoding submodule for decoding. As the entry point to the decoding submodule, the initial decoding layer employs a specific neural network structure (such as a deconvolutional layer or a recurrent neural network layer) to perform preliminary decoding on the input feature data, transforming it from a high-dimensional abstract feature space into first-dimensional feature data with certain semantics and structure. For example, in a mobile robot navigation scenario, the initial decoding layer might decode the concatenated feature data into a two-dimensional feature map containing information about the target object's position and shape, providing a foundation for further processing.

[0119] The decoding submodule contains multiple cascaded decoding layers, including a starting decoding layer, at least one intermediate decoding layer, and a final decoding layer. The intermediate decoding layers further refine the features and extract semantics from the first two-dimensional feature data obtained from the previous layer. Each intermediate decoding layer, based on the output of the previous layer, uses specific convolutional operations or other feature transformations to mine deeper information from the feature data, continuously enriching and improving the feature representation. For example, in an autonomous drone flight scenario, an intermediate decoding layer might further extract key features related to flight path planning, such as traversable areas and obstacle boundaries, from the first two-dimensional feature data containing terrain and landmark information.

[0120] The final decoding layer performs the final decoding operation on the feature data processed by the intermediate decoding layers to obtain a final two-dimensional feature representation suitable for input into the prediction submodule. It integrates and adjusts the feature data to better meet the input requirements of the prediction submodule, while retaining information crucial to the final inference result.

[0121] During the decoding process in the decoding submodule, the first two-dimensional feature data output from each decoding layer and the reference three-dimensional feature data are input into the cross-view three-dimensional geometry enhancement submodule. This submodule first expands the reference three-dimensional feature data to obtain the second two-dimensional feature data, making it easier to fuse with the first two-dimensional feature data in terms of dimension. Then, the first two-dimensional feature data and the second two-dimensional feature data are subjected to dimensionality compression, for example, through fully connected layers or pooling operations, to reduce the data dimensionality, reduce the amount of computation, and retain key information.

[0122] Next, attention is calculated based on the compressed first and second two-dimensional feature data. The attention mechanism can dynamically calculate the importance weights between different feature elements to form an attention matrix. For example, in intelligent driving scenarios, the attention mechanism can highlight the weights of three-dimensional spatial features (such as the distance and speed of obstacles ahead) relevant to the current decision in feature fusion, based on the vehicle's driving status and command requirements.

[0123] The compressed first and second two-dimensional feature data are weighted and fused using an attention matrix, enhancing features more relevant to the current task. Then, a dimensionality restoration operation is performed to restore the fused data to appropriate dimensions, resulting in fused three-dimensional feature data. This fused three-dimensional feature data fully combines the advantages of two-dimensional image features and three-dimensional spatial features, providing a more powerful feature representation for subsequent predictions.

[0124] Finally, the fused 3D feature data is input into either the decoding or prediction submodule for processing. If input into the decoding submodule, it is further fused with the first 2D feature data output from the decoding layer and then input into the next adjacent decoding layer for decoding. Through this repeated fusion and decoding operation, the feature representation is continuously optimized, and deeper connections between features are uncovered. When the fused 3D feature data is input to the final decoding layer, the first 2D feature data output from the final decoding layer is fused again with the corresponding fused 3D feature data from the final decoding layer and then input into the prediction submodule for classification. The prediction submodule analyzes and predicts the fused feature data according to specific task requirements, such as determining the type of traffic signs in autonomous driving or determining the location of target objects in mobile robot navigation, to obtain the final inference result.

[0125] Beneficial effects Deep Feature Fusion: Through the collaborative work of the decoding submodule, the cross-view 3D geometry enhancement submodule, and the prediction submodule, deep fusion of semantic features, image features, and 3D geometric features is achieved. This fusion method fully leverages the complementary information and complex relationships between different features, enabling the model to more comprehensively and accurately understand scene information, thereby improving the accuracy and reliability of inference results. In scenarios with extremely high requirements for decision-making accuracy, such as intelligent driving, mobile robot navigation, and autonomous drone flight, it can make decisions that are more in line with the actual situation.

[0126] Flexible Feature Processing: The cascaded design of multiple decoding layers in the decoding submodule, along with the iterative fusion and decoding of features between the decoding layers and the cross-view 3D geometry enhancement submodule, enables the model to flexibly process feature data. This flexibility helps the model adapt to the needs of different scenarios and tasks, improving its generalization ability through continuous optimization of feature representation. Whether in complex urban traffic scenarios, diverse indoor and outdoor working environments, or flight areas with different terrains, the model can effectively process feature data and derive accurate inference results.

[0127] Efficient information utilization: Operations such as dimensionality compression, attention calculation, and weighted fusion in the cross-view 3D geometry enhancement submodule effectively utilize information from feature data, reduce redundancy, and improve computational efficiency. Simultaneously, through reasonable module design and data flow, the entire feature analysis and processing process can operate efficiently, meeting the real-time requirements of scenarios such as intelligent driving, mobile robot navigation, and autonomous drone flight while ensuring inference accuracy.

[0128] In an optional embodiment, in an intelligent driving scenario, when a vehicle is driving at an intersection, the driver issues a command to "identify the color of the traffic light ahead and determine whether it is safe to pass." The system extracts semantic feature data from the command information via a text encoding module, extracts first image feature data from the target image data of the intersection ahead via a visual encoding module, and extracts reference three-dimensional feature data from the three-dimensional representation of the vehicle's surrounding environment via a three-dimensional feature extraction module.

[0129] The semantic feature data and the first image feature data are concatenated and input into the initial decoding layer of the decoding submodule. The initial decoding layer outputs the first two-dimensional feature data containing the location of the traffic light and preliminary visual features. This data, along with the unfolded and processed reference three-dimensional feature data, enters the cross-view three-dimensional geometry enhancement submodule. Through operations such as dimensionality compression, attention calculation, and weighted fusion, fused three-dimensional feature data is obtained.

[0130] The fused 3D feature data is fed back to the intermediate decoding layer of the decoding submodule. It is then fused with the first 2D feature data output from the intermediate decoding layer and further decoded to extract features related to traffic light color and safe passage judgment. After processing by the final decoding layer, the first 2D feature data output from the final decoding layer is fused again with the corresponding fused 3D feature data and input into the prediction submodule. Based on information such as traffic light color and the spatial relationship between vehicles and intersections, the prediction submodule determines the inference result of "current green light, safe passage," providing a decision-making basis for the driver.

[0131] In a mobile robot navigation scenario, the robot performs a task in a warehouse and receives the instruction "find the red cube-shaped goods and move next to it." The system acquires the semantic feature data of the instruction, the first image feature data of the target area within the warehouse, and the reference 3D feature data of the warehouse environment.

[0132] After concatenating semantic features and first image features, the first two-dimensional feature data is obtained by decoding at the starting decoding layer of the decoding submodule. This data is then fused with the processed reference three-dimensional feature data in the cross-view three-dimensional geometry enhancement submodule to form fused three-dimensional feature data. The fused three-dimensional feature data is further decoded at the intermediate decoding layer of the decoding submodule, where it is fused with the first two-dimensional feature data output from that layer to refine the position and shape features of the goods. Finally, the first two-dimensional feature data output from the final decoding layer is fused with the corresponding fused three-dimensional feature data and input into the prediction submodule. The prediction submodule derives inference results such as the direction and distance the robot should move, guiding the robot to accurately reach the red cube-shaped goods. Optionally, concatenating the semantic feature data and the first image feature data and inputting them into the decoding layer of the decoding submodule for decoding includes: The semantic feature data and the first image feature data are concatenated and input into the starting decoding layer of the decoding module for decoding to obtain the first two-dimensional feature data; wherein, the decoding submodule includes multiple decoded layers in series, and the decoded layers include a starting decoding layer, an ending decoding layer and at least one intermediate decoding layer.

[0133] In this embodiment, in applications such as intelligent driving, mobile robot navigation, and autonomous drone flight, the operation of the decoding submodule within the feature analysis and processing module plays a crucial role in obtaining accurate inference results. Traditional decoding methods may lack systematicity and specificity when processing semantic feature data and first image feature data, resulting in an inability to fully extract useful information from the data and affecting the accuracy of inference. For example, the structure and parameters of the decoding layer may not be reasonably designed, making the decoding process too simplistic and unable to effectively extract task-related information from the concatenated feature data. Alternatively, the differences and complementarities between different types of feature data may not be considered during the decoding process, leading to poor information fusion. Therefore, an optimized decoding method is needed that can precisely decode the concatenated feature data to provide high-quality feature representations for subsequent feature fusion and inference.

[0134] The system concatenates semantic feature data and first image feature data, a step that integrates crucial information from different data sources. Semantic feature data represents the intent and requirements of the instruction, while the first image feature data contains visual information about the target scene. For example, in an autonomous driving scenario, semantic feature data might indicate the need to identify a specific type of vehicle, while the first image feature data contains the visual features of various vehicles on the road. Concatenating them together forms a comprehensive feature vector, providing a rich information foundation for the decoding operation.

[0135] Then, the concatenated feature data is input into the initial decoding layer of the decoding module for decoding. The initial decoding layer employs a specific neural network structure, such as a deconvolution layer, whose purpose is to convert the high-dimensional concatenated feature data into first-dimensional feature data with certain spatial structure and semantic information. The deconvolution layer recovers the spatial information of the image by progressively expanding the size of the feature map, while transforming and combining the features so that the information implicit in the high-dimensional feature data can be presented in the form of a two-dimensional feature map. For example, in a mobile robot navigation scenario, the initial decoding layer may decode the concatenated feature data into a two-dimensional feature map, where different regions represent preliminary information such as the position and shape of the target object (such as cargo).

[0136] The decoding submodule comprises multiple cascaded decoding layers, including a starting decoding layer, a ending decoding layer, and at least one intermediate decoding layer. Each decoding layer has a unique function, processing the feature data sequentially to gradually refine and enrich the feature representation. The starting decoding layer, as the starting point of the entire decoding process, lays the foundation for subsequent decoding operations, initially extracting task-related feature information. For example, in an autonomous drone flight scenario, the starting decoding layer might extract features related to landmark locations and approximate shapes from the stitched feature data, outputting them as a two-dimensional feature map, thus providing a basis for subsequent intermediate decoding layers to further mine detailed information.

[0137] Beneficial effects Fine-grained feature extraction: By concatenating semantic feature data and first image feature data and then decoding them in the initial decoding layer, the complementary information of both can be fully utilized to finely extract task-related features. This approach avoids the information omission problem that may occur when processing feature data separately, improving the accuracy and completeness of feature extraction. In scenarios such as intelligent driving, mobile robot navigation, and autonomous drone flight, it can more accurately identify target objects, plan paths, or judge environmental states, providing a more reliable basis for subsequent decision-making.

[0138] Building the Decoding Foundation: The first two-dimensional feature data output by the initial decoding layer provides a solid foundation for the processing of subsequent intermediate and final decoding layers. It transforms complex high-dimensional feature data into a form that is easier for subsequent decoding layers to process, enabling the entire decoding process to proceed in an orderly manner. Each decoding layer can build upon the output of the previous layer to progressively extract deeper information from the feature data, thereby improving the efficiency and effectiveness of decoding and ultimately enhancing the quality of the inference results.

[0139] Adaptable to Complex Tasks: This decoding method can adapt to various complex task requirements in applications such as intelligent driving, mobile robot navigation, and autonomous drone flight. Whether it is a simple target recognition task or a complex path planning and environmental perception task, by reasonably designing the structure and parameters of the initial decoding layer and subsequent decoding layers, it is possible to extract feature information that meets the task requirements from the spliced ​​feature data, thereby enhancing the system's adaptability and flexibility to different tasks.

[0140] In an optional embodiment, in an intelligent driving scenario, the vehicle is traveling on a highway, and the driver issues the instruction "identify the police car appearing ahead and maintain a safe distance." The system uses a text encoding module to obtain semantic feature data from the instruction, which includes information such as "ahead," "police car," and "maintain a safe distance." The system also uses a visual encoding module to obtain first image feature data from the target image data of the road ahead, which includes visual information such as the vehicle's shape and color.

[0141] Semantic feature data and first image feature data are concatenated and input into the initial decoding layer. The initial decoding layer uses deconvolution to decode the concatenated high-dimensional feature data into first two-dimensional feature data, forming a two-dimensional feature map. In this feature map, specific regions may represent information such as the position, color, and approximate shape of vehicles ahead, initially displaying features related to police cars. For example, color features in the feature map can initially determine whether a vehicle is a police car (police cars usually have specific color markings), and position features can determine the specific location of the police car on the road ahead, providing a foundation for subsequent intermediate decoding layers to further accurately identify police cars and plan safe distances.

[0142] Optionally, the step of fusing the feature data output from each decoding layer in the decoding submodule with the reference 3D feature data into the cross-view 3D geometry enhancement submodule to obtain fused 3D feature data includes: For each of the decoding layers, the first two-dimensional feature data output by the decoding layer is input into the fusion with the reference three-dimensional feature data to obtain the fused three-dimensional feature data corresponding to the decoding layer.

[0143] In this embodiment, within the context of applications such as intelligent driving, mobile robot navigation, and autonomous drone flight, effectively fusing different types of feature data is a key issue in improving the accuracy of system decision-making. In the feature analysis and processing module, fusing the feature data output from each decoding layer in the decoding submodule with reference 3D feature data is crucial for fully utilizing the spatial information and other relevant information of the scene. However, traditional fusion methods are often too simplistic, potentially involving only direct addition or splicing, failing to delve into the intrinsic connections and complementarities between different feature data. This results in the fused feature data not accurately reflecting the true situation of the scene, thus affecting the accuracy of the inference results. Therefore, a more refined and efficient fusion method is needed to fully leverage the advantages of different feature data and improve the fusion effect.

[0144] For each decoding layer in the decoding submodule, the system fuses the first two-dimensional feature data output by the decoding layer with the reference three-dimensional feature data to obtain the fused three-dimensional feature data corresponding to the decoding layer. This process is divided into multiple steps to fully explore the potential of the two types of feature data.

[0145] First, the reference 3D feature data is unfolded to obtain the second 2D feature data. Since the reference 3D feature data is usually represented in the form of a 3D tensor, it needs to be unfolded into a 2D form to facilitate fusion with the first 2D feature data. For example, in an intelligent driving scenario, the reference 3D feature data may contain the 3D position and shape information of objects in the vehicle's surrounding environment. After unfolding, this information can be presented in the form of a 2D matrix, making it more dimensionally compatible with the first 2D feature data.

[0146] Next, dimensionality compression is performed on the first and second 2D feature data. This step reduces the dimensionality of the data and the computational load by using techniques such as fully connected layers or pooling operations, while retaining key information. Dimensionality compression not only improves computational efficiency but also removes some potentially redundant information, making subsequent fusion more effective. For example, in mobile robot navigation scenarios, pooling operations are used to compress the dimensionality of the first 2D feature data containing the target object's position and shape information, as well as the unfolded reference 3D feature data, retaining the most important feature information, such as the target object's key contours and approximate location.

[0147] Then, attention is calculated based on the compressed first and second two-dimensional feature data to obtain the attention matrix. The attention mechanism can dynamically calculate the importance weights between different feature elements, highlighting key features according to the needs of the task. In intelligent driving scenarios, the attention mechanism can determine which three-dimensional spatial features (such as the distance to obstacles ahead) and two-dimensional image features (such as the shape of obstacles) are more important to the current decision based on the vehicle's driving status and command requirements, thereby assigning corresponding weights to each feature element to form an attention matrix.

[0148] Finally, the compressed first and second two-dimensional feature data are weighted and fused based on the attention matrix, and then dimension restoration is performed to obtain fused three-dimensional feature data. Weighted fusion enhances features more relevant to the current task, thus better reflecting key information of the scene. Dimension restoration restores the fused data to appropriate dimensions for subsequent processing. For example, in an autonomous drone flight scenario, the fused three-dimensional feature data obtained after weighted fusion and dimension restoration can accurately represent the relationship between landmarks and surrounding obstacles in three-dimensional space, as well as their relevance to mission instructions, providing strong support for flight path planning.

[0149] Beneficial effects: Deep Feature Fusion: Through a series of operations including unfolding, dimensionality compression, attention computation, and weighted fusion, deep fusion of the first-dimensional feature data and the reference three-dimensional feature data is achieved. This fusion method can fully explore the inherent connections and complementarities between different feature data, making the fused feature data more comprehensive and accurate in reflecting scene information, thereby improving the accuracy and reliability of inference results. In scenarios with extremely high requirements for decision-making accuracy, such as intelligent driving, mobile robot navigation, and autonomous drone flight, it can make decisions that are more in line with the actual situation.

[0150] Efficient computation and information utilization: Dimensionality compression reduces data dimensionality, lowers computational load, and improves computational efficiency. Simultaneously, the attention mechanism enables the model to selectively utilize key information from feature data, avoiding redundancy and waste, further enhancing information utilization efficiency. This efficient computation and information utilization approach allows the system to process data rapidly in scenarios with high real-time requirements, meeting the decision-making needs of intelligent agents.

[0151] Enhanced Adaptability: The application of the attention mechanism enables the fusion process to dynamically adjust the fusion method of feature data according to different task requirements and scenario characteristics, highlighting features relevant to the current task. This enhances the system's adaptability to different application scenarios and tasks. Whether in complex urban traffic environments, diverse indoor and outdoor work sites, or flight areas with different terrains, the system can effectively fuse feature data and improve decision-making capabilities.

[0152] In an optional embodiment, in an intelligent driving scenario, the vehicle is driving on a city street, and the driver issues the instruction to "identify the traffic lights at the intersection ahead and determine whether it is safe to pass." The system obtains the first two-dimensional feature data at a certain decoding layer of the decoding submodule, which includes the position and preliminary visual features of the traffic lights. At the same time, it acquires reference three-dimensional feature data of the vehicle's surrounding environment, including the three-dimensional spatial layout of the intersection and the vehicle's spatial position.

[0153] After the reference 3D feature data is expanded into a second 2D feature data, it is compressed together with the first 2D feature data. Through pooling operations, the key visual features of the traffic lights and important information about the 3D spatial layout of the intersection are preserved. Next, attention is calculated. Based on factors such as the current vehicle speed and distance from the intersection, the importance weights of 3D spatial features (such as the spatial distance between the traffic lights and vehicles) and 2D image features (such as the color of the traffic lights) are determined to form an attention matrix.

[0154] The two types of compressed feature data are weighted and fused using an attention matrix to strengthen features relevant to determining whether it is safe to pass, such as the correlation between traffic light color and vehicle distance. Finally, dimensionality restoration is performed to obtain fused 3D feature data. This data accurately reflects the state of the traffic lights in 3D space and their relationship with vehicles, providing crucial information for subsequent judgments on safe passage.

[0155] Optionally, the step of inputting the fused 3D feature data into the decoding submodule or the prediction module for processing to obtain the inference result includes: The fused three-dimensional feature data corresponding to the decoding layer is fused with the first two-dimensional feature data output by the decoding layer, and then input into the next adjacent decoding layer for decoding. The first two-dimensional feature data output by the last decoding layer is fused with the fused three-dimensional feature data corresponding to the last decoding layer, and then input into the prediction submodule for classification to obtain the inference result.

[0156] In this embodiment, in the fields of intelligent driving, mobile robot navigation, and autonomous drone flight, accurately deriving inference results based on fused feature data is crucial for the system to achieve effective decision-making. Traditional methods often lack systematicity and rationality in processing feature data output from the decoding layer and fused 3D feature data, leading to inaccurate or unreliable inference results. For example, they may not fully utilize feature information from different stages of the decoding layer, or they may not perform appropriate processing when inputting the fused feature data into the prediction submodule, making the prediction results unable to accurately reflect the needs of the actual scenario. Therefore, a scientific and reasonable processing flow is needed that can fully integrate feature data from each decoding layer and accurately transform the fused feature data into the final inference result.

[0157] The system further fuses the fused 3D feature data corresponding to the decoding layer with the first 2D feature data output by the decoding layer, and then inputs this data into the next adjacent decoding layer for decoding. This step aims to further explore the potential connections between feature data and continuously optimize the feature representation through multiple fusion and decoding operations. For example, in an intelligent driving scenario, the fused 3D feature data of the current decoding layer contains the fusion result of the 3D spatial information of the vehicle's surrounding environment and image visual features, while the first 2D feature data retains the image semantic information at a specific stage during the decoding process. After fusing them again and inputting them into the next decoding layer, the next decoding layer can further extract information related to driving decisions based on this richer feature representation, such as the drivable area of ​​the road ahead and the location of potential hazards.

[0158] After processing through multiple decoding layers, the first two-dimensional feature data output by the final decoding layer is fused with the corresponding fused three-dimensional feature data from the final decoding layer. The output of the final decoding layer has already undergone multiple feature optimizations and refinements; merging it again with the fused three-dimensional feature data at this point integrates information from all the decoding layers, as well as the fusion information from the three-dimensional space and image features. For example, in a mobile robot navigation scenario, the first two-dimensional feature data output by the final decoding layer may already accurately represent the target object's position, shape, and grasping point, while the fused three-dimensional feature data supplements this with the target object's pose in three-dimensional space and its spatial relationship with the surrounding environment.

[0159] Finally, the fused feature data is input into the prediction submodule for classification to obtain inference results. The prediction submodule analyzes and judges the fused feature data based on the specific task type, such as target recognition and path planning in intelligent driving, target grasping decisions in mobile robot navigation, and flight attitude adjustment in autonomous drone flight. For example, in an autonomous drone flight scenario, the prediction submodule, based on the fused feature data, determines whether the drone is approaching the target area and whether it needs to adjust its flight altitude or direction, thus outputting corresponding inference results, such as "maintain current altitude, slightly adjust 5 degrees to the left".

[0160] Beneficial effects: Optimized feature representation: By fusing feature data from multiple decoding layers, complementary information between features of different levels and types can be fully explored, continuously optimizing feature representation and making the feature data more comprehensive and accurate in reflecting the actual scene. This helps improve the accuracy and reliability of inference results, providing stronger support for agents to make decisions in complex environments.

[0161] Improving decision-making accuracy: The feature data from the final decoding layer is fused again with the fused 3D feature data before being input into the prediction submodule. This ensures that the prediction submodule can obtain the most comprehensive and representative feature information. Based on this information, classification and reasoning enable the agent to make decisions that better meet actual needs, ensuring driving safety in autonomous driving, improving task execution efficiency in mobile robot navigation, and ensuring flight stability and mission completion quality in autonomous drone flight.

[0162] Enhanced system adaptability: This systematic processing flow is applicable to various application scenarios and task types, such as intelligent driving, mobile robot navigation, and autonomous drone flight. By flexibly adjusting the structure of the decoding layer and the algorithm of the prediction submodule, it can adapt to different environmental conditions and task requirements, thus enhancing the system's versatility and adaptability.

[0163] Optionally, the step of inputting the first two-dimensional feature data output by the decoding layer into the fused three-dimensional feature data corresponding to the decoding layer by fusing it with the reference three-dimensional feature data includes: The reference three-dimensional feature data is unfolded to obtain the second two-dimensional feature data; The first two-dimensional feature data and the second two-dimensional feature data are dimensionally compressed, and attention is calculated based on the compressed first two-dimensional feature data and the second two-dimensional feature data to obtain an attention matrix; The compressed first two-dimensional feature data and the second two-dimensional feature data are weighted and fused according to the attention matrix, and the dimensions are restored to obtain the fused three-dimensional feature data.

[0164] In this embodiment, in applications such as intelligent driving, mobile robot navigation, and autonomous drone flight, the process of fusing the first two-dimensional feature data output from the decoding layer with the reference three-dimensional feature data to obtain fused three-dimensional feature data is crucial. Traditional fusion methods are often too simplistic and direct, failing to fully exploit the potential of the two types of feature data, resulting in poor fusion effects. For example, simple concatenation or addition operations cannot effectively integrate feature information of different dimensions and properties, making the fused feature data unable to accurately reflect the real situation of the scene, thus affecting the accuracy and reliability of the inference results. Therefore, a more refined and efficient fusion method is needed that can fully utilize the advantages of the two types of feature data, improve the fusion effect, and provide a higher quality feature representation for subsequent inference.

[0165] The system first unfolds the reference 3D feature data to obtain the second 2D feature data. Since the reference 3D feature data is typically stored as a 3D tensor, its dimensions and structure do not match the first 2D feature data, hindering direct fusion. The unfolding operation transforms the 3D feature data into a 2D form, making it more dimensionally compatible with the first 2D feature data. For example, in intelligent driving scenarios, the reference 3D feature data may contain information such as the 3D coordinates and shapes of objects in the vehicle's surrounding environment. After unfolding, this information can be presented as a 2D matrix, facilitating fusion with the first 2D feature data.

[0166] Next, dimensionality compression is performed on the first and second 2D feature data. This step reduces the dimensionality of the data by using techniques such as fully connected layers or pooling operations. Dimensionality compression has two important functions: first, it reduces computational load and improves the efficiency of the fusion process; second, it removes redundant information from the data and retains key features. For example, in mobile robot navigation scenarios, pooling operations are used to compress the dimensionality of the first 2D feature data containing the target object's position and shape information, as well as the unfolded reference 3D feature data, retaining only the key outline and approximate position of the target object, making subsequent fusion more efficient and targeted.

[0167] Then, attention is calculated based on the compressed first and second two-dimensional feature data to obtain the attention matrix. The attention mechanism is a method that dynamically allocates weights, automatically determining the importance of different feature elements according to the needs of the current task. In intelligent driving scenarios, the attention mechanism can assign corresponding weights to each feature element based on the vehicle's driving state and command requirements. For example, in overtaking scenarios, it may focus more on features such as the speed and distance of vehicles ahead, forming an attention matrix. This matrix reflects the relative importance of different features in the current task.

[0168] Finally, the compressed first and second two-dimensional feature data are weighted and fused based on the attention matrix, and then dimension restoration is performed to obtain fused three-dimensional feature data. The weighted fusion operation applies the attention matrix to both types of feature data, assigning greater weight to features with higher importance, thereby highlighting key feature information. For example, in an autonomous drone flight scenario, if the current task is to avoid obstacles, features related to the obstacle's position and shape will be given higher weights. After weighted fusion, a dimension restoration operation is performed to restore the fused data to a dimension suitable for subsequent processing, resulting in fused three-dimensional feature data. This data format fully integrates the advantages of the first two-dimensional feature data and the reference three-dimensional feature data, providing richer and more accurate feature information for subsequent inference.

[0169] Beneficial effects: Deep Feature Integration: Through a series of operations including unfolding, dimensionality compression, attention computation, and weighted fusion, deep integration of the first-dimensional feature data and the reference three-dimensional feature data is achieved. This fusion method can fully explore the complementarity and potential connections between the two types of feature data, making the fused feature data more comprehensive and accurate in reflecting the actual situation of the scene, thereby improving the accuracy and reliability of the inference results. In scenarios with extremely high requirements for decision-making accuracy, such as intelligent driving, mobile robot navigation, and autonomous drone flight, it can provide more reliable decision-making basis for intelligent agents.

[0170] Efficient computation and information utilization: Dimensional compression reduces computational load while improving information utilization efficiency, removing redundant information and enabling the model to focus on key features. The attention mechanism further optimizes the feature fusion process, dynamically adjusting feature weights according to task requirements to ensure that important features are fully utilized. This efficient computation and information utilization approach allows the system to process data rapidly in scenarios with high real-time requirements, meeting the decision-making needs of intelligent agents.

[0171] Adaptive Task Requirements: The application of the attention mechanism enables the fusion process to automatically adjust the feature fusion method and weight allocation according to different task requirements and scenario characteristics. Whether it is intelligent driving in complex urban traffic environments, or mobile robot navigation or drone autonomous flight in diverse indoor and outdoor environments, the system can adaptively highlight features relevant to the current task, enhancing the system's adaptability and flexibility to different application scenarios.

[0172] Optionally, the step of obtaining the three-dimensional representation corresponding to the candidate image data includes: Perform 3D modeling based on the candidate image data to obtain an initial 3D representation; Feature extraction is performed on the camera parameters corresponding to the candidate image data to obtain parameter features; The parameter features are fused with the initial three-dimensional representation to obtain the fused three-dimensional representation.

[0173] In this embodiment, in the fields of intelligent driving, mobile robot navigation, and autonomous drone flight, obtaining accurate 3D representations of candidate image data is a crucial foundation for understanding the surrounding environment and making correct decisions. Traditional methods often suffer from insufficient accuracy or inability to fully utilize image data information when generating 3D representations. For example, in intelligent driving, it may be impossible to accurately construct a 3D model of the vehicle's surrounding environment, leading to inaccurate judgments of obstacle positions and distances; in mobile robot navigation, it cannot effectively combine image data from the robot's perspective to generate accurate 3D scene representations, affecting the accuracy of path planning; in autonomous drone flight, the generated 3D representations cannot accurately reflect the spatial relationships between terrain features and target objects, causing deviations in flight mission execution. Therefore, a method is needed that can more accurately and comprehensively obtain 3D representations of candidate image data.

[0174] The system first performs 3D modeling based on candidate image data to obtain an initial 3D representation. In intelligent driving scenarios, this may involve using multi-view geometry algorithms, combining candidate image data captured from different perspectives by multiple cameras on the vehicle, and constructing a 3D point cloud model of the vehicle's surrounding environment as the initial 3D representation through techniques such as feature matching and triangulation. This point cloud model initially depicts the spatial distribution of objects in the scene, but may contain some noise and incompleteness.

[0175] Next, feature extraction is performed on the camera parameters corresponding to the candidate image data to obtain parametric features. Camera parameters contain rich information, such as focal length, optical center position, rotation, and translation parameters. These parameters determine how the image is projected from the 3D scene onto the 2D plane. By extracting features from the camera parameters, parametric features reflecting key information such as camera viewpoint, shooting position, and orientation can be obtained. For example, in mobile robot navigation scenarios, the camera parameters of the robot's vision sensor can help determine the robot's position and orientation relative to its surroundings. After feature extraction, these parameters can provide important basis for subsequent 3D representation optimization.

[0176] Finally, the parametric features are fused with the initial 3D representation to obtain the fused 3D representation. The fusion process can employ various methods, such as encoding the parametric features into vector form, concatenating them with the point cloud data or other geometric representations of the initial 3D representation, and then processing them through neural networks or other optimization algorithms. This allows the information from the parametric features to be integrated into the initial 3D representation, thereby optimizing the 3D model. In autonomous UAV flight scenarios, camera parametric features can correct information such as terrain features and the position and shape of target objects in the initial 3D representation, making the fused 3D representation more accurately reflect the actual scene and providing a more reliable foundation for UAV path planning and flight control.

[0177] Beneficial effects: Improving 3D representation accuracy: By optimizing the initial 3D representation incorporating camera parameter features, the true position, shape, and spatial relationships of objects in a scene can be more accurately reflected. In autonomous driving, it can precisely determine the distance and position of obstacles, providing more reliable information for safe driving; in mobile robot navigation, it helps generate more accurate maps and optimize path planning; in autonomous drone flight, it enables drones to perceive their surroundings more accurately, avoid collisions, and complete tasks efficiently.

[0178] Fully utilizing image information: This method not only uses the candidate image data itself to construct the 3D model, but also mines the information implicit in the camera parameters, achieving comprehensive utilization of image data. Compared with methods that rely solely on image content for 3D modeling, it can obtain richer and more accurate 3D scene information, improving the system's ability to perceive the environment.

[0179] Enhancing System Adaptability: This method is applicable to various scenarios such as intelligent driving, mobile robot navigation, and autonomous drone flight. Regardless of the complexity of the scenario, it can optimize the 3D representation through the analysis and fusion of camera parameters. This allows the system to maintain high performance in different application environments, enhancing its versatility and adaptability.

[0180] Figure 2 This is a schematic diagram of the structure of a first model provided in an embodiment of this application. Figure 2 As shown, the first model is mainly applied to practical autonomous driving scenarios such as end-to-end perception of key targets in complex and dynamic urban road environments, state prediction of the vehicle and other intelligent agents, and trajectory planning of the vehicle. The first model includes: a 3D feature extraction module, a visual encoding module, a text encoding module, and a feature analysis and processing module. The feature analysis and processing module includes: a decoding submodule, a cross-view 3D geometry enhancement submodule, and a prediction submodule. The decoding layer includes a starting decoding layer, an ending decoding layer, and at least one intermediate decoding layer.

[0181] The key technologies used in the first model include: 1. Multi-view data input: Multiple cameras (front-view, side-view, and rear-view) of the vehicle simultaneously capture candidate image data from the surroundings, forming a multi-view visual input.

[0182] 2. Fusion of 3D Geometric Perception and Inference: These candidate image data are simultaneously input into the frozen 3D feature extraction module and visual encoding module. The 3D feature extraction module is responsible for extracting accurate 3D geometric features (such as depth, camera pose, and point cloud information) from the multi-view candidate image data, thereby understanding the physical spatial structure of the scene.

[0183] 3. The visual encoding module extracts the 2D semantic features (first image feature data) of the target image data. The core innovative module—Cross-view 3D Geometric Enabler (CVGE)—begins to work. It decodes the reference 3D feature data output by the 3D feature extraction module and deeply integrates it into the 2D visual features through a hierarchical adaptive injection mechanism.

[0184] 4. Decision generation with a geometric foundation: After being empowered by CVGE, the features processed by the decoding and prediction submodules are visual features that have already incorporated precise geometric information. In this way, the first model can perform reasoning based on precise spatial relationships.

[0185] 5. The first model can directly output two forms of decision results: Natural language decision and interpretation: For example, it may generate a reasoning chain like this: "Identified that an oncoming silver car is turning left, and its trajectory may intersect with our current planned path; at the same time, a pedestrian on the right has entered the crosswalk, but is still far away; and cross-view detection of key risk targets, the status of the vehicle and other obstacles; Embodied actions or trajectories: More importantly, it can directly output specific, quantified control commands or future trajectory points, such as a smooth, safe driving trajectory that takes into account the future occupancy areas of all traffic participants."

[0186] The first model is a powerful multimodal large language model. In this invention, the first model takes as input a set of panoramic images (the aforementioned candidate image data) and corresponding language task instructions (the aforementioned instruction information). In benchmark testing, we use six panoramic images from the current frame. For the trajectory prediction task, we use three forward-looking angles. For the trajectory planning task, vehicle status information and navigation instructions are also incorporated into the text instructions. In the feature extraction stage, the multi-view images are processed by a visual encoding module to generate 2D visual embeddings (first image feature data), while the language instructions are converted into text embeddings (semantic feature data) by a text encoding module. These embeddings are concatenated and input into a feature analysis module, which generates corresponding text tags. The model is optimized by minimizing the standard cross-entropy loss.

[0187] 2. Hierarchical Adaptive Injection Mechanism To fully leverage cross-view geometric modeling capabilities and effectively enhance the first model to meet the stringent accuracy and robustness requirements of complex intelligent driving scenarios, we designed a hierarchical adaptive injection mechanism. Specifically, this mechanism employs a frozen 3D feature extraction module to perform cross-view 3D geometric modeling on a set of surround-view images, extracting visual features as reference 3D feature data. In particular, we preserve the original camera parameters and registered embeddings in the 3D features; these embeddings encode accurate multi-view information, crucial for accurate scene geometric representation. Furthermore, our CVGE allows 2D visual embeddings to query 3D representations, thereby capturing necessary cross-view geometric information. We then decouple the architecture of the basic feature analysis module to ensure that the hidden states in each decoder layer can be extracted. The 2D visual representation of each layer (the aforementioned first 2D feature data) is extracted from the hidden states using a fixed image ID position mask. Next, the aforementioned first 2D feature data and the aforementioned reference 3D feature data are input into CVGE to obtain the geometrically enhanced 3D visual embedding (the aforementioned fused 3D feature data). Considering the differences in embedded representations between network layers and their sensitivity to 3D information, CVGE employs a modular design, with each layer having a consistent structure but independent parameters. This allows the 2D visual features of each layer to adaptively learn and extract the most relevant geometric information. Finally, the first 2D feature data in the hidden state is replaced with a mask, and the input to the next decoder layer is obtained through residual connections.

[0188] 3. Cross-view 3D geometry enhancement module Existing integration schemes (such as simple feature stitching or addition) cannot fully leverage the advantages of VLM visual embedding to capture rich scene geometric information, thus limiting its adaptability in complex and highly dynamic intelligent driving scenarios. To address this issue, we designed a cross-view 3D geometry enhancement submodule, which establishes a learnable cross-modal interaction mechanism that enables 2D visual embedding to autonomously mine and integrate key information from 3D geometric features, thereby achieving deep geometry enhancement of visual representation.

[0189] The cross-view 3D geometry enhancement submodule receives first-dimensional (2D) feature data and shared reference 3D feature data as input, and outputs geometrically enhanced fused 3D feature data. First, the shared reference 3D feature data is flattened to align with the labels across all views, matching the total number of labels in the first-dimensional (2D) feature data to facilitate cross-view information interaction. Next, to optimize computational efficiency and reconcile the dimensionality differences between the two embeddings, we employ two independent multilayer perceptron (MLP) models for dimensionality compression. The compressed reference 2D feature data features serve as the query vector Q, while the reconstructed and compressed reference 3D feature data features serve as the key K and value V vectors.

[0190] In autonomous driving tasks, camera intrinsic and extrinsic parameters are typically used as known prior information. These parameters are crucial for trajectory planning tasks that rely on full 3D scene mapping. Therefore, we explicitly encode camera parameters and incorporate them into the generated key K and value V vectors. Finally, based on the obtained Q, K, and V representations, we design a cross-modal geometric attention fusion module. Unlike traditional static fusion methods, the multi-head attention mechanism enables the model to autonomously discover long-range deep relationships between 2D visual features and 3D geometric representations, while performing on-demand information fusion through dynamic weight calculation.

[0191] The fused 3D feature data is then augmented using an MLP to ensure that the final dimension is aligned with the first 2D feature data, thereby obtaining a 3D visual representation enhanced by cross-view geometric information (the aforementioned fused 3D feature data).

[0192] To implement the above embodiments, this application also proposes an instruction processing device.

[0193] Figure 3 This is a schematic diagram of the structure of an instruction processing device provided in an embodiment of this application.

[0194] like Figure 3 As shown, the device may include: Image selection module 310 is configured to, in response to receiving input instruction information, determine target image data from at least one candidate image data according to the instruction information; wherein, the at least one candidate image data is image data taken from different perspectives; The instruction processing module 320 is used to analyze and process the three-dimensional representation corresponding to the candidate image data and the target image data based on the instruction information, and obtain the reasoning result corresponding to the instruction information.

[0195] Optionally, the image selection module includes: The extraction submodule is used to extract the azimuth information contained in the instruction information and obtain the extraction result; An image selection submodule is used to determine target image data from at least one candidate image data based on the extraction results.

[0196] Optionally, the image selection submodule includes: The first selection unit is configured to respond to the instruction information containing the azimuth information, select a target viewpoint matching the target azimuth according to the target azimuth corresponding to the azimuth information, and use the candidate image data corresponding to the target viewpoint as the target image data; The second selection unit is configured to, in response to the instruction information not containing the orientation information, select each of the candidate image data as the target image data.

[0197] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, and will not be repeated here.

[0198] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing method embodiments.

[0199] To implement the above embodiments, this application also proposes a computer program product having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the foregoing method embodiments.

[0200] To implement the above embodiments, this application also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method described in the foregoing method embodiments.

[0201] Figure 4This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example, the electronic device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0202] Reference Figure 4 The electronic device 800 may include one or more of the following components: processing component 802, memory 804, power component 806, multimedia component 808, audio component 810, input / output (I / O) interface 812, sensor component 814, and communication component 816.

[0203] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0204] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of such data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0205] Power component 806 provides power to various components of electronic device 800. Power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.

[0206] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0207] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0208] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0209] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0210] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0211] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0212] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of an electronic device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0213] To implement the above embodiments, this application also proposes a chip, including: the chip includes a processing circuit configured to perform the methods provided in the foregoing embodiments.

[0214] Figure 5 This is a schematic diagram of the structure of a chip according to an embodiment of this application. See also... Figure 5 The diagram shown is a schematic representation of the structure of chip 1100, but it is not limited to this.

[0215] Chip 1100 includes processing circuitry 1101, which is configured to perform any of the above methods.

[0216] In some embodiments, chip 1100 further includes one or more interface circuits 1102. Optionally, the interface circuit 1102 is connected to memory 1103, and the interface circuit 1102 can be used to receive signals from memory 1103 or other devices, and the interface circuit 1102 can be used to send signals to memory 1103 or other devices. For example, the interface circuit 1102 can read instructions stored in memory 1103 and send the instructions to processing circuit 1101.

[0217] In some embodiments, the interface circuit 1102 performs at least one of the communication steps such as sending and / or receiving in the above method, while the processing circuit 1101 performs other steps.

[0218] In some embodiments, the terms interface circuit, interface, transceiver pin, transceiver, etc., can be used interchangeably.

[0219] In some embodiments, chip 1100 further includes one or more memories 1103 for storing instructions. Optionally, all or part of the memories 1103 may be located outside of chip 1100.

[0220] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0221] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0222] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0223] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0224] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0225] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes one or a combination of the steps of the method embodiments.

[0226] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0227] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. An instruction processing method, characterized in that, include: In response to receiving input instruction information, a target image data is determined from at least one candidate image data according to the instruction information; wherein, the at least one candidate image data is image data taken from different perspectives; Based on the instruction information, the three-dimensional representation corresponding to the candidate image data and the target image data are analyzed and processed to obtain the inference result corresponding to the instruction information.

2. The method according to claim 1, characterized in that, The step of determining the target image data from at least one candidate image data according to the instruction information includes: Extract the azimuth information contained in the instruction information to obtain the extraction result; Based on the extraction results, target image data is determined from at least one candidate image data.

3. The method according to claim 2, characterized in that, The step of determining the target image data from at least one candidate image data based on the extraction result includes any one of the following: In response to the instruction information containing the azimuth information, a target viewpoint matching the target azimuth is selected according to the target azimuth corresponding to the azimuth information, and the candidate image data corresponding to the target viewpoint is used as the target image data; In response to the fact that the instruction information does not contain the orientation information, each of the candidate image data is used as the target image data.

4. The method according to any one of claims 1-3, characterized in that, The step of analyzing and processing the 3D representation corresponding to the candidate image data and the target image data based on the instruction information to obtain the inference result corresponding to the instruction information includes: The instruction information, the candidate image data, and the target image data are input into a pre-trained first model to extract the three-dimensional representation corresponding to the candidate image data, and inference is performed based on the instruction information, the three-dimensional representation corresponding to the candidate image data, and the target image data to obtain the inference result corresponding to the instruction information; wherein, the first model includes: a three-dimensional feature extraction module, a visual encoding module, a text encoding module, and a feature analysis and processing module.

5. The method according to claim 4, characterized in that, The step of extracting the three-dimensional representation corresponding to the candidate image data and performing inference based on the instruction information, the three-dimensional representation corresponding to the candidate image data, and the target image data to obtain the inference result corresponding to the instruction information includes: The instruction information is input into the text encoding module for feature extraction to obtain semantic feature data corresponding to the instruction information; The target image data is input into the visual encoding module for feature extraction to obtain the first image feature data corresponding to the target image data; The three-dimensional representation corresponding to the candidate image data is input into the three-dimensional feature extraction module for feature extraction to obtain reference three-dimensional feature data; The reference 3D feature data, the semantic feature data, and the first image feature data are input into the feature analysis and processing module for fusion and decoding to obtain the inference result.

6. The method according to claim 5, characterized in that, The feature analysis and processing module includes: a decoding submodule, a cross-view 3D geometry enhancement submodule, and a prediction submodule; the step of inputting the reference 3D feature data, the semantic feature data, and the first image feature data into the feature analysis and processing module for fusion and decoding to obtain the inference result includes: The semantic feature data and the first image feature data are concatenated and input into the decoding layer in the decoding submodule for decoding. The feature data output from each decoding layer in the decoding submodule and the reference 3D feature data are input into the cross-view 3D geometry enhancement submodule for fusion to obtain fused 3D feature data. The fused 3D feature data is input into the decoding submodule or the prediction module for processing to obtain the inference result.

7. The method according to claim 6, characterized in that, The step of concatenating the semantic feature data and the first image feature data and inputting them into the decoding layer of the decoding submodule for decoding includes: The semantic feature data and the first image feature data are concatenated and input into the starting decoding layer of the decoding module for decoding to obtain the first two-dimensional feature data; wherein, the decoding submodule includes multiple decoded layers in series, and the decoded layers include a starting decoding layer, an ending decoding layer and at least one intermediate decoding layer.

8. The method according to claim 7, characterized in that, The step of fusing the feature data output from each decoding layer in the decoding submodule with the reference 3D feature data into the cross-view 3D geometry enhancement submodule to obtain fused 3D feature data includes: For each of the decoding layers, the first two-dimensional feature data output by the decoding layer is input into the fusion with the reference three-dimensional feature data to obtain the fused three-dimensional feature data corresponding to the decoding layer.

9. The method according to claim 8, characterized in that, The step of inputting the fused 3D feature data into the decoding submodule or the prediction module for processing to obtain the inference result includes: The fused three-dimensional feature data corresponding to the decoding layer is fused with the first two-dimensional feature data output by the decoding layer, and then input into the next adjacent decoding layer for decoding. The first two-dimensional feature data output by the last decoding layer is fused with the fused three-dimensional feature data corresponding to the last decoding layer, and then input into the prediction submodule for classification to obtain the inference result.

10. The method according to claim 8, characterized in that, The step of inputting the first two-dimensional feature data output by the decoding layer into the fused three-dimensional feature data corresponding to the decoding layer by fusing it with the reference three-dimensional feature data includes: The reference three-dimensional feature data is unfolded to obtain the second two-dimensional feature data; The first two-dimensional feature data and the second two-dimensional feature data are dimensionally compressed, and attention is calculated based on the compressed first two-dimensional feature data and the second two-dimensional feature data to obtain an attention matrix; The compressed first two-dimensional feature data and the second two-dimensional feature data are weighted and fused according to the attention matrix, and the dimensions are restored to obtain the fused three-dimensional feature data.

11. The method according to claim 4, characterized in that, The steps for obtaining the three-dimensional representation corresponding to the candidate image data include: Perform 3D modeling based on the candidate image data to obtain an initial 3D representation; Feature extraction is performed on the camera parameters corresponding to the candidate image data to obtain parameter features; The parameter features are fused with the initial three-dimensional representation to obtain the fused three-dimensional representation.

12. An instruction processing device, characterized in that, include: An image selection module is configured to, in response to receiving input instruction information, determine target image data from at least one candidate image data according to the instruction information; wherein, the at least one candidate image data is image data taken from different perspectives; The instruction processing module is used to analyze and process the three-dimensional representation corresponding to the candidate image data and the target image data based on the instruction information, and obtain the inference result corresponding to the instruction information.

13. The apparatus according to claim 12, characterized in that, The image selection module includes: The extraction submodule is used to extract the azimuth information contained in the instruction information and obtain the extraction result; An image selection submodule is used to determine target image data from at least one candidate image data based on the extraction results.

14. The apparatus according to claim 13, characterized in that, The image selection submodule includes: The first selection unit is configured to respond to the instruction information containing the azimuth information, select a target viewpoint matching the target azimuth according to the target azimuth corresponding to the azimuth information, and use the candidate image data corresponding to the target viewpoint as the target image data; The second selection unit is configured to, in response to the instruction information not containing the orientation information, select each of the candidate image data as the target image data.

15. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method as described in any one of the preceding claims 1-11.

16. A vehicle, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method as described in any one of the preceding claims 1-11.

17. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of the preceding claims 1-11.

18. A chip, characterized in that, The chip includes processing circuitry configured to perform the method described in any one of claims 1-11.

19. A computer program product, characterized in that, It includes a computer program, which, when executed by a processor, implements the method as described in any one of claims 1-11.