Guidance method and device, vehicle, medium and program product

Through multi-view cameras and multi-modal large models, the problem of finding cars in large parking lots is solved, and efficient and accurate vehicle positioning and navigation are achieved.

CN120279472APending Publication Date: 2025-07-08XIAOMI EV TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510412529.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In large parking lots, it is difficult for users to quickly find parked vehicles. The existing technologies such as inaccurate GPS positioning, limited coverage of Bluetooth beacons, and unintuitive picture display, resulting in low efficiency and difficulty in finding a car.

Method used

通过多视角摄像头拍摄环境图像,利用多模态大模型生成结构化环境描述信息,生成指引信息,指导用户快速找到目标对象。

Benefits of technology

It improves the accuracy and reliability of environmental data, reduces manual labeling costs, enhances the ability to adapt to complex environments, and improves the convenience of car search and user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279472A_ABST
    Figure CN120279472A_ABST
Patent Text Reader

Abstract

The invention relates to a guiding method and device, a vehicle, a medium and a program product in the technical field of computer vision and artificial intelligence, and the method comprises the steps: carrying out the processing of environment images shot at multiple visual angles, and generating a target environment image; inputting the target environment image into a pre-trained target model to obtain environment description information output by the target model; according to the environment description information, guidance information is generated, and the guidance information is used for describing features of a target object in the environment. The guidance information is generated after the environment images shot at multiple visual angles are processed, the problems that single-visual-angle image information is insufficient, blocked or distorted can be solved, the accuracy and reliability of the environment data are improved, the guidance information for the target object is generated, different application scenes can be adapted, and the adaptability to the complex environment is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of computer vision and artificial intelligence, and particularly relates to a guiding method, device, vehicle, medium and program product. Background Art

[0002] When a vehicle is parked in a large parking lot such as an underground parking lot or an outdoor parking lot of a shopping mall, it is very inconvenient for the vehicle owner to find the vehicle. For example, a user drives a vehicle into a large underground parking lot (such as a shopping mall parking lot) and parks it in a certain area. Due to the large scale of the parking lot, complex floor structures, and usually small or blocked parking space numbers, it is difficult for the user to quickly find their own vehicle when returning. Therefore, it takes a lot of energy and time to find the vehicle. Summary of the Invention

[0003] To overcome the problems existing in the related art, the present disclosure provides a guiding method, device, vehicle, medium and program product.

[0004] According to a first aspect of an embodiment of the present disclosure, a guiding method is provided, including: Processing environmental images captured from multiple perspectives to generate a target environmental image; Inputting the target environmental image into a pre-trained target model to obtain environmental description information output by the target model; Generating guiding information according to the environmental description information, where the guiding information is used to describe the characteristics of a target object in the environment.

[0005] Optionally, the environmental image includes a first environmental image captured by the vehicle from multiple perspectives, and the guiding information includes first guiding information, where the first guiding information is used to be pushed to a mobile terminal to display the characteristics of a first target object in the environment where the vehicle is located.

[0006] Optionally, the processing the environmental images captured from multiple perspectives to generate a target environmental image includes: Responding to a vehicle finding request of the mobile terminal, processing the environmental images captured from multiple perspectives to generate a target environmental image, where the vehicle finding request is generated according to a user operation on the mobile terminal.

[0007] Optionally, the first environmental image is captured by at least one of the following methods: capturing in response to the user leaving the vehicle, capturing in response to receiving the vehicle finding request.

[0008] Optionally, the first target object includes at least one of the following: parking space number, wall, pillar, facility, ground.

[0009] Optionally, the environmental image includes a second environmental image captured by a home device from multiple perspectives, the guiding information includes second guiding information, and the second guiding information is used to be pushed to a mobile terminal to display the characteristics of a second target object in the environment where the home device is located.

[0010] Optionally, the second target object includes at least one of the following: a pet, a door or window, a home device.

[0011] Optionally, the environmental image includes a third environmental image captured by a robot from multiple perspectives, the guiding information includes third guiding information, and the third guiding information is used to characterize the characteristics of a third target object in the environment where the robot is located.

[0012] Optionally, the environmental image includes a fourth environmental image captured by a mobile terminal from multiple perspectives, the guiding information includes fourth guiding information, and the fourth guiding information is used to display the characteristics of a fourth target object in the environment where the mobile terminal is located.

[0013] Optionally, generating the guiding information according to the environmental description information includes: Generating a feature description of a target object in the environment according to the environmental description information; Verifying the feature description according to the preset verification information corresponding to each target object; Generating the guiding information according to the verification results of each feature description.

[0014] Optionally, the preset verification information includes at least one of the following: format information, feature range information.

[0015] Optionally, inputting the target environmental image into a pre-trained target model to obtain environmental description information output by the target model includes: Inputting the target environmental image into a pre-trained target model, and the target model extracts spatial features from the target environmental image; Determining a mapping relationship from the target environmental image from multiple perspectives to the environmental description according to the spatial features; Obtaining the environmental description information output by the target model according to the mapping relationship.

[0016] Optionally, processing the environmental images captured from multiple perspectives to generate a target environmental image includes: Determining validity information of the environmental images captured from multiple perspectives; When the validity information of the environmental image indicates that the environmental image meets the validity requirements, performing pixel normalization on each environmental image; Unifying the sizes of the normalized environmental images; Crop the environment image with unified dimensions to generate the target environment image.

[0017] According to the second aspect of the embodiments of the present disclosure, there is provided a guiding device, including: A first generation module, configured to process environment images captured from multiple perspectives to generate a target environment image; An input module, configured to input the target environment image into a multi-modal large model to obtain environment description information output by the multi-modal large model; A second generation module, configured to generate guiding information according to the environment description information, where the guiding information is used to describe the characteristics of a target object in the environment.

[0018] According to the third aspect of the embodiments of the present disclosure, there is provided a vehicle, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of the first aspect.

[0019] According to the fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to any one of the first aspect are implemented.

[0020] According to the fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of the first aspect are implemented.

[0021] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: By processing environment images captured from multiple perspectives to generate guiding information, problems such as insufficient information, occlusion, or distortion in single-perspective images can be overcome, improving the accuracy and reliability of environmental data. By parsing the target environment image through a pre-trained target model and automatically outputting structured and semantic environment description information, problems such as image blur caused by factors such as light can be effectively resisted, reducing the cost of manual annotation, and improving the efficiency and automation level of environmental analysis. Based on the environment description information, guiding information for the target object is generated, which can adapt to different application scenarios and enhance the adaptability to complex environments. According to the guiding information, users can quickly locate, improving the convenience of environmental scene identification and the user interaction experience in complex scenarios.

[0022] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings

[0023] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0024] Figure 1 is a flowchart of a guidance method shown according to an exemplary embodiment.

[0025] Figure 2 is a schematic diagram of an application scenario of a guidance method shown according to an exemplary embodiment.

[0026] Figure 3 is a flowchart of another guidance method shown according to an exemplary embodiment.

[0027] Figure 4 is a flowchart of another guidance method shown according to an exemplary embodiment.

[0028] Figure 5 is a block diagram of a guidance device shown according to an exemplary embodiment.

[0029] Figure 6 is a block diagram of a vehicle shown according to an exemplary embodiment.

[0030] Figure 7 is a block diagram of a cloud server shown according to an exemplary embodiment. Detailed Description of the Embodiments

[0031] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0032] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining the authorization given by the owner of the corresponding device.

[0033] Before introducing a car-finding method provided by an embodiment of the present disclosure, first introduce the technical problems existing in the related scenarios. To improve the efficiency of finding a car in a parking lot or the like, a function of viewing parking space photos is provided. However, the car owner needs to view the pictures to determine the position where the vehicle is parked. For vehicle positioning technologies based on GPS or Bluetooth, although the vehicle can be positioned, affected by the weak signal in the underground parking lot, the vehicle positioning is inaccurate, or the user terminal cannot receive the vehicle positioning information in the underground parking lot. Moreover, the deployment cost of Bluetooth beacons is high and the coverage range is limited, and only rough positioning information can be provided, and the detailed features around the parking space cannot be described, so the difficulty of finding a car is still large and the efficiency is low.

[0034] Regarding the function of parking and taking pictures, when the user leaves the vehicle, four panoramic views of the front, rear, left, and right of the vehicle and the top view according to the reverse image are automatically taken. However, there are the following three problems with only providing pictures to the user: 1. The pictures are displayed in the terminal application, and the picture is small. The user needs to click to zoom in to view, and the operation is cumbersome, and the interaction is not convenient or intuitive; 2. Due to the complex parking space environment (such as small text on the wall, being blocked, etc.), it is difficult for the user to quickly identify key content such as floor numbers and area information, or the parking space number is blocked by the parked vehicle, and the floor information and area information on the ground are blocked, making information identification difficult. 3. The panoramic pictures are divided into front, rear, left, and right, and the multi-view pictures are displayed separately. The user needs to have good spatial imagination to understand the spatial relationship and find the car in combination with the spatial relationship, and the difficulty and convenience are both low.

[0035] In view of this, an embodiment of the present disclosure provides a guiding method, aiming to improve the convenience and efficiency of, for example, finding a car, while reducing the difficulty of finding a car and improving the visual understanding of finding a car, so as to improve the user's car-finding experience.

[0036] Figure 1 It is a flowchart of a guiding method shown according to an exemplary embodiment. It can be explained that the guiding method provided by the embodiment of the present disclosure can be applied to the intelligent cockpit of a vehicle, for example, the cockpit domain controller of the vehicle, or can also be applied to a cloud server, etc. For the sake of clear and simple description of the implementation scenario, this method is taken as an example of being applied to the cockpit domain controller for exemplary illustration. As Figure 1 shown, the guiding method may include the following steps.

[0037] In step S11, the environmental images captured from multiple perspectives are processed to generate a target environmental image; Among them, multiple perspectives may be multiple views of the same environment captured by one or more cameras from different angles or positions. The environmental image is an image including various shooting objects in the scene. The target environmental image is an image after the environmental image is screened, processed, de-duplicated, etc.

[0038] In the embodiments of the present disclosure, environmental images are captured from different perspectives by a camera. These images may contain overlapping parts, or may present the characteristics of different photographed objects due to perspective differences. Next, image processing techniques (such as image registration, stitching, fusion, etc.) are used to process these images.

[0039] Among them, image registration can align images from different perspectives. Images captured by different cameras at the same time can be registered, or images captured at different times can be arranged and registered in chronological order. Image stitching can merge the registered images into a larger image. For example, environmental images captured at two different times actually reflect different regions of the same wall, so an image of the same wall can be stitched. Image fusion can adjust parameters such as brightness and contrast to make the stitched image look more natural and consistent. Finally, through these processing steps, a target environmental image containing all important features in the environment is generated.

[0040] In step S12, the target environmental image is input into a pre-trained target model to obtain environmental description information output by the target model. Among them, the target model can be a model trained based on deep learning. For example, it can be a multimodal large model, which can be trained with images collected in various scenarios to learn the mapping relationship from multi-perspective images to scene feature descriptions. For example, the multimodal large model can be a multimodal Transformer model. The environmental description information can be used to describe the target objects existing in the environment in words and describe the characteristics of the target objects. For example, in a parking scenario, the environmental description information can describe in words the column existing on the right side of the vehicle and the characteristics such as the shape and color of the column. Exemplarily, it can be described by the following words: Target object - column, characteristics - round, yellow.

[0041] In the embodiments of the present disclosure, the target environmental image is used as an input and passed to a pre-trained target model. The target model can be based on a deep learning architecture (such as the convolutional neural network CNN) and has powerful image analysis and understanding capabilities. The target model performs operations such as feature extraction, classification, and recognition on the input image through its internal parameters and algorithms, and finally generates an output containing environmental description information. These information may include various objects, as well as the positions, sizes, categories, etc. of various objects, and may also include the overall layout of the environment, lighting conditions, etc.

[0042] In step S13, guidance information is generated according to the environmental description information, and the guidance information is used to describe the feature description of the target object in the environment.

[0043] Among them, the guidance information can guide or describe the characteristics of specific target objects in the environment, so as to be used to find the target object according to these characteristics. The target object is an object that is particularly concerned or needs to be identified and located in the environment. It should be noted that for the environmental images collected in different scenarios, the target object can be different. For example, in the environmental images collected in the parking lot scenario, the target object can include vehicles with special colors, walls, columns, etc.; in the environmental images collected in the home scenario, the target object can include people at home, household appliances, doors and windows, etc.

[0044] In the embodiments of the present disclosure, by using the environmental description information obtained from the target model, the characteristics of a specific target object are further analyzed and extracted. For example, some objects are deleted to obtain the characteristic description of the specific target object. This may also involve operations such as screening, summarizing, and reasoning on the description information to extract the key information related to the target object. Then, according to these key information, the guidance information is generated.

[0045] The guidance information can include characteristic descriptions such as the name, location, shape, and color of the target object, and can also include other useful information related to the target object (such as the relative position relationship with the surrounding environment, possible movement trajectories, etc.). These information can be used for tasks such as navigation, positioning, and identification, providing accurate descriptions and guidance about the target object for users or systems.

[0046] In the above technical solution, by processing the environmental images taken from multiple perspectives to generate the guidance information, the problems of insufficient information, occlusion, or distortion in a single perspective image can be overcome, improving the accuracy and reliability of the environmental data. By parsing the target environmental images through a pre-trained target model and automatically outputting structured and semantic environmental description information, the problems such as image blurring caused by factors such as light can be effectively resisted, reducing the manual annotation cost, and improving the efficiency and automation level of environmental analysis. Based on the environmental description information, generating the guidance information for the target object can adapt to different application scenarios and enhance the adaptability to complex environments. According to the guidance information, users can quickly locate, improving the convenience of environmental scene recognition and the user interaction experience in complex scenarios.

[0047] Optionally, the environmental image includes a first environmental image taken by the vehicle from multiple perspectives, the guidance information includes first guidance information, and the first guidance information is used to be pushed to the mobile terminal to display the characteristics of the first target object in the environment where the vehicle is located.

[0048] Among them, the first target object is a specific object with recognition in the environment where the vehicle is located. For example, walls, column heads, ground, vehicles, etc. Furthermore, the characteristics of the first target object can be displayed to the user, and the user does not need to combine spatial imagination to judge the position of the vehicle.

[0049] See Figure 2 As shown, the vehicle can capture a first environmental image when preset conditions are met. For example, it can capture images when it recognizes a change in vehicle parking around the vehicle. For instance, when a vehicle parked in an adjacent parking space drives out of the parking space, 4 panoramic views and 1 top view are captured, or when a vehicle parks in an adjacent vacant parking space, 4 panoramic views and 1 top view are captured. In this way, as the surrounding objects change, environmental images including different objects in the environment can be captured, which is conducive to improving the accuracy of the output environmental description information.

[0050] Furthermore, the vehicle can generate guidance information locally based on the guidance method of the present disclosure for the captured first environmental image and push the push information to the mobile terminal, or it can upload the captured first environmental image to the cloud server, and the cloud server generates guidance information based on the guidance method of the present disclosure and pushes the push information to the mobile terminal.

[0051] It can be noted that the timing of pushing the guidance information in the present disclosure can be flexibly arranged. For example, when the user leaves the vehicle, after capturing the first environmental image, the guidance information is generated and pushed to the mobile terminal. In this way, when the user needs to view it, they can directly open the application to view. It can also be, for example, in the foregoing embodiment, when it recognizes a change in vehicle parking around the vehicle, the first environmental image is captured, and then the guidance information is generated in real time and pushed to the mobile terminal. In this way, the user can conveniently view the feature description information of the target object in the environment where the vehicle is located at any time, improving the convenience and efficiency of finding the vehicle.

[0052] Optionally, the processing of the environmental images captured from multiple perspectives to generate a target environmental image includes: In response to a vehicle-finding request from the mobile terminal, the environmental images captured from multiple perspectives are processed to generate a target environmental image, and the vehicle-finding request is generated according to a user operation on the mobile terminal.

[0053] Among them, the vehicle-finding request can be initiated by the user through the mobile terminal and is used to request to display the feature description of the target object in the parked vehicle environment.

[0054] In the embodiments of the present disclosure, when a user initiates a vehicle search request through a mobile terminal (such as a mobile phone application), multi-view environment images stored on a cloud server or a cockpit domain controller can be retrieved first. These images are captured from different angles by multiple cameras installed around the vehicle or in the parking lot when the vehicle is parked or when the user leaves the vehicle. Subsequently, image processing techniques (such as image stitching, denoising, enhancement, etc.) are used to process these images to generate a clearer and more comprehensive target environment image, which can more accurately reflect the environment when the vehicle is parked. Then, cross-modal environmental description information is generated and output through a multi-modal large model, and finally guidance information is obtained, which helps the user quickly find the vehicle and improves the efficiency and convenience of vehicle search.

[0055] Optionally, the first environmental image is captured by at least one of the following methods: capturing in response to the user leaving the vehicle; capturing in response to receiving the vehicle search request.

[0056] It should be noted that capturing in response to the user leaving the vehicle can be that when the user finishes parking and leaves the vehicle, the camera is automatically triggered to capture the current environmental image. This can ensure that there is the latest environmental image for reference when the user needs to search for the vehicle. It can also be that when the user leaves the vehicle with automatic parking enabled, during or after the vehicle's automatic parking process, the camera is automatically triggered to capture the current environmental image.

[0057] Capturing in response to receiving the vehicle search request is to trigger the camera to capture the current environmental image after receiving the user's vehicle search request. This can reflect the environmental conditions when the vehicle search request is issued, which helps the user find the vehicle based on the current environment and avoid changes in the target objects in the environmental image caused by the parking and leaving, or the parking and driving away of surrounding vehicles during the process.

[0058] Optionally, the first target object includes at least one of the following: parking space number, wall, pillar, facility, ground.

[0059] In the embodiments of the present disclosure, the parking space number can be used to identify and distinguish each parking space in the parking lot. Each parking space has a unique number, which can facilitate the vehicle owner to quickly find the parking space. The wall can enclose and separate spaces. It can be an indoor wall or an outdoor wall, and the characteristics of the wall material, color, height, etc. can be described through the first guidance information.

[0060] In the embodiments of the present disclosure, columns can be characterized by shapes such as circular columns, square columns, and polygonal columns, or can be characterized by colors, materials, etc. Facilities can be various equipment and devices equipped in buildings or sites. For example, the facilities in a parking lot may include lighting equipment, monitoring equipment, fire-fighting facilities, signboards, etc. The characterization of these facilities can improve the convenience for users to find their vehicles. The ground can be characterized by materials such as cement ground and resin ground, or can be characterized by colors, etc.

[0061] Combined with Figure 2 , it is illustrated through the following scenario example. Suppose a user drives a vehicle into an underground parking lot (such as a shopping mall parking lot) and parks the vehicle in a certain area of the parking lot. Due to the large scale of the parking lot, complex floor structures, and usually small or blocked parking space numbers, it may be difficult for the user to quickly find their vehicle when returning. Even if the user is not far from the vehicle, it is difficult to pre-know the characteristics of target objects such as the walls of the parking space or the column heads. Therefore, the user can open the application on the mobile terminal (such as a mobile phone), and by selecting "assist in finding the vehicle" on the application, the first guidance information can be displayed on the mobile phone, or the user can assist in finding the vehicle by pressing the vehicle-finding button on the mobile terminal (such as a smart car key), and then the first guidance information can be displayed on the display screen of the smart car key.

[0062] In this example, when the user leaves the vehicle, before locking the vehicle and cutting off the power, the vehicle can automatically take four panoramic views of the front, rear, left, and right around the vehicle and the top view of the reverse image through the in-vehicle camera. Then, these images are uploaded to the cloud server as environmental images. After receiving the vehicle-finding request, the target model deployed on the cloud server processes them based on multi-view visual understanding to generate the environmental description information shown in Table 1. Of course, it can also be directly saved in the vehicle's cockpit domain controller, and the cockpit domain controller processes it based on multi-view visual understanding after receiving the vehicle-finding request to generate the environmental description information shown in Table 1:

[0063] Table 1 Furthermore, the environmental description information can be sorted out to obtain the guidance information for finding the vehicle as shown in Table 2:

[0064] Table 2 In this way, the above guidance information can be displayed on the display screen of the mobile terminal. Based on this guidance information, the user can search for the vehicle. By outputting environmental description information through a model based on multi-perspective environmental images, problems such as image blurring caused by factors such as light can be effectively resisted. Through evaluation, it is expected that the indoor parking space number recognition accuracy rate will reach 91%, the outdoor parking space number recognition accuracy rate will reach 98%, and the target object feature description accuracy rate will be above 88%. The user can not only conveniently view the features of the target object, and then quickly find the location of the vehicle according to the features, improving the travel convenience of the vehicle owner, but also with relatively high accuracy.

[0065] Optionally, the environmental image includes a second environmental image captured by a home device from multiple perspectives, and the guidance information includes second guidance information for pushing to the mobile terminal to display the feature description of the second target object in the environment where the home device is located.

[0066] Among them, the second environmental image can be a home environmental image captured by a home device (such as a smart camera, a smart camera, etc.) from different perspectives.

[0067] In the embodiments of the present disclosure, the home device captures the home environment from multiple different perspectives through a built-in camera or other image acquisition devices. By performing multi-modal model processing on these multi-perspective second environmental images, the features of the second target object in the home environment can be extracted, and then the second guidance information can be generated.

[0068] In the embodiments of the present disclosure, through the feature extraction algorithm, the position and contour of the second target object in the image can be automatically recognized, and the corresponding feature description can be generated. Furthermore, the second guidance information representing the features of the second target object is generated. For example, the position and contour of the express delivery on the sofa are recognized, and then the user can be guided that the express delivery has been taken home and there is no need to pick up the express delivery again.

[0069] Optionally, the second target object includes at least one of the following: a pet, a door or window, a home device.

[0070] In the embodiments of the present disclosure, if the target object is a pet, the guidance information may include the position of the pet and whether it is in an active state; if the target object is a door or window, the guidance information may include the state (open or closed) of the door or window and the position information.

[0071] Exemplarily, in the field of smart home or a home security system, the feature description of the distribution and state of objects or people in the room can be generated through the environmental images in the home environment collected by multiple cameras. For example, when the user is not at home, a text description such as "there is a package on the living room sofa" or "the kitchen window is not closed" can be automatically generated and then pushed to the mobile terminal for prompting.

[0072] In a possible implementation manner, abnormal accidents or events in public places can be monitored through multi-perspective environmental images in public places, abnormal events or accidents can be identified, and detailed scene descriptions can be provided.

[0073] Optionally, the environmental image includes a third environmental image captured by a robot from multiple perspectives, the guidance information includes third guidance information, and the third guidance information is used to characterize the feature description of a third target object in the environment where the robot is located.

[0074] In the embodiments of the present disclosure, in the field of service robots, it can be used for robots to understand the surrounding environment. For example, in hotel or hospital scenarios, robots can identify floor information, room numbers, landmarks, etc. through multi-perspective environmental images, so as to better complete delivery or guidance tasks. The third guidance information can also be pushed to the user terminal to facilitate quickly finding the robot according to the feature description of the third target object in case of, for example, the robot malfunctioning and being unable to return.

[0075] Optionally, the environmental image includes a fourth environmental image captured by a mobile terminal from multiple perspectives, the guidance information includes fourth guidance information, and the fourth guidance information is used to display the feature description of a fourth target object in the environment where the mobile terminal is located.

[0076] In the embodiments of the present disclosure, environmental images can be captured from multiple perspectives by a mobile terminal, and then spatial features and location information can be identified. For example, in an indoor navigation system, such as in shopping malls, airports, exhibition halls and other scenarios, it can help users quickly find the target object. For example, in the airport pick-up scenario, A can capture multi-angle fourth environmental images, and then generate fourth guidance information to be pushed to B's mobile terminal, and then B can quickly find A's location according to the fourth guidance information.

[0077] Optionally, referring to Figure 3 As shown, in step S13, generating the guidance information according to the environmental description information includes: In step S131, generating the feature description of the target object in the environment according to the environmental description information; In the embodiments of the present disclosure, the feature information of the target object can be extracted from the environmental description information. These feature information may come from pixel values, shape contours, color distributions, etc. in the image, or may come from keywords, phrases, etc. in the descriptive text.

[0078] In the embodiments of the present disclosure, for example, a JSON parsing tool can be used to extract fields to parse the original output of the multi-modal large model and generate a structured feature description, such as: { "Parking space number": "B3-125", "Wall color": "Grey", "Column color": "Yellow", "Facility information": ["Elevator", "Charging pile"] } In step S132, according to the preset verification information corresponding to each of the target objects, verify the feature description; Among them, the preset verification information may be information preset for verifying whether the feature description of the target object is accurate, and may include standard feature values, feature ranges, allowable errors, etc.

[0079] In the embodiments of the present disclosure, the extracted feature description is compared with the preset verification information to verify the accuracy and reliability of the feature description. The preset verification information may come from known features of the target object, verification information input by the user, historical data, etc.

[0080] Among them, for the feature description of each target object, it is compared with the preset verification information one by one. The comparison may involve matching of feature values, judgment of feature ranges, similarity calculation, etc. According to the comparison results, determine the accuracy of the feature description.

[0081] In step S133, generate the guiding information according to the verification results of each of the feature descriptions.

[0082] In the implementation of the present disclosure, based on the verification results of the feature description, generate guiding information. The guiding information may include the location, status, operation suggestions, etc. of the target object. Classify the target object according to the verification results, such as verified, unverified, suspected error, etc. Then, according to the classification results and user needs, generate corresponding guiding information. For example, for a verified target object, location information, status information, etc. can be generated; for an unverified or suspected error target object, guiding information can be generated and output after performing default processing operations.

[0083] Optionally, the preset verification information includes at least one of the following: format information, feature range information.

[0084] Among them, the format information may be whether the current target object and the feature description of the target object are output in a preset format. For example, when the target object is a parking space number, the format information may be "letter + number". If the feature description is number + letter, guiding information can be generated after swapping the positions.

[0085] Among them, the feature range information may be whether the primary color values representing the color are within the normal range, or whether the feature values output for the wall height or material are within the normal range.

[0086] For example, the rationality of the feature description of the target object in the generated vehicle parking environment is verified to exclude abnormal results. For example, the verification of the parking space number can be whether the parking space number conforms to the common format (such as floor number + digital number). The color verification can be whether the colors of the walls and columns are within a reasonable range (such as gray, yellow, white, etc.). If the verification result indicates that the preset verification information is not satisfied, the default information can be selected for output. If the verification result indicates that the preset verification information is satisfied, the output string can be directly formed by rules to generate the guidance information for pushing or displaying.

[0087] Optionally, in step S12, the inputting the target environment image into the pre-trained target model to obtain the environment description information output by the target model includes: Inputting the target environment image into the pre-trained target model, and the target model extracts spatial features from the target environment image; In the embodiments of the present disclosure, the target environment image is used as input data and sent into the model through the input layer of the model for spatial feature extraction. The training process of the target model may involve the input of a large amount of labeled data and the optimization of model parameters through deep learning algorithms, so that it can accurately identify the spatial features in the image.

[0088] Among them, the target model can use, for example, a convolutional neural network (CNN) or other deep learning architectures to perform multi-level feature extraction on the input image. These features usually include low-level features such as edges, textures, shapes, and high-level features such as objects and scenes. The spatial features specifically refer to the features that can reflect the positional relationship of objects in the image. For example, the target model can gradually extract the spatial features of the image through structures such as convolutional layers and pooling layers. These features exist in the form of feature maps, and each feature map corresponds to a specific feature in the image.

[0089] According to the spatial features, determine the mapping relationship from the target environment image from multiple perspectives to the environment description; In the embodiments of the present disclosure, since the target environment may consist of multiple images taken from different perspectives, the model needs to be able to integrate the information in these images to generate a unified environment description. This usually involves complex operations such as feature fusion and perspective transformation to ensure that the model can extract consistent environmental features from multiple perspectives.

[0090] Among them, the target model may adopt an attention mechanism, a recurrent neural network (RNN), or a graph neural network (GNN), etc., to integrate the spatial features in the multi-perspective images. The target model can learn the mapping relationship from the multi-perspective images to the environment description.

[0091] According to the mapping relationship, obtain the environment description information output by the target model.

[0092] In the embodiments of the present disclosure, after the target model learns the mapping relationship from multi-view images to environmental descriptions, it can utilize this relationship to generate environmental description information. Such information may include attributes such as the positions, shapes, sizes, colors of objects in the environment, as well as the spatial relationships between them.

[0093] Among them, the target model can convert the integrated spatial features into environmental description information through a fully connected layer or other output layer structures. Such information usually exists in the form of text, vectors, or graph structures.

[0094] For example, the generated target environmental image after processing is input into a multi-modal large model, and the multi-modal large model performs reasoning to generate environmental description information. Among them, the architecture of the large model may include: Visual encoder: used to extract the spatial features of the input image.

[0095] Cross-modal fusion module: used to combine visual information and text prompts to generate a unified feature description.

[0096] Decoder: used to generate a text description of the structured output.

[0097] The environmental description information output by the target model may be: Parking space number: the current parking space number (if any), otherwise output the numbers of nearby parking spaces such as adjacent or spaced parking spaces.

[0098] Wall: describe the color of the wall closest to the vehicle.

[0099] Column: describe the color of the nearest column.

[0100] Facility information: determine whether there are elevators, stairs, charging piles, etc. nearby.

[0101] Optionally, in step S11, the processing of the environmental images captured from multiple perspectives to generate a target environmental image includes: Determine the validity information of the environmental images captured from multiple perspectives; In the embodiments of the present disclosure, the number of environmental images can be preset in advance, and then check whether the number of environmental images is accurate to determine the validity information of the environmental images captured from multiple perspectives. For example, check whether the 5 uploaded environmental images (4 panoramic views and 1 top view) are complete and the number is correct, and then determine the validity information of the environmental images.

[0102] In the embodiments of the present disclosure, the grayscale information of the environmental image is determined, and the validity information of each environmental image is determined according to the grayscale information of each environmental image and the grayscale requirement. Alternatively, the clarity of each environmental image is determined, and the validity information of each environmental image is determined according to the clarity of each environmental image and the clarity requirement. Alternatively, an edge detection algorithm (such as Canny) can be used to determine whether there are significant target objects in the environmental image. If there are no significant target objects, the validity information of the environmental image can be determined to be invalid.

[0103] Wherein, when the validity information of the environmental image indicates that the environmental image meets the validity requirement, a prompt can be given to re-take the environmental image or supplement information.

[0104] When the validity information of the environmental image indicates that the environmental image meets the validity requirement, pixel normalization is performed on each environmental image; In the embodiments of the present disclosure, the mean pixel value and the standard deviation of the pixel values of the pixel points in the environmental image are calculated; according to the pixel values of the pixel points in the environmental image, the mean pixel value, and the standard deviation of the pixel values, pixel normalization is performed on each environmental image.

[0105] Exemplarily, the normalization of each environmental image can be: normalizing the pixel values of the pixel points in each environmental image from the original range (such as 0 - 255) to the range [-1, 1] to reduce the influence of illumination changes and device differences on model inference. The formula is as follows:

[0106] Wherein, represents the pixel values of each pixel point in each environmental image; represents the mean pixel value of each pixel point in each environmental image; represents the standard deviation of the pixel values of each pixel point in each environmental image.

[0107] The sizes of the normalized environmental images are unified; In the embodiments of the present disclosure, the size of each normalized environmental image can be adjusted: all the normalized environmental images are uniformly adjusted to a fixed size (such as 224×224 or 512×512) to adapt to the input requirements of the multi-modal large model. For example, when the size of the normalized environmental image is smaller than the fixed size, a bilinear interpolation algorithm can be used for image expansion, which can maintain image details and avoid information loss caused by scaling.

[0108] In the embodiments of the present disclosure, channel alignment can also be performed on each normalized environmental image to ensure that all images are in the RGB three-channel format. For example, if the normalized environmental image is a grayscale image, it can be extended to a three-primary-color image with three channels by, for example, repeating the single-channel primary color value three times.

[0109] Crop the environmental image with unified size to generate the target environmental image.

[0110] Exemplarily, the effective area in the environmental image can be cropped according to the vehicle position. This reduces the interference of irrelevant information and improves the inference speed. For example, the cropping strategy can be: for the top view, retain the ground area within a certain range around the vehicle (such as within 1 meter of the vehicle's exterior). For the panoramic view, retain the fixed area including walls, columns, or other prominent objects. The generated target environmental image after cropping can reduce the interference of redundant information and improve the efficiency of model inference.

[0111] See Figure 4 As shown, the guiding method of the present disclosure is exemplarily illustrated through a flowchart. The multi-view environmental images input at the beginning (such as 4 panoramic views and 1 top view captured by a vehicle) are determined for validity information. For example, the target objects in the images are determined through the Canny edge detection algorithm. When the validity information indicates that the environmental image is valid, preprocessing is performed on the valid image. The preprocessing can include pixel value normalization and size unification.

[0112] Further, after the size unification is completed, the effective area of the image is cropped, and the effective area including the target object in the image can be cropped, and then the target environmental image is generated from the effective area including the target object. Then the target environmental image is input into the target model, for example, input into a multi-modal large model for inference to obtain environmental description information. The format of the environmental description information is analyzed, for example, structured information is extracted for analysis, and the format analysis result representing the correct format of the environmental description information is verified for rationality, for example, the rationality of the feature range is verified, so as to exclude the abnormal environmental description information. Finally, according to the environmental description information excluding the abnormal ones, through operations such as rule splicing, the features of the target object in the environmental image are obtained as guiding information (such as the features of walls, columns, ground, facilities, etc. in the vehicle parking environment), and finally the guiding information can be pushed to the mobile terminal for display.

[0113] Based on the same inventive concept, the embodiments of the present disclosure also provide a guiding device. See Figure 5 As shown, the guiding device includes: A first generation module 510, configured to process the environmental images captured from multiple perspectives to generate a target environmental image; An input module 520, configured to input the target environmental image into a multi-modal large model to obtain environmental description information output by the multi-modal large model; A second generation module 530, configured to generate guidance information according to the environmental description information, where the guidance information is used to describe the characteristics of a target object in the environment.

[0114] Optionally, the environmental image includes a first environmental image captured by a vehicle from multiple perspectives, and the guidance information includes first guidance information, where the first guidance information is used to be pushed to a mobile terminal to display the characteristics of a first target object in the environment where the vehicle is located.

[0115] Optionally, the first generation module 510 is configured to: In response to a vehicle search request from the mobile terminal, process environmental images captured from multiple perspectives to generate a target environmental image, where the vehicle search request is generated according to a user operation on the mobile terminal.

[0116] Optionally, the first environmental image is captured by at least one of the following methods: capturing in response to a user leaving the vehicle, capturing in response to receiving the vehicle search request.

[0117] Optionally, the first target object includes at least one of the following: parking space number, wall, pillar, facility, ground.

[0118] Optionally, the environmental image includes a second environmental image captured by a home device from multiple perspectives, and the guidance information includes second guidance information, where the second guidance information is used to be pushed to a mobile terminal to display the characteristics of a second target object in the environment where the home device is located.

[0119] Optionally, the second target object includes at least one of the following: pet, door and window, home device.

[0120] Optionally, the environmental image includes a third environmental image captured by a robot from multiple perspectives, and the guidance information includes third guidance information, where the third guidance information is used to characterize the characteristics of a third target object in the environment where the robot is located.

[0121] Optionally, the environmental image includes a fourth environmental image captured by a mobile terminal from multiple perspectives, and the guidance information includes fourth guidance information, where the fourth guidance information is used to display the characteristics of a fourth target object in the environment where the mobile terminal is located.

[0122] Optionally, the second generation module 530 is configured to: Generate a feature description of the target object in the environment according to the environmental description information; Verify the feature description according to the preset verification information corresponding to each of the target objects; Generate the guidance information according to the verification results of the feature descriptions.

[0123] Optionally, the preset verification information includes at least one of the following: format information, feature range information.

[0124] Optionally, the input module 520 is configured to: Input the target environment image into a pre-trained target model, and the target model extracts spatial features from the target environment image; Determine the mapping relationship from the target environment image from multiple perspectives to the environment description according to the spatial features; Obtain the environment description information output by the target model according to the mapping relationship.

[0125] Optionally, the first generation module 510 is configured to: Determine the validity information of the environment images captured from multiple perspectives; When the validity information of the environment images indicates that the environment images meet the validity requirements, perform pixel normalization on each of the environment images; Unify the sizes of the normalized environment images; Crop the environment images with unified sizes to generate the target environment images.

[0126] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0127] An embodiment of the present disclosure further provides a vehicle, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of the foregoing embodiments.

[0128] An embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to any one of the foregoing embodiments are implemented.

[0129] An embodiment of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of the foregoing embodiments are implemented.

[0130] Figure 6is a block diagram of a vehicle 600 shown according to an exemplary embodiment. For example, the vehicle 600 can be a hybrid vehicle, or a non - hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 600 can be an autonomous vehicle, a semi - autonomous vehicle, or a non - autonomous vehicle.

[0131] Referring to Figure 6 , the vehicle 600 can include various subsystems. For example, the infotainment system 610, the perception system 620, the decision - making and control system 630, the drive system 640, and the computing platform 650. Among them, the vehicle 600 can also include more or fewer subsystems, and each subsystem can include multiple components. In addition, each subsystem and each component of the vehicle 600 can be interconnected by wired or wireless means.

[0132] In some embodiments, the infotainment system 610 can include a communication system, an entertainment system, and a navigation system, etc.

[0133] The perception system 620 can include several sensors for sensing information about the environment around the vehicle 600. For example, the perception system 620 can include a global positioning system (the global positioning system can be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter - wave radar, ultrasonic radar, and a camera device.

[0134] The decision - making and control system 630 can include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.

[0135] The drive system 640 can include components that provide motive power for the vehicle 600. In one embodiment, the drive system 640 can include an engine, an energy source, a transmission system, and wheels. The engine can be one or a combination of an internal combustion engine, an electric motor, and an air - compression engine. The engine can convert the energy provided by the energy source into mechanical energy.

[0136] Some or all functions of the vehicle 600 are controlled by the computing platform 650. The computing platform 650 can include at least one processor 651 and a memory 652. The processor 651 can execute instructions 653 stored in the memory 652.

[0137] The processor 651 can be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.

[0138] The memory 652 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0139] In addition to the instructions 653, the memory 652 can also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 652 can be used by the computing platform 650.

[0140] In an embodiment of the present disclosure, the processor 651 can execute the instructions 653 to complete all or part of the steps of the above-mentioned guidance method.

[0141] Figure 7 It is a block diagram of a guidance device 700 shown according to an exemplary embodiment. For example, the device 700 can be provided as a server that can receive environment images captured from multiple angles sent by a vehicle and push guidance information to a mobile terminal for display. Refer to Figure 7 As shown, the device 700 includes a processing component 722, which further includes one or more processors, and memory resources represented by a memory 732 for storing instructions executable by the processing component 722, such as application programs. The application programs stored in the memory 732 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 722 is configured to execute instructions to perform the above-mentioned guidance method.

[0142] The device 700 may further include a power component 726 configured to perform power management of the device 700, a wired or wireless network interface 750 configured to connect the device 700 to a network, and an input / output interface 758. The device 700 can operate based on an operating system stored in the memory 732, such as Windows Server TM , Mac OS X TM, Unix TM , Linux TM , FreeBSD TM or the like.

[0143] Other embodiments of the present disclosure will be readily apparent to those skilled in the art in view of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

[0144] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A guiding method, characterized in that including: processing environmental images captured from multiple perspectives to generate a target environmental image; inputting the target environmental image into a pre-trained target model to obtain environmental description information output by the target model; generating guidance information according to the environmental description information, where the guidance information is used to describe the characteristics of a target object in the environment.

2. The method according to claim 1, wherein The environmental image includes a first environmental image captured by a vehicle from multiple perspectives, and the guidance information includes first guidance information, where the first guidance information is used to be pushed to a mobile terminal to display the characteristics of a first target object in the environment where the vehicle is located.

3. The method according to claim 2, wherein The processing of the environmental images captured from multiple perspectives to generate a target environmental image includes: responding to a vehicle search request of the mobile terminal, processing the environmental images captured from multiple perspectives to generate a target environmental image, where the vehicle search request is generated according to a user operation on the mobile terminal.

4. The method according to claim 3, wherein The first environmental image is captured by at least one of the following methods: capturing in response to a user leaving the vehicle, capturing in response to receiving the vehicle search request.

5. The method according to claim 2, wherein The first target object includes at least one of the following: parking space number, wall, pillar, facility, ground.

6. The method according to claim 1, wherein The environmental image includes a second environmental image captured by a home device from multiple perspectives, and the guidance information includes second guidance information, where the second guidance information is used to be pushed to a mobile terminal to display the characteristics of a second target object in the environment where the home device is located.

7. The method according to claim 6, wherein The second target object includes at least one of the following: pet, door and window, home device.

8. The method according to claim 1, wherein The environmental image includes a third environmental image captured by a robot from multiple perspectives, and the guidance information includes third guidance information, where the third guidance information is used to characterize the characteristics of a third target object in the environment where the robot is located.

9. The method according to claim 1, wherein The environmental image includes a fourth environmental image captured by a mobile terminal from multiple perspectives, and the guidance information includes fourth guidance information, where the fourth guidance information is used to display the characteristics of a fourth target object in the environment where the mobile terminal is located.

10. The method according to claim 1, characterized in that, The generating of the guidance information according to the environmental description information includes: generating a feature description of the target object in the environment according to the environmental description information; verifying the feature description according to the preset verification information corresponding to each target object; generating the guidance information according to the verification results of each feature description.

11. The method according to claim 10, characterized in that, The preset verification information includes at least one of the following: format information, feature range information.

12. The method according to claim 1, wherein The inputting of the target environmental image into a pre-trained target model to obtain the environmental description information output by the target model includes: inputting the target environmental image into a pre-trained target model, where the target model extracts spatial features from the target environmental image; determining a mapping relationship from the target environmental image from multiple perspectives to the environmental description according to the spatial features; obtaining the environmental description information output by the target model according to the mapping relationship.

13. The method according to any one of claims 1-12, characterized in that, The processing of the environmental images captured from multiple perspectives to generate a target environmental image includes: determining validity information of the environmental images captured from multiple perspectives; When the validity information of the environmental image indicates that the environmental image meets the validity requirements, perform pixel normalization on each of the environmental images; Perform size unification on the normalized environmental image; Crop the environmental image with unified size to generate the target environmental image.

14. A guiding device, characterized in that, It includes: The first generation module is configured to process the environmental images captured from multiple perspectives to generate the target environmental image; The input module is configured to input the target environmental image into the multi-modal large model to obtain the environmental description information output by the multi-modal large model; The second generation module is configured to generate guidance information according to the environmental description information, and the guidance information is used to describe the characteristics of the target object in the environment.

15. A vehicle, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of claims 1-13.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it realizes the steps of the method according to any one of claims 1-13.

17. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by the processor, it realizes the steps of the method according to any one of claims 1-13.