A parking lot queuing method and device, electronic equipment and storage medium

By acquiring and fusing multimodal information of parking lots and navigation destination association information, and using a deep learning model to rank parking lots, the problem of parking difficulties caused by insufficient parking spaces is solved, and the accuracy and convenience of parking lot recommendations are achieved.

CN115062240BActive Publication Date: 2025-11-07BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210639478.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-07
Publication Date
2025-11-07
Estimated Expiration
2042-06-07

AI Technical Summary

Technical Problem

With the increase in car ownership and the shortage of parking spaces, existing technologies are unable to accurately recommend suitable parking lots, leading to a serious parking problem.

Method used

By acquiring multimodal information of candidate parking lots, including text attributes, image information, and surrounding video information, and combining it with the association information of navigation destinations, a deep machine learning model is used for feature extraction and similarity calculation to achieve accurate ranking of parking lots.

Benefits of technology

It improves the accuracy and timeliness of parking lot sorting, ensuring that users can quickly find suitable parking lots and solving the parking problem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115062240B_ABST
    Figure CN115062240B_ABST
Patent Text Reader

Abstract

The present disclosure provides a parking lot ranking method and device, electronic equipment and storage medium, relates to the field of artificial intelligence, in particular to intelligent transportation technology and deep machine learning technology. The specific implementation scheme comprises: obtaining multi-modal information of a candidate parking lot recalled based on a navigation destination; and ranking the candidate parking lot according to the multi-modal information of the candidate parking lot. The present disclosure ranks the recalled parking lot by using the multi-modal information of the parking lot, thereby improving the accuracy of the parking lot ranking.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of artificial intelligence, in particular to intelligent transportation technology and deep machine learning technology, and specifically to a parking lot ranking method and device, an electronic device, a storage medium and a computer program product. BACKGROUND

[0002] With the rapid economic development and the acceleration of urbanization, people's living standards have been greatly improved. In order to facilitate travel, more and more families purchase cars, resulting in a sharp increase in the number of cars. At the same time, users' parking demand is increasing, and parking resources are severely insufficient, and the problem of parking difficulty is becoming more and more serious. SUMMARY

[0003] The present disclosure provides a parking lot ranking method, device, electronic device, storage medium and computer program product.

[0004] According to an aspect of the present disclosure, a parking lot ranking method is provided, comprising:

[0005] For the candidate parking lot recalled based on the navigation destination, the multi-modal information of the candidate parking lot is obtained;

[0006] According to the multi-modal information of the candidate parking lot, the candidate parking lot is ranked.

[0007] According to an aspect of the present disclosure, a parking lot ranking device is provided, comprising:

[0008] An information acquisition module, for the candidate parking lot recalled based on the navigation destination, acquires the multi-modal information of the candidate parking lot;

[0009] A ranking module, configured to rank the candidate parking lot according to the multi-modal information of the candidate parking lot.

[0010] According to another aspect of the present disclosure, an electronic device is provided, comprising:

[0011] At least one processor; and

[0012] A memory in communication connection with the at least one processor; wherein

[0013] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the parking lot ranking method of any embodiment of the present disclosure.

[0014] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, and the computer instructions are used to make a computer execute the parking lot ranking method of any embodiment of the present disclosure.

[0015] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the parking lot ranking method of any of the embodiments of the present disclosure.

[0016] According to the technology of the present disclosure, the recalled parking lots are ranked by using the multi-modal information of the parking lot, which improves the accuracy of the parking lot ranking.

[0017] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:

[0019] Figure 1 is a flowchart of a parking lot ranking method provided by an embodiment of the present disclosure;

[0020] Figure 2 is a flowchart of another parking lot ranking method provided by an embodiment of the present disclosure;

[0021] Figure 3a is a flowchart of another parking lot ranking method provided by an embodiment of the present disclosure;

[0022] Figure 3b is a structural diagram of a feature extraction model provided by an embodiment of the present disclosure;

[0023] Figure 4 is a flowchart of another parking lot ranking method provided by an embodiment of the present disclosure;

[0024] Figure 5 is a structural diagram of a parking lot ranking device provided by an embodiment of the present disclosure;

[0025] Figure 6 is a block diagram of an electronic device for implementing the parking lot ranking method of an embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, the description below omits the description of well-known functions and structures.

[0027] Figure 1 A flowchart of a parking lot ranking method according to an embodiment of the present disclosure is shown in FIG. 1. The embodiment can be applied to ranking the recalled parking lots in a navigation parking scenario. The method can be executed by a parking lot ranking device, which is implemented in software and / or hardware and integrated on an electronic device, such as a smart terminal, for example, a smartphone or a vehicle-mounted terminal.

[0028] Specifically, referring to FIG. 1, the parking lot ranking method includes the following steps. Figure 1

[0029] S101, for the candidate parking lots recalled based on a navigation destination, obtaining multi-modal information of the candidate parking lots.

[0030] S102, ranking the candidate parking lots according to the multi-modal information of the candidate parking lots.

[0031] The navigation destination can be a navigation end point input by a user in a navigation system. In order to facilitate the user to select a suitable parking lot, the present disclosure automatically recalls a plurality of candidate parking lots for the user after the user inputs the navigation destination. The process of recalling the candidate parking lots based on the navigation destination can be as follows: a plurality of parking lots within a preset range around the navigation destination are recalled as candidate parking lots. Specifically, the coordinate position of the navigation destination can be used to determine the plurality of candidate parking lots within the preset range around the coordinate position. In an example, the coordinate positions of the parking lots are determined in advance, and when the candidate parking lots within the preset range around the navigation destination are determined, the coordinate position of the navigation destination and the coordinate positions of the parking lots determined in advance can be used to query the parking lots within the preset range (for example, 3 kilometers) around the navigation destination, to obtain the plurality of candidate parking lots.

[0032] In order to recommend the most suitable parking lot to the user, the plurality of recalled parking lots need to be ranked. In order to ensure the accuracy of the ranking of the candidate parking lots, the present disclosure proposes a method of ranking based on multi-modal information of the parking lots. First, for the candidate parking lots recalled based on the navigation destination, the multi-modal information of the candidate parking lots is obtained. The multi-modal information of the candidate parking lots includes at least two of text attribute information of the candidate parking lot, picture information of the candidate parking lot, and surrounding video information of the candidate parking lot. For example, the text attribute information of the candidate parking lot includes the coordinate position of the candidate parking lot, the vertical class of the parking lot, the access popularity, the opening attribute of the parking lot, the number of parking spaces of the parking lot, the charging standard of the parking lot, the free policy, etc.; the picture information of the candidate parking lot includes panoramic pictures of the candidate parking lot, which can be collected by the parking lot in advance; and the surrounding video information of the candidate parking lot can be obtained by connecting a surrounding monitoring system of the parking lot.

[0033] ​The collection, storage, use, processing, transmission, provision and disclosure of the multi-modal information in the technical solutions of the present disclosure comply with relevant laws and regulations and do not violate public order and good customs.

[0034] After obtaining the multi-modal information, the information of the candidate parking lots in different modalities has some intersection and complement because the different modalities have different ways of expression and different perspectives of looking at things. Therefore, the candidate parking lots can be sorted based on three modalities of the candidate parking lot at the same time, or sorted based on any two of the three modalities, which is not limited here.

[0035] It should be noted that the candidate parking lots are not sorted based on the information of a modality because the sorting result based on a modality is inaccurate and biased. For example, the candidate parking lots recalled based on a navigation destination include A and B, where the text attributes of the candidate parking lot A are as follows: 600 meters away from the navigation destination, 0.7 of access heat, open to the outside world and 100 idle parking spaces; the text attributes of the candidate parking lot B are as follows: 650 meters away from the navigation destination, 0.7 of access heat, open to the outside world and 90 idle parking spaces. At this time, if the sorting is performed only according to the text attribute information of the candidate parking lot, the candidate parking lot A is placed in front of the candidate parking lot B. At this time, if other modalities of the candidate parking lot are considered, for example, the surrounding monitoring video of the candidate parking lot is considered, it is found that the road around the candidate parking lot A is congested, while the road around the candidate parking lot B is smooth. At this time, the candidate parking lot B is actually more suitable for the user, so the accurate sorting should be that the candidate parking lot B is placed in front of the candidate parking lot A. Therefore, as shown in the above example, the sorting of the recalled candidate parking lot based on multi-modal information is more accurate. In addition, the real-time monitoring video around the candidate parking lot is considered during the sorting, so the timeliness of the sorting result can be ensured.

[0036] In the embodiments of the present disclosure, the candidate parking lots are sorted based on the multi-modal information of the candidate parking lots, which improves the accuracy of the sorting and provides a guarantee for subsequent recommendation of suitable parking lots to the user. In this way, the user can complete convenient parking based on the recommended parking lot.

[0037] Figure 2 is a flowchart of another parking lot sorting method according to an embodiment of the present disclosure. Referring to Figure 2 , the parking lot sorting method is as follows:

[0038] S201, for the candidate parking lots recalled based on a navigation destination, multi-modal information of the candidate parking lots is obtained.

[0039] The multi-modal information of the candidate parking lot includes at least two of text attribute information of the candidate parking lot, picture information of the candidate parking lot, and surrounding video information of the candidate parking lot.

[0040] S202. According to the multi-modal information of the candidate parking lot, the candidate parking lot is sorted in combination with the associated information of the navigation destination.

[0041] In the embodiments of the present disclosure, in order to further improve the accuracy of the sorting result of the candidate parking lot, the associated information of the navigation destination can also be combined in the sorting process. The associated information of the navigation destination can be single-modal information or multi-modal information. If the associated information is single-modal information, for example, text-modal information, it can include attribute information (such as location, name, etc.) and navigation track information of the navigation destination. If it is multi-modal information, the associated information of the navigation destination can include picture information of the navigation destination, surrounding monitoring video information of the navigation destination, etc. in addition to the above information.

[0042] In an optional embodiment, the similarity of the multi-modal information of each candidate parking lot and the associated information of the navigation destination can be calculated, and then the candidate parking lot is sorted according to the similarity calculation result.

[0043] In the embodiments of the present disclosure, the candidate parking lot is sorted according to the multi-modal information of the candidate parking lot in combination with the associated information of the navigation destination, which further improves the accuracy of the sorting of the candidate parking lot.

[0044] Figure 3a is a flowchart of another parking lot sorting method according to the embodiments of the present disclosure. Referring to Figure 3a , the parking lot sorting method is as follows:

[0045] S301. For the candidate parking lot recalled based on the navigation destination, multi-modal information of the candidate parking lot is obtained.

[0046] The multi-modal information of the candidate parking lot includes at least two of text attribute information of the candidate parking lot, picture information of the candidate parking lot, and surrounding video information of the candidate parking lot. In the embodiments of the present disclosure, three modal information is taken as an example for illustration.

[0047] S302. Based on the multi-modal information of the candidate parking lot, multi-modal feature representation of the candidate parking lot is determined.

[0048] In an optional implementation, a parking lot feature extraction model can be pre-constructed, which includes four parts in order to be able to extract features from multi-modal information, namely a text feature extraction model, a feature extraction network in a pre-trained picture target detection model, a feature extraction network in a pre-trained video target detection model, and a multilayer perceptron (MLP, Multilayer Perceptron). The picture target detection model and the video target detection model are optionally trained based on a Transformer model. Based on this, the process of determining the multi-modal feature representation of the candidate parking lot based on the multi-modal information of the candidate parking lot is as follows: based on the constructed text feature extraction model, text feature representation is extracted from the text attribute information of the candidate parking lot; wherein the text feature model includes a word embedding model and a feature extractor based on a convolutional neural network; the feature extraction network in the pre-trained picture target detection model is used to extract the first visual feature representation from the picture information of the candidate parking lot; the feature extraction network in the pre-trained video target detection model is used to extract the second visual feature representation from the surrounding video information of the candidate parking lot; the text feature representation, the first visual feature representation and the second visual feature representation are fused by the multilayer perceptron to obtain the multi-modal feature representation of the candidate parking lot, wherein the multi-modal feature representation can exist in the form of a feature matrix. In this way, through the constructed parking lot feature extraction model, the multi-modal feature representation fused with multi-modal information can be quickly obtained.

[0049] S303, determine the feature representation of the navigation destination based on the associated information of the navigation destination.

[0050] The association information includes attribute information of the navigation destination and navigation track information. To determine the feature representation of the navigation destination, a point of interest feature extraction model can be constructed, which can include a text feature model and a multi-layer perceptron, and the text feature model includes a word embedding model and a feature extractor based on a convolutional neural network. Based on the association information of the navigation destination, the feature representation of the navigation destination is determined, including: based on the constructed text feature extraction model, text feature representation is extracted from the attribute information and the navigation track information of the navigation destination; specifically, the attribute information is encoded based on the word embedding model, and the encoding result is input into the feature extractor for feature extraction; the navigation track information is encoded based on the word embedding model, and the encoding result is input into the feature extractor for feature extraction; the two feature extraction results are fused through the multi-layer perceptron to obtain the feature representation of the navigation destination. It should be noted that if the association information of the navigation destination is also multi-modal information, for example, in addition to the above information, it also includes pictures, videos, etc. of the navigation destination, the constructed point of interest feature extraction model should also include the feature extraction network in the pre-trained picture target detection model and the feature extraction network in the pre-trained video target detection model. For details of the feature extraction process, please refer to the extraction process of the multi-modal feature representation of the candidate parking lot.

[0051] It should be noted that the structures of the parking lot feature extraction model and the point of interest feature extraction model can be referred to Figure 3b .

[0052] S304, based on the multi-modal feature representation of the candidate parking lot and the feature representation of the navigation destination, the similarity between the navigation destination and the candidate parking lot is determined.

[0053] S305, the candidate parking lot is sorted according to the similarity.

[0054] In an optional embodiment, the cosine angle or distance between the multi-modal feature representation of the candidate parking lot and the feature representation of the navigation destination can be calculated, and the calculation result is taken as the value of the similarity between the navigation destination and the candidate parking lot. In this way, after calculating the similarity between each candidate parking lot and the navigation destination respectively, the candidate parking lots can be sorted according to the size of the similarity value.

[0055] In the embodiments of the present disclosure, the multi-modal information of the candidate parking lot is fused by calculating the multi-modal feature representation of the candidate parking lot, so that the obtained parking lot feature is more rich, which is the basis for accurate sorting of the parking lot. Further, the similarity between the multi-modal feature representation of the candidate parking lot and the feature representation of the navigation destination is calculated, and the candidate parking lot is sorted according to the similarity. In this way, the candidate parking lot most similar to the navigation destination can be placed in the front, which can ensure the accuracy of the sorting.

[0056] Figure 4 is a flowchart of another parking lot ranking method according to an embodiment of the present disclosure. Referring to Figure 4 , the parking lot ranking method is as follows:

[0057] S401, for the candidate parking lot recalled based on the navigation destination, obtaining multi-modal information of the candidate parking lot.

[0058] Among them, the multi-modal information of the candidate parking lot includes at least two of the candidate parking lot text attribute information, the candidate parking lot picture information and the candidate parking lot surrounding video information.

[0059] S402, ranking the candidate parking lot according to the multi-modal information of the candidate parking lot.

[0060] The specific implementation process of steps S401-S402 can be referred to the above-mentioned embodiments, which will not be repeated here.

[0061] S403, re-ranking the candidate parking lot after ranking according to the user set navigation starting point, and recommending the candidate parking lot after re-ranking to the user.

[0062] In the embodiment of the present disclosure, the candidate parking lot after ranking is re-ranked according to the user's navigation starting point, in order to find the relevant parking lot with the shortest route, and then the candidate parking lot after re-ranking is recommended to the user, so that the user can quickly determine the parking lot suitable for himself.

[0063] Figure 5 is a structural schematic diagram of a parking lot ranking device according to an embodiment of the present disclosure, which can be applied to the case of ranking the recalled parking lot in the navigation parking scene. Referring to Figure 5 , it comprises:

[0064] The information acquisition module 501 is used for the user to obtain multi-modal information of the candidate parking lot recalled based on the navigation destination;

[0065] The ranking module 502 is used for ranking the candidate parking lot according to the multi-modal information of the candidate parking lot.

[0066] On the basis of the above-mentioned embodiments, the ranking module can optionally comprise:

[0067] The ranking unit is used for ranking the candidate parking lot according to the multi-modal information of the candidate parking lot in combination with the associated information of the navigation destination.

[0068] On the basis of the above-mentioned embodiments, the ranking unit can further comprise:

[0069] The first feature extraction subunit is configured to determine a multi-modal feature representation of the candidate parking lot based on multi-modal information of the candidate parking lot.

[0070] The second feature extraction subunit is configured to determine a feature representation of the navigation destination based on the associated information of the navigation destination.

[0071] The similarity calculation subunit is configured to determine a similarity between the navigation destination and the candidate parking lot based on the multi-modal feature representation of the candidate parking lot and the feature representation of the navigation destination.

[0072] The ranking subunit is configured to rank the candidate parking lots according to the similarity.

[0073] In the above embodiment, optionally, the multi-modal information of the candidate parking lot includes at least two of text attribute information of the candidate parking lot, picture information of the candidate parking lot, and surrounding video information of the candidate parking lot.

[0074] In the above embodiment, optionally, the first feature extraction subunit is further configured to:

[0075] extract a text feature representation from the text attribute information of the candidate parking lot based on the constructed text feature extraction model; wherein the text feature model includes a word embedding model and a feature extractor based on a convolutional neural network.

[0076] extract a first visual feature representation from the picture information of the candidate parking lot by using a feature extraction network in a pre-trained picture target detection model;

[0077] extract a second visual feature representation from the surrounding video information of the candidate parking lot by using a feature extraction network in a pre-trained video target detection model.

[0078] fuse the text feature representation, the first visual feature representation, and the second visual feature representation to obtain the multi-modal feature representation of the candidate parking lot.

[0079] In the above embodiment, optionally, the associated information includes attribute information and navigation trajectory information of the navigation destination.

[0080] The second feature extraction subunit is further configured to:

[0081] extract a text feature representation from the attribute information and the navigation trajectory information of the navigation destination based on the constructed text feature extraction model; wherein the text feature model includes a word embedding model and a feature extractor based on a convolutional neural network.

[0082] In the above embodiment, optionally, the method further includes:

[0083] The reordering module is configured to reorder the sorted candidate parking lots according to a navigation starting point set by the user, and recommend the reordered candidate parking lots to the user.

[0084] The parking lot sorting apparatus provided in the embodiments of the present disclosure can execute the parking lot sorting method provided in any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method. The contents not described in detail in the embodiments can be referred to the description in any of the method embodiments of the present disclosure.

[0085] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0086] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0087] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0088] As shown in Figure 6 The device 600 includes a computing unit 601 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0089] Various components in the device 600 are connected to the I / O interface 605, including an input unit 606 such as a keyboard, a mouse, etc., an output unit 607 such as various types of displays, a speaker, etc., a storage unit 608 such as a magnetic disk, an optical disk, etc., and a communication unit 609 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0090] The computing unit 601 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs various methods and processes described above, such as the parking lot ordering method. For example, in some embodiments, the parking lot ordering method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded onto the RAM 603 and executed by the computing unit 601, one or more steps of the parking lot ordering method described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the parking lot ordering method by any other appropriate means, such as by means of firmware.

[0091] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0092] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0093] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0094] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0095] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0096] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0097] It should be understood that the various forms of flow shown above can be used to reorder, add, or remove steps. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.

[0098] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A parking lot ranking method, comprising: obtaining multi-modal information of a candidate parking lot recalled based on a navigation destination; wherein the multi-modal information of the candidate parking lot comprises candidate parking lot text attribute information, candidate parking lot picture information, and candidate parking lot surrounding video information; ranking the candidate parking lot according to the multi-modal information of the candidate parking lot in combination with associated information of the navigation destination; wherein the associated information of the navigation destination is multi-modal information comprising a picture of the navigation destination, surrounding monitoring video information of the navigation destination, attribute information of the navigation destination, and navigation trajectory information; re-ranking the ranked candidate parking lot according to a user-set navigation starting point, and recommending the re-ranked candidate parking lot to the user.

2. The method of claim 1, wherein, ranking the candidate parking lot according to the multi-modal information of the candidate parking lot in combination with the associated information of the navigation destination comprises: determining multi-modal feature representation of the candidate parking lot based on the multi-modal information of the candidate parking lot; determining feature representation of the navigation destination based on the associated information of the navigation destination; determining similarity between the navigation destination and the candidate parking lot based on the multi-modal feature representation of the candidate parking lot and the feature representation of the navigation destination; ranking the candidate parking lot according to the similarity.

3. The method of claim 2, wherein, determining the multi-modal feature representation of the candidate parking lot based on the multi-modal information of the candidate parking lot comprises: extracting text feature representation from the candidate parking lot text attribute information based on a constructed text feature extraction model; wherein the text feature extraction model comprises a word embedding model and a feature extractor based on a convolutional neural network; extracting first visual feature representation from the candidate parking lot picture information using a feature extraction network in a pre-trained picture target detection model; extracting second visual feature representation from the candidate parking lot surrounding video information using a feature extraction network in a pre-trained video target detection model; fusing the text feature representation, the first visual feature representation, and the second visual feature representation to obtain the multi-modal feature representation of the candidate parking lot.

4. The method of claim 2, wherein, the associated information comprises attribute information and navigation trajectory information of the navigation destination; determining the feature representation of the navigation destination based on the associated information of the navigation destination comprises: extracting text feature representation from the attribute information and navigation trajectory information of the navigation destination based on a constructed text feature extraction model; wherein the text feature extraction model comprises a word embedding model and a feature extractor based on a convolutional neural network.

5. A parking lot ranking device, comprising: an information obtaining module, which obtains multi-modal information of a candidate parking lot recalled based on a navigation destination in response to a user; wherein the multi-modal information of the candidate parking lot comprises candidate parking lot text attribute information, candidate parking lot picture information, and candidate parking lot surrounding video information; a ranking module, which ranks the candidate parking lot according to the multi-modal information of the candidate parking lot. The reordering module is configured to reorder the sorted candidate parking lots according to a navigation starting point set by a user, and recommend the reordered candidate parking lots to the user. The sorting module comprises: The sorting unit is configured to sort the candidate parking lots according to the multi-modal information of the candidate parking lots and the associated information of the navigation destination; the associated information of the navigation destination is multi-modal information, including a picture of the navigation destination, surrounding monitoring video information of the navigation destination, attribute information of the navigation destination, and navigation trajectory information.

6. The apparatus of claim 5, wherein, The sorting unit further comprises: The first feature extraction subunit is configured to determine the multi-modal feature representation of the candidate parking lot based on the multi-modal information of the candidate parking lot; The second feature extraction subunit is configured to determine the feature representation of the navigation destination based on the associated information of the navigation destination; The similarity calculation subunit is configured to determine the similarity between the navigation destination and the candidate parking lot based on the multi-modal feature representation of the candidate parking lot and the feature representation of the navigation destination; The sorting subunit is configured to sort the candidate parking lots according to the similarity.

7. The apparatus of claim 6, wherein, The first feature extraction subunit is further configured to: extract a text feature representation from the text attribute information of the candidate parking lot based on a constructed text feature extraction model; the text feature extraction model comprises a word embedding model and a feature extractor based on a convolutional neural network; extract a first visual feature representation from the picture information of the candidate parking lot by using a feature extraction network in a pre-trained picture target detection model; extract a second visual feature representation from the surrounding video information of the candidate parking lot by using a feature extraction network in a pre-trained video target detection model; fuse the text feature representation, the first visual feature representation, and the second visual feature representation to obtain the multi-modal feature representation of the candidate parking lot.

8. The apparatus of claim 6, wherein, The associated information comprises attribute information and navigation trajectory information of the navigation destination. The second feature extraction subunit is further configured to: extract a text feature representation from the attribute information and the navigation trajectory information of the navigation destination based on a constructed text feature extraction model; the text feature extraction model comprises a word embedding model and a feature extractor based on a convolutional neural network. 9.An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the parking lot sorting method of any one of claims 1-4.

10. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable a computer to perform the parking lot sorting method of any one of claims 1-4. 11.A computer program product comprising a computer program which, when executed by a processor, implements the parking lot sorting method of any one of claims 1-4.

Citation Information

Patent Citations

  • Parking lot attribute prediction model training method, and parking lot recommendation method and device

    CN113257030A

  • Parking lot recommendation method and device, electronic equipment and storage medium

    CN113849746A