Multi-modal data query method and device, equipment and medium

By obtaining the main query data modality of the query task in the tourism system, extracting location-related information and indexing it into the spatial grid, the problem of low data management and retrieval efficiency in the existing technology is solved, and efficient multimodal data query and recommendation are achieved.

CN120804374APending Publication Date: 2025-10-17WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510824337.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing tourism systems find it difficult to effectively manage and retrieve multimodal travel data, which affects tourists' experience. In addition, large language models in existing technologies face challenges in data integration and query efficiency improvement.

Method used

By obtaining the main query data modality of the query task, extracting location-related information, and indexing the corresponding spatial grid based on this information, other modal data related to the main query data are obtained for aggregate sorting and structured recommendation.

Benefits of technology

It improves the efficiency of data retrieval, can quickly find the multimodal data most relevant to user requirements, and improves query efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804374A_ABST
    Figure CN120804374A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode data query method and device, equipment and a medium, and the method comprises the steps: obtaining a query task, and recognizing a main query data mode of the query task; executing the query task based on a main query data mode to obtain main query data; extracting position associated information in the main query data; indexing to a corresponding space grid based on the position association information, and obtaining query data of other modals related to the main query data in the space grid; the space grid corresponds to an actual geographic area, and multi-modal data in the corresponding geographic area is stored; and aggregating and sorting the obtained main query data and the query data of other modalities related to the main query data, and forming structured recommendation content to be returned to the user. According to the invention, the query efficiency of the user during query can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data indexing, and in particular to a multi-modal data query method, device, equipment and medium. BACKGROUND

[0002] The rapid development of the tourism industry has brought a series of problems to the city's tourism system. Among them, the more prominent problem is that the existing tourism system obtains data from complex sources, with huge volume and diverse modalities, which is difficult to effectively obtain and manage. The travel system is difficult to provide multi-modal travel data that travelers want, which has a great negative impact on the travel experience of tourists. Moreover, when tourists go out, they will choose tourist destinations and determine travel arrangements according to personal preferences, travel time budget, etc., which also puts higher requirements on the retrieval ability of the tourism system. In the prior art, artificial intelligence, especially large language models, have made significant progress in understanding natural language and generating text content, providing new technical means to solve the above problems. However, how to efficiently integrate various data sources and improve the query efficiency of users when querying is still a technical challenge. SUMMARY

[0003] The present application provides a multi-modal data query method, device, equipment and medium, aiming to provide multi-modal data that meets personal preferences for travelers and improve the query efficiency of users when querying.

[0004] According to an aspect of the present application, a multi-modal data query method is provided, comprising: obtaining a query task and identifying the main query data modality of the query task; executing the query task based on the main query data modality to obtain main query data; extracting location-related information in the main query data; indexing the location-related information to the corresponding spatial grid and obtaining query data of other modalities related to the main query data in the spatial grid; the spatial grid corresponds to an actual geographic area and stores multi-modal data within the corresponding geographic area; aggregating and sorting the obtained main query data and query data of other modalities related to the main query data, and forming structured recommendation content to return to the user.

[0005] Optionally, the identification of the main query data modality of the query task comprises: determining the main query data modality in the query task based on a pre-defined user intent and modality mapping relationship.

[0006] Optionally, the location-related information at least includes address, coordinate.

[0007] Optionally, the multi-modal data stored in the spatial grid includes at least vector data, graph model data, and document data; The vector data includes at least travel guides, user reviews, and question-and-answer dialogue record converted vector data; The graph model data includes at least topological graph data converted from city subway networks, highway bus route maps, and aviation / high-speed rail transportation hubs; The document data includes static structured entities; the static structured entities include at least restaurants, scenic spots, accommodations, and shopping districts.

[0008] Optionally, the method further comprises: vectorizing travel guides, user reviews, and question-and-answer dialogue records obtained in a geographical area through a natural language model, and storing the data in a vector database corresponding to the spatial grid; constructing a traffic topological graph from city subway networks, highway bus route maps, and aviation / high-speed rail transportation hubs in a geographical area, and storing the data in a graph model database corresponding to the spatial grid; storing static structured entities in a document database corresponding to the spatial grid in a geographical area. Optionally, the method further comprises: indexing the location association information to the spatial grid corresponding to the geographical area, and obtaining other modal query data related to the main query data in the corresponding spatial grid and / or surrounding spatial grids.

[0009] Optionally, the method further comprises: extracting features of the main query data and other modal query data related to the main query data to obtain feature vectors; performing feature fusion on the feature vectors of different modal data to obtain fusion vectors; calculating the similarity between the fusion vectors, and performing recommendation sorting on different modal data based on the similarity; structuring and displaying the recommendation sorting results.

[0010] According to another aspect of the present application, a multi-modal data query device is provided, comprising: a main modal recognition unit configured to obtain a query task and recognize the main query data modal of the query task; a data query unit configured to execute the query task based on the main query data modality and obtain main query data; an information extraction unit configured to extract location-related information in the main query data; a data retrieval unit configured to index to a corresponding spatial grid based on the location-related information and obtain query data of other modalities related to the main query data in the spatial grid; the spatial grid corresponds to an actual geographic area and stores multi-modal data in the corresponding geographic area; a result output unit configured to aggregate and sort the obtained main query data and query data of other modalities related to the main query data, and form structured recommendation content to return to the user.

[0011] According to another aspect of the present application, an electronic device is provided, which comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the multi-modal data query method according to any one of the embodiments of the present application.

[0012] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the multi-modal data query method according to any one of the embodiments of the present application when executed by the processor.

[0013] The technical solution of the embodiments of the present application finds the main query data corresponding to the query task by obtaining the main query data modality, extracts the location-related information in the main query data, indexes to the corresponding spatial grid based on the location-related information, and obtains query data of other modalities related to the main query data in the spatial grid, aggregates and sorts the obtained main query data and query data of other modalities related to the main query data, and forms structured recommendation content to return to the user, thereby finding multi-modal data most relevant to the user's required query task. In addition, by making the spatial grid correspond to an actual geographic area and storing multi-modal data in the corresponding geographic area, global search is no longer needed when retrieving information, and only the required query data of other modalities needs to be found in the spatial grid corresponding to the geographic area, thereby greatly improving the efficiency of data retrieval.

[0014] It is to be understood that the embodiments described herein are merely exemplary of the application and that a myriad of modifications, both as to the nature and number of elements within the execution of the application and as to the modes of execution thereof, can be made by those skilled in the art, and that in its broader aspect, the application is directed to those who have, or promise to have the skills in the art. Any and all such modifications are intended to be included within the scope of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.

[0016] Figure 1 is a flow chart of a multi-modal data query method according to an embodiment of the present application; Figure 2 is a flow chart of a multi-modal data query method according to an embodiment of the present application; Figure 3 is a schematic diagram of a multi-modal grid space index module architecture in an embodiment of the present application; Figure 4 is a flow chart of a multi-modal data query device according to an embodiment of the present application; Figure 5 is a structural schematic diagram of an electronic device implementing a multi-modal data query method according to an embodiment of the present application. DETAILED DESCRIPTION

[0017] In order to make the technical personnel in the art better understand the present application scheme, the following will combine the drawings in the embodiments of the present application, and the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort should be within the scope of the present application.

[0018] It should be noted that the terms "first", "second", and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0019] Embodiment one Figure 1 A flowchart of a multi-modal data query method is provided for the embodiment one of the present application. As shown in the figure, the method comprises: Figure 1 S101, acquiring a query task and identifying a main query data modality of the query task.

[0020] The query task is determined by a search sentence input by a user, and the data modality can include vector type data, graph model data and document type data, etc. The main query data modality can be one of the vector type data, the graph model data and the document type data. In this embodiment, the main query data modality corresponding to the query task can be identified. For example, the query task is: “How to get to a station”, it can be identified that the main query data of the query task is “station”, and then it can be further identified that the data modality of the “station” is the graph model data, and through searching the graph model data, the specific station information can be found.

[0021] S102, executing the query task based on the main query data modality to acquire main query data.

[0022] After the main query data modality is determined, the query task can be executed, and the corresponding main query data in the database corresponding to the main query data modality can be searched, i.e. for the query task: “How to get to a station”, the specific station information can be found in the graph model database. S103, extracting location association information in the main query data.

[0023] The main query data can include location association information, and the location association information at least includes coordinates and places, and can also include information related to the coordinates and places, such as a station, a scenic spot, an airport and a railway station, and then the location information can be extracted from the station, the scenic spot, the airport and the railway station.

[0024] S104, indexing to a corresponding spatial grid based on the location association information, and acquiring query data of other modalities related to the main query data in the spatial grid; the spatial grid corresponds to an actual geographical area, and stores multi-modal data in the corresponding geographical area.

[0025] In this embodiment, the spatial grid can correspond to an actual geographical area, for example, a city can be divided into multiple rectangular areas in a grid form, and each rectangular area can be associated with a spatial grid. The multi-modal data in the corresponding geographical area of the spatial grid is stored in each spatial grid.

[0026] ​When executing a query task, after finding the main query data, you can query data of other modalities related to the main query data within the spatial grid where the main query data resides. For example, after querying site data in the graph model data, you can query document-type data within the spatial grid where the site resides, such as tourist attraction data, restaurant data, etc.

[0027] S105: Aggregate and sort the acquired main query data and query data of other modalities related to the main query data, and form structured recommendation content to return to the user.

[0028] In this embodiment, the main query data and query data of other modalities related to the main query data can be aggregated and sorted. For example, vector data, graph model data, and document data can be mapped to a unified feature space to capture multidimensional associations and perform cross-modal fusion of multimodal data. The fused features can then be comprehensively sorted and displayed in a structured manner based on the sorting results.

[0029] The technical solution of the embodiment of the present invention obtains the main query data modality of the query task, finds the corresponding main query data, extracts the location association information in the main query data, and then indexes the corresponding spatial grid based on the location association information, and obtains query data of other modalities related to the main query data in the spatial grid, aggregates and sorts the obtained main query data and query data of other modalities related to the main query data, and forms structured recommendation content and returns it to the user, thereby finding the multimodal data most relevant to the search task required by the user; in addition, by making the spatial grid correspond to an actual geographical area and storing the multimodal data in the corresponding geographical area, when retrieving information, it is no longer necessary to perform a global search, and it is only necessary to find the required query data of other modalities in the spatial grid corresponding to the geographical area, thereby greatly improving the efficiency of data retrieval.

[0030] Example 2 Figure 2 This is a flow chart of a multimodal data query method provided by the second embodiment of the present invention. Figure 1 As shown, the method includes: S201: Acquire a query task, and determine a main query data modality in the query task based on a predefined user intent and modality mapping relationship; The predefined user intent and modal mapping relationship can be: urban subway network, highway bus route map, aviation / high-speed rail transportation hub usually corresponds to graph model data; travel guide, user review, question and answer dialogue record, etc. usually correspond to vector type data; and static structured entities such as restaurants, scenic spots, accommodations, shopping districts usually correspond to document type data. Based on semantic retrieval, it is determined whether the user intent is vector type data, based on keyword retrieval, it is determined whether the user intent is document type data, and based on graph path retrieval, it is determined whether the user intent is graph model data. Based on this correspondence, the main query data modality in the user's query task can be determined.

[0031] S202, performing the query task based on the main query data modality to obtain main query data; After determining the main query data modality, the query task can be performed, such as determining that the query task is: "how to go to a station", it can be identified that the main query data of the query task is "station", then it can be further identified that the data modality of "station" is graph model data, and by performing path search in the graph model data, the station can be found.

[0032] S203, extracting location associated information in the main query data; The main query data can include location associated information, and the location associated information at least includes coordinates and places, and can also include information related to the coordinates and places, such as a station, a scenic spot, an airport, and a railway station. Then the location information can be extracted from the station, the scenic spot, the airport, and the railway station. In a specific embodiment, the location associated information at least includes an address and coordinates.

[0033] S204, indexing to a corresponding space grid of a local geographical area based on the location associated information, and obtaining query data of other modalities related to the main query data in the corresponding space grid and / or surrounding space grid; In this embodiment, the space grid can correspond to an actual geographical area, for example, a city can be divided into multiple rectangular areas in a grid form, and each rectangular area can be associated with a space grid. Each space grid stores multi-modal data in the geographical area corresponding to the space grid.

[0034] In performing the query task, after the main query data is queried, other modal data related to the main query data in the space grid where the main query data is located can be queried. For example, when the station data in the graph model data needs to be queried, the document type data in the space grid where the station is located can be queried, such as scenic spot data, restaurant data, etc.

[0035] In an embodiment, the main query data can also query the surrounding space grids of the corresponding space grid, for example, when searching for food near a location, not only the food in the space grid where the location is located, but also the food in the adjacent space grids of the space grid where the location is located.

[0036] In an embodiment, the method for constructing a multi-modal database comprises: The travel guides, user reviews, and question-and-answer conversation records obtained in a geographical area are vectorized by a natural language model and stored in the vector database corresponding to the space grid. For example, the obtained natural language content can be vectorized by a Transformer-based model (such as BERT or Sentence-BERT) and stored in a vector database (such as FAISS or Milvus). The established vector database can support semantic approximate search.

[0037] A traffic topology graph is constructed based on the urban subway network, road and bus route map, and aviation / high-speed rail transportation hub in a geographical area, and a graph model database corresponding to the space grid is stored. For example, the obtained urban subway network, road and bus route map, and aviation / high-speed rail transportation hub data can be used to construct a traffic topology graph, with stations as nodes and running paths as edges. The constructed graph model database can support multiple graph search queries such as shortest path and least transfer.

[0038] Static structured entities in a geographical area are stored in the document-type database corresponding to the space grid. For example, the obtained static structured entities such as restaurants, scenic spots, accommodations, and shopping districts can be used to construct a document-type database. The constructed document-type database can support location aggregation and grid-level search based on the keyword search of the Elasticsearch-based inverted engine and the spatial partitioning technology such as GeoHash or QuadTree.

[0039] It should be noted that in the embodiment, GeoHash, Quadtree and other spatial division techniques can be used for high-precision grid processing of the city map, each grid unit is used as a spatial grid, and the spatial grid is associated with three types of data index pointers: vector data, document data and graph model node data, to realize efficient query scheduling under spatial restriction. In the index organization mode, the vector data uses GeoHash grid number and semantic vector as a joint index key and is stored in a vector database such as Milvus, which is used for semantic-driven deep interest retrieval; the document data uses entity name and spatial grid number to construct an inverted index and is deployed in a full-text search system such as Elasticsearch, which supports structured and keyword search; and the graph model data constructs a traffic network based on an adjacency list structure and labels each node with a grid number, to realize path-based topological calculation and station aggregation recommendation.

[0040] S205, aggregating and sorting the obtained main query data and other modal query data related to the main query data, and forming structured recommendation content to return to the user.

[0041] In the embodiment, based on the retrieved data of different modalities, the data of different modalities can be aggregated and sorted, and the structured recommendation content is returned to the user.

[0042] The method for aggregating and sorting the multi-modal data and forming the structured recommendation content in the embodiment is as follows: S2051, respectively extracting features of the main query data and other modal query data related to the main query data to obtain feature vectors.

[0043] For the extracted vector data, NLP techniques (such as BERT and TF-IDF) can be used to convert text into dense vectors (such as multi-dimensional semantic vectors) to capture semantic features (such as “food recommendation”, “scenic spot check-in” and “convenient transportation”).

[0044] For the document data, the data is stored in an inverted index engine such as Elasticsearch, an inverted index is established based on keywords (such as “hot pot” and “5A scenic spot”), and fast retrieval is supported. The geographic location (latitude and longitude) is converted into a grid code by GeoHash or QuadTree, and spatial range query (such as “scenic spots within 3 kilometers around the current location of the user”) is supported.

[0045] For graph model data, algorithms such as GraphSAGE and DeepWalk can be used to convert the graph structure into low-dimensional vectors (such as multi-dimensional node embedding vectors) to capture the spatial correlation and accessibility of nodes (such as the importance of “subway transfer hubs”).

[0046] S2052, the feature vectors of different modal data are fused to obtain a fusion vector.

[0047] The search results of each modality can be converted into a unified feature vector, such as a text semantic vector (from vector data), an entity attribute feature (such as a score, a price, a distance, from document data), and a traffic accessibility vector (such as bus travel time to the user's location, from graph model data).

[0048] S2053, the similarity between the fusion vectors is calculated, and different modal data is recommended and sorted based on the similarity.

[0049] The similarity between the fusion vectors can be calculated, and different modal data can be recommended and sorted based on the similarity. Specifically, the fused vectors can be stored in a vector database (such as Milvus, Faiss), an approximate nearest neighbor (ANN) index can be established, and fast similarity retrieval can be supported. Elasticsearch is used for structured attributes (such as categories, location ranges) to build an inverted index to achieve accurate queries under filter conditions. Through the inverted index, entities that do not meet the conditions are filtered (such as filtering entities of the “scenic spot” category and located in the current grid); for the filtered results, the Top-K candidates (such as Top 100) most similar to the user query vector are queried through the vector database. A ranking model (such as LambdaMART, LightGBM) is introduced to reorder the results in combination with artificial features (such as user historical clicks, entity ratings) to improve relevance.

[0050] S2054, the recommended sorting results are structured and displayed.

[0051] In this embodiment, the results can be divided into “food and beverage”, “scenic spot”, “accommodation”, “shopping” and other blocks, and Top N results are displayed in each block according to the sorting score; or, based on the GeoHash grid division results, dense areas are displayed in the form of a heat map or a grid block on the map (such as “3 recommended restaurants within 500 meters around People's Square”).

[0052] In a specific application example, the method for querying multi-modal data can be as shown in Figure 3 Figure 3 ​In the embodiment, a city is divided into multiple locations, each location can be divided into multiple geographic areas, each area corresponds to a spatial grid, and each spatial grid can index data of three modalities: vector data, graph model data and document data.

[0053] When the main query data modality in the query task is determined and the main query data retrieved in the database corresponding to the modality, the location association information in the main query data can be extracted, and the location association information is indexed to the spatial grid corresponding to the geographic area, and the query data of other modalities related to the main query data in the corresponding spatial grid and / or surrounding spatial grid is obtained. Figure 3 In the embodiment, after the document data is retrieved, the guide (vector data) or traffic data (graph model data) can be retrieved according to the location of the spatial grid, or after the vector data is retrieved, the restaurant (document data) or traffic data (graph model data) can be retrieved according to the location of the spatial grid; or after the graph model data is retrieved, the restaurant (document data) or review data (vector data) can be retrieved according to the location of the spatial grid.

[0054] Specifically, when a traffic transfer scheme for a certain place is requested, the system performs a shortest path / least transfer query in the traffic graph model data; the coordinates of the passing stations are found according to the query result, and then the corresponding spatial grid is indexed, and the document data (such as scenic spots, restaurants) in the corresponding grid or surrounding grid in the spatial grid is searched; finally, the system can sort and recommend the searched multi-modal data to the user according to categories, scores, and heat.

[0055] The technical scheme of the embodiment of the application acquires the main query data modality of the query task, finds the corresponding main query data, extracts the location association information in the main query data, and then indexes the corresponding spatial grid based on the location association information, and acquires the query data of other modalities related to the main query data in the spatial grid, aggregates and sorts the acquired main query data and the query data of other modalities related to the main query data, and forms structured recommendation content to return to the user, so as to find the multi-modal data most related to the user's query task; in addition, the spatial grid corresponds to an actual geographic area, and the multi-modal data in the corresponding geographic area is stored, so that when the information is retrieved, global search is no longer needed, and only the required other modalities query data needs to be searched in the spatial grid corresponding to the geographic area, thereby greatly improving the efficiency of data retrieval.

[0056] Embodiment three Figure 4 A structural schematic diagram of a multi-modal data query device provided by the embodiment three of the application is shown in FIG. 1. Figure 4 As shown in the figure, the device comprises: The main modal recognition unit 401 is configured to acquire a query task and recognize a main query data mode of the query task. The data query unit 402 is configured to execute the query task based on the main query data mode and acquire main query data. The information extraction unit 403 is configured to extract location-related information in the main query data. The data query unit 404 is configured to index to corresponding spatial grids based on the location-related information and acquire query data of other modes related to the main query data in the spatial grids; the spatial grids correspond to an actual geographic area and store multi-modal data in the corresponding geographic area. The result output unit 405 is configured to aggregate and sort the acquired main query data and query data of other modes related to the main query data, and form structured recommendation content to return to a user.

[0057] The multi-modal data query device provided in the embodiments of the present application can execute the multi-modal data query device provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0058] Embodiment four Figure 5 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.

[0059] As Figure 5As shown, the electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., communicatively connected to the at least one processor 11, where the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0060] Various components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, a speaker, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0061] The processor 11 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as a resource scheduling method for message data processing.

[0062] In some embodiments, a resource scheduling method for message data processing can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the resource scheduling method for message data processing described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform a resource scheduling method for message data processing by any other appropriate means, such as by means of firmware.

[0063] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip systems (SOCs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0064] Computer programs used to implement the processes of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program

[0065] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0066] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0067] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0068] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0069] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.

[0070] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the disclosure. Accordingly, the disclosure is not limited to the specific embodiments described above, but only by the scope of the appended claims.

Claims

1. A multimodal data query method, characterized in that: include: Obtaining a query task and identifying a primary query data modality of the query task; Execute the query task based on the main query data modality to obtain the main query data; extracting location-related information from the main query data; Indexing to a corresponding spatial grid based on the location association information, and obtaining query data of other modalities related to the main query data in the spatial grid; The spatial grid corresponds to an actual geographical area and stores multimodal data within the corresponding geographical area; The obtained main query data and query data of other modes related to the main query data are aggregated and sorted, and structured recommendation content is formed and returned to the user.

2. The multimodal data query method according to claim 1, characterized in that: The identifying of the main query data modality of the query task includes: The main query data modality in the query task is determined based on a predefined user intention and modality mapping relationship.

3. The multimodal data query method according to claim 1, characterized in that: The location-related information includes at least an address and coordinates.

4. The multimodal data query method according to claim 3, characterized in that: The multimodal data stored in the spatial grid includes at least vector data, graph model data and document data; The vector data at least includes travel guides, user comments, and vector data converted from question-and-answer dialogue records; The graph model data at least includes topological graph data converted from urban subway networks, highway bus route maps, and aviation / high-speed rail transportation hubs; The document-type data includes static structured entities; the static structured entities include at least restaurants, scenic spots, accommodations, and shopping districts.

5. The method for storing multimodal data according to claim 4, wherein: Also includes: Vectorize travel guides, user reviews, and question-and-answer conversation records obtained in a geographic area using a natural language model, and store them in the vector database corresponding to the spatial grid; Constructing a transportation topology map based on the urban subway network, highway bus route map, and aviation / high-speed rail transportation hub in a geographical area, and storing a graph model database corresponding only to the spatial grid; Static structured entities in a geographical area are stored in a document-type database corresponding to the spatial grid.

6. The multimodal data query method according to claim 5, characterized in that: The indexing to the corresponding spatial grid based on the position association information and obtaining query data of other modalities related to the main query data in the spatial grid includes: The spatial grid corresponding to the geographical area is indexed based on the location association information, and query data of other modalities related to the main query data in the corresponding spatial grid and / or surrounding spatial grids are obtained.

7. The multimodal data query method according to claim 1, characterized in that: The step of aggregating and sorting the acquired main query data and query data of other modalities related to the main query data, and forming structured recommendation content to return to the user includes: Extracting features of the main query data and query data of other modalities related to the main query data respectively to obtain feature vectors; Perform feature fusion on the feature vectors of different modal data to obtain a fusion vector; Calculating the similarity between the fused vectors, and recommending and ranking the different modal data based on the similarity; The recommended ranking results are displayed in a structured manner.

8. A multimodal data query device, characterized in that: include: A main mode identification unit, configured to obtain a query task and identify a main query data mode of the query task; A data query unit, configured to execute the query task based on the main query data modality and obtain the main query data; an information extraction unit, configured to extract location-related information from the primary query data; a data retrieval unit, configured to index a corresponding spatial grid based on the position association information and obtain query data of other modalities related to the main query data in the spatial grid; The spatial grid corresponds to an actual geographical area and stores multimodal data within the corresponding geographical area; The result output unit is used to aggregate and sort the acquired main query data and query data of other modalities related to the main query data, and form structured recommendation content to return to the user.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the multimodal data query method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the multimodal data query method according to any one of claims 1 to 7 when executed.