POI (Point of Interest) attribute generation method, encoder training method and encoder training device
By introducing multimodal information and large language models into the electronic map, combining position encoder and attribute prediction model, the problems of low coverage and low accuracy of mining point-of-interest attributes are solved, and efficient and accurate updates of point-of-interest attributes are achieved.
Patent Information
- Application Number
- CN202510273087.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-08-01
Smart Images

Figure CN120407894A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of electronic maps, and in particular, to a method for generating point-of-interest attributes, an encoder training method, and an apparatus therefor. Background Art
[0002] In an electronic map, a large number of points of interest (POIs) with various types are displayed. Whether the attributes of the points of interest, such as spatial coordinates, are accurate or not is an important factor affecting the user experience of LBS (Location Based Services).
[0003] The attributes of points of interest mainly include two methods: external acquisition and internal mining. External acquisition is usually provided by a third party or collected by means of task distribution. Although the information obtained in this way is highly accurate, the coverage is limited and the efficiency is low, and it is not applicable to the extraction of attributes of a large area of points of interest. The internal mining method mainly mines the attributes of points of interest through various sources of materials, usually only for single-modal materials, with low utilization rate of the materials, resulting in low accuracy of attribute mining. At the same time, complex feature engineering is required to implement attribute mining, and the calculation cost is relatively high.
[0004] Therefore, there is an urgent need to provide a high-accuracy point-of-interest attribute mining solution. Summary of the Invention
[0005] The present application provides a method for generating point-of-interest attributes, an encoder training method, and an apparatus therefor. By using multi-modal information of points of interest, including spatial positions, as well as multimedia information such as images and texts, the attribute prediction model is used to realize the mining of point-of-interest attributes, improving the mining accuracy, and using multi-modal information to improve the coverage range of attribute mining.
[0006] In a first aspect, the present application provides a method for generating point-of-interest attributes, including:
[0007] Obtain multi-modal information of a target point of interest; wherein, the multi-modal information includes spatial position information and multimedia information;
[0008] Encode the spatial position information based on a pre-trained position encoder to obtain a position encoding feature;
[0009] Input the position encoding feature and the multimedia information into a pre-trained attribute prediction model to obtain the attributes of the target point of interest, where the attributes of the target point of interest include position.
[0010] In a second aspect, the present application provides an encoder training method for training the position encoder provided in the first aspect of the present application. The method includes:
[0011] Obtain a plurality of spatial location information samples and a plurality of multimedia information samples;
[0012] Based on the plurality of spatial location information samples and the plurality of multimedia information samples, train a pre-trained position encoder connected to an attribute prediction model.
[0013] In a third aspect, the present application provides a generating device for point-of-interest attributes, including:
[0014] A multimodal information obtaining module, configured to obtain multimodal information of a target point of interest; wherein, the multimodal information includes spatial location information and multimedia information;
[0015] A position encoding module, configured to encode the spatial location information based on a pre-trained position encoder to obtain position encoding features;
[0016] An attribute prediction module, configured to input the position encoding features and the multimedia information into a pre-trained attribute prediction model to obtain the attributes of the target point of interest, and the attributes of the target point of interest include location.
[0017] In a fourth aspect, the present application provides an encoder training device for training the position encoder provided in the first aspect of the present application, and the device includes:
[0018] A sample obtaining module, configured to obtain a plurality of spatial location information samples and a plurality of multimedia information samples;
[0019] An encoder training module, configured to train a pre-trained position encoder connected to an attribute prediction model based on the plurality of spatial location information samples and the plurality of multimedia information samples.
[0020] In a fifth aspect, the present application provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to execute the method provided in the first aspect or the second aspect of the present application.
[0021] In a sixth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when a processor executes the computer-executable instructions, the method provided in the first aspect or the second aspect of the present application is implemented.
[0022] In a seventh aspect, the present application provides a program product, including a computer program, and when the computer program is executed by a processor, the method provided in the first aspect or the second aspect of the present application is implemented.
[0023] The method for generating point of interest attributes, the encoder training method and device provided by this application introduce spatial location information in addition to multimedia materials such as images and texts during the mining of point of interest attributes. As a modal information, the spatial location information is used to mine the point of interest attributes through multi-modal information, with a wide coverage range and high accuracy. At the same time, the overall framework takes the original multi-modal materials as input, without complex feature engineering, reducing the computational cost and time consumption. The corresponding encoder encodes the input multi-modal information, and the powerful understanding and perception ability of the attribute prediction model is used to mine the point of interest attributes, further improving the accuracy of the point of interest attribute mining. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0025] Figure 1 It is a schematic diagram of a point of interest attribute generation process provided by an embodiment of this application;
[0026] Figure 2 It is a schematic flow diagram of a method for generating point of interest attributes provided by an embodiment of this application;
[0027] Figure 3 It is a schematic structural diagram of an attribute prediction model provided by an embodiment of this application;
[0028] Figure 4 It is a schematic flow diagram of another method for generating point of interest attributes provided by an embodiment of this application;
[0029] Figure 5 It is a schematic diagram of the interface of a shooting task in an electronic map software provided by an embodiment of this application;
[0030] Figure 6 It is a schematic diagram of the update process of a point of interest marked in an electronic map provided by an embodiment of this application;
[0031] Figure 7 It is a schematic flow diagram of an encoder training method provided by an embodiment of this application;
[0032] Figure 8 It is a schematic structural diagram of a pre-training framework provided by an embodiment of this application:
[0033] Figure 9 It is a schematic flow diagram of an encoder pre-training method provided by an embodiment of this application;
[0034] Figure 10 It is a schematic structural diagram of an electronic device provided by an embodiment of this application.
[0035] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and will be described in more detail hereinafter. These drawings and the written description are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of the Embodiments
[0036] Here, exemplary embodiments will be described in detail, and examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0037] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0038] First, some terms related to the present application are explained:
[0039] POI (Point of Interest): In an electronic map, it usually refers to a geographical location, or an entity such as a building or a scenic spot.
[0040] ViT (Vision Transformer): A neural network model for processing images, which only uses the encoder in the basic Transformer to perform image feature extraction tasks.
[0041] Modal alignment task: It refers to aligning data of different modalities, mapping them to the same space, so as to obtain a direct connection between multi-modal data, in order to achieve effective fusion and utilization of multi-modal data.
[0042] The attributes of existing POIs in the electronic map, such as names, addresses, coordinates, etc., may change over time, such as store relocations, name changes, cancellations, etc. At the same time, some new POIs will also continuously appear in the electronic map. In order to adapt to the above changes and update the POIs marked in the electronic map in a timely manner, it is necessary to mine the attributes of the POIs.
[0043] Exemplarily, Figure 1 is a schematic diagram of a process for generating POI attributes provided by an embodiment of the present application, as Figure 1As shown in the figure, the attributes of the point of interest can be reported to the attribute mining device in the electronic map system by collectors, relevant personnel of a third party, etc. through the way of active acquisition. In addition to the active reporting method, in order to improve the coverage and efficiency of the generation of point of interest attributes, the attribute mining device can also use the collected internal materials to obtain the attributes of the point of interest through data analysis and mining, realize the update of the attributes of the point of interest stored in the storage server, and then update or mark the attributes of the point of interest on the electronic map in a timely manner, so as to provide services such as travel and entertainment for users through accurate point of interest attributes.
[0044] In the related art, when mining the attributes of points of interest, single-modal materials are usually used, and a large number of single-modal data are directionally mined through an expert model to obtain the attributes of points of interest. This method fails to make full use of multi-modal information, resulting in low utilization rate of materials, and requires complex feature engineering on the materials in advance, resulting in large information loss, thus resulting in poor accuracy of mining the attributes of points of interest and limited improvement in the coverage range.
[0045] In order to improve the accuracy and coverage of point of interest attribute mining, the present application provides a method for generating point of interest attributes. By using multi-modal information including spatial location information and multimedia information, each modal information is encoded through a corresponding encoder. After modal alignment, the attributes of the point of interest are predicted through a large language model. Through the fusion and alignment of multi-modal materials, the full utilization of multi-modal materials is realized, and the attributes of points of interest are mined by using multi-modal materials, with a wide coverage range and high accuracy; at the same time, taking the original materials as input, there is no need to go through complex feature engineering, reducing the computing cost and improving the efficiency; at the same time, by using the powerful learning and understanding ability of a mature large language model to predict the attributes of points of interest, the accuracy of point of interest attribute mining is further improved.
[0046] Figure 2 The following is a schematic flowchart of a method for generating point of interest attributes provided by an embodiment of the present application. This method can be executed by an electronic device with corresponding data processing capabilities, such as an attribute mining device in an electronic map system, and this device can be a server, a computer or other devices.
[0047] As Figure 2 shown, the method for generating point of interest attributes includes the following steps:
[0048] Step S201, obtaining multi-modal information of the target point of interest; the multi-modal information includes spatial location information and multimedia information.
[0049] Among them, the target point of interest can be any one or more points of interest, which can be newly added points of interest in the electronic map or points of interest that have been published or launched. The target point of interest can also be a point of interest of a specified type, such as a point of interest with a relatively high probability of attribute change.
[0050] The spatial location information is used to describe the coordinates of one or more location points associated with the point of interest. The location points associated with the point of interest can also be called the associated location points of the point of interest, which are the possible location points representing the location of the point of interest, such as coordinates. The multimedia information can be single-modal information, for example, it can only include images or only include text, or it can include multi-modal information, for example, it includes images and text.
[0051] The spatial location information can include at least one of the location information of point elements, line elements, and surface elements. The location information of point elements is used to describe the coordinates of the corresponding points, such as longitude and latitude; the location information of line elements is used to describe the coordinates of multiple points on a curve; the location information of surface elements is used to describe the coordinates of multiple points on the contour line of a closed surface composed of multiple curves.
[0052] The point element in the spatial location information of the target point of interest is the possible location point of the target point of interest. The line element is used to indicate the road near the target point of interest, and the surface element is used to represent the projection of the building of the target point of interest on the ground.
[0053] When the multimedia information only includes images or text, the spatial location information in the multi-modal information is one type of modal information, and the image or text is the other type of modal information. When the multimedia information includes images and text, the spatial location information in the multi-modal information is one type of modal information, and the images and text in the multimedia information are the other two types of modal information, that is, the multi-modal information includes three types of modal information.
[0054] The image in the multimedia information of the target point of interest is also called the associated image, which can include the image of the sign carrying the target point of interest, the image of the building block containing the target point of interest, the image determined by parameters such as shooting position, orientation, and depth of field with the viewpoint located in the coverage area of the target point of interest, etc.
[0055] The text in the multimedia information of the target point of interest can include the text carrying information such as the address and name of the target point of interest, such as the waybill of the target point of interest.
[0056] In the database or storage device, the multi-modal information of multiple points of interest, including spatial location information and multimedia information, can be pre-stored. When mining the attributes of the target point of interest, the multi-modal information of the target point of interest can be read from the database or storage device, and subsequent steps can be executed to obtain the attributes of the target point of interest.
[0057] In some embodiments, the spatial location information of the target point of interest can be extracted from the multimedia information of the target point of interest. Taking the associated image of the target point of interest as the multimedia information as an example, the spatial location information of the target point of interest can be obtained based on the shooting location and / or the viewpoint location of the associated image. Taking the associated text of the target point of interest as the multimedia information as an example, the spatial location information of the target point of interest can be determined based on the coordinates corresponding to keywords such as the address or the name of the point of interest in the associated text.
[0058] Optionally, the multimedia information includes at least one image and at least one text; obtaining the multimodal information of the target point of interest includes:
[0059] Obtaining at least one image and at least one text of the target point of interest; and obtaining the spatial location information based on the at least one image and the at least one text.
[0060] Taking the point of interest as a dimension, the multimedia information of each target point of interest can be counted to obtain the image set and the text set of each target point of interest, where the image set is used to store the images corresponding to the target point of interest, and the text set is used to store the texts corresponding to the point of interest.
[0061] At least one image of the target point of interest may include an image whose content or shooting location is associated with the target point of interest, that is, the associated image of the target point of interest, such as an image whose shooting location is near the target point of interest, an image whose content includes a part of the building of the target point of interest or a part of the signboard, etc.
[0062] At least one text of the target point of interest can be any type of information containing the address of the target point of interest, such as a waybill, a news report, a transaction record, etc.
[0063] At least one text of the target point of interest may include the associated text of the target point of interest, such as an associated waybill. The attributes of the target point of interest described in the associated text satisfy the association conditions with the attributes of the stored target point of interest, such as one or more of a high address similarity, a high name similarity, and a consistent contact information.
[0064] The associated waybill of the target point of interest is a waybill whose parameters such as the waybill address, the contact information, and the sender name are associated with the target point of interest. For example, the waybill address (the consignee address or the shipper address) has a high similarity with the address of the stored target point of interest, the contact information is consistent with the contact information of the stored target point of interest, and the sender or recipient name has a high similarity with the name of the stored target point of interest.
[0065] For the acquired image, based on the recognized content in the image, as well as the shooting address, camera parameters, etc. of the image, the corresponding point of interest of the image can be determined, so as to obtain an associated image of the point of interest. For the acquired text, the key information describing the attributes of the point of interest in the text can be extracted, and based on the extracted key information, the corresponding point of interest of the text can be determined, and an associated text of the corresponding point of interest can be obtained.
[0066] Taking the waybill as an example, for the acquired waybill, parameters such as the waybill address, contact information, sender name, etc. in the waybill can be extracted, and based on these parameters, the corresponding point of interest of the waybill can be determined, so as to obtain an associated waybill of the corresponding point of interest.
[0067] For each target point of interest, based on the associated image and associated text of the target point of interest, the spatial position information of the target point of interest can be obtained.
[0068] Specifically, based on the shooting position or viewpoint position of the associated image of the target point of interest, the coordinates corresponding to the waybill address (the address with a high similarity to the stored address of the target point of interest) in the associated waybill, and the coordinates of some elements such as buildings and roads recognized in the associated image, the spatial position information of the target point of interest can be obtained.
[0069] The first candidate position can be determined based on the shooting position, camera orientation and camera field of view of the associated image, and the second candidate position can be determined based on the coordinates of buildings, roads, etc. recognized in the associated image. Exemplarily, the spatial position information can be determined as the average value of the first candidate position, the second candidate position, the viewpoint position of the associated image and the coordinates corresponding to the waybill address in the associated waybill. Or, the spatial position information can be determined as including the average value (the position information of the point element) of the first candidate position, the viewpoint position of the associated image and the coordinates corresponding to the waybill address in the associated waybill, and the position information of the surface element and the position information of the line element composed of the coordinates of buildings, roads, etc. recognized in the associated image.
[0070] In some embodiments, when determining the spatial position information, some of the above features can also be omitted, such as omitting the second candidate position, omitting the viewpoint position, etc.
[0071] The automatic extraction of the spatial position information through multimedia information including two modalities of text and images improves the automation degree of obtaining the spatial position information, and at the same time enriches the content of the spatial position information, providing sufficient position information for the position prediction of the point of interest.
[0072] Step S202, encode the spatial position information based on a pre-trained position encoder to obtain position encoding features.
[0073] Input the spatial location information in the multimodal information into a pre-trained location encoder, and encode the spatial location information through the location encoder to obtain location encoding features.
[0074] Exemplarily, the location encoder can be an encoder based on the attention mechanism, such as the Transformer model.
[0075] In step S203, input the location encoding features and the multimedia information into a pre-trained attribute prediction model to obtain the attributes of the target point of interest, where the attributes of the target point of interest include location.
[0076] The attribute prediction model is a multimodal model for processing multimodal data such as input location encoding, text, images, etc. Through cross-modal fusion and understanding, it realizes attribute prediction including the location of the target point of interest.
[0077] Exemplarily, the attribute prediction model can be a multimodal large model, specifically a multimodal large language model.
[0078] The core of the attribute prediction model can be a large language model (Large Language Model, LLM). Since the large language model can directly process text, the input text is processed into an embedding vector by the embedding layer of the large language model. If the multimedia information only includes text, the attribute prediction model is a large language model connected with a modality alignment layer. The text of the target point of interest can be input into the embedding layer of the large language model to obtain an embedding vector, and the location encoding features are converted to the space where the embedding vector is located through the modality alignment layer connected to the large language model to achieve modality alignment. Then, the output of the modality alignment layer and the output of the embedding layer are input into the subsequent network layers of the large language model, such as the encoder, decoder, etc., to obtain the attributes of the target point of interest.
[0079] If the multimedia information includes images, the attribute prediction model is a large language model connected with a visual encoder and a modality alignment layer. It is necessary to encode the images through the visual encoder to obtain image encoding features, and then perform modality alignment on the location encoder features and the image encoder features through the modality alignment layer, that is, both are mapped to the space where the embedding vector (the vector output by the embedding layer of the large language model) is located.
[0080] If the multimedia information only includes images, the position encoder features and image encoder features after modality alignment are input into the large language model, and the large language model is used to analyze the input multi-modal data to predict the attributes of the target interest point. If the multimedia information includes text in addition to images, the position encoder features and image encoder features after modality alignment, as well as the text, are input into the large language model, and the large language model is used to analyze the input multi-modal data to predict the attributes of the target interest point.
[0081] Exemplarily, Figure 3 is a schematic structural diagram of the attribute prediction model provided by the embodiment of the present application. As Figure 3 shown, the attribute prediction model is a multi-modal large language model that can process data in three modalities: images, text, and positions. The attribute prediction model includes a large language model, a visual encoder, and a modality alignment layer. Specifically, the internal architecture of the large language model includes an embedding layer, an encoder, a decoder, etc. It should be noted that the specific architecture of the large language model is not limited in the solution of the present application. Even if some large models do not have an encoder or a decoder, they can still meet the technical requirements of this embodiment, thus ensuring the flexibility and applicability of the solution.
[0082] In some embodiments, the layer where the large language model is located is denoted as the large language model layer.
[0083] The spatial position information includes the position information of three elements: points, lines, and planes. The point element is specifically a coordinate point, the line element is a line composed of multiple point elements, which can be used to describe roads, and the plane element is used to describe the target interest point, which can be a polygon obtained by looking down on the target interest point; after the spatial position information is input into the position encoder, it is processed into position encoding feature F p ; the image is processed into image encoder feature F img after passing through the visual encoder; the position encoder feature and the image encoder feature are subjected to modality alignment through the modality alignment layer to obtain the position encoding feature Fˊ[[ID=__18]] p after modality alignment and the image encoding feature Fˊ img after modality alignment; the text is input into the embedding layer of the large language model to obtain the embedding vector E txt ; the position encoding feature Fˊ p after modality alignment, the image encoding feature Fˊ img after modality alignment, and the embedding vector E txt are used as the input of the encoder of the large language model, and through the encoder and decoder of the large language model, the attributes of the predicted target interest point are output.
[0084] In addition to predicting the position of the interest point, the attribute prediction model can also predict the name, contact information, etc.
[0085] By providing accurate attributes of points of interest, such as coordinates, it can help users find points of interest more accurately and quickly. At the same time, during the navigation process, more accurate location information of points of interest can provide more accurate route planning and navigation guidance.
[0086] The method for generating local point-of-interest attributes provided in this application, when mining point-of-interest attributes, in addition to multimedia materials such as images and texts, also introduces spatial location information. As a modal information, the spatial location information is used to mine point-of-interest attributes through multi-modal information, with a wide coverage range and high accuracy. At the same time, the overall framework takes the original multi-modal materials as input, without complex feature engineering, reducing the computational cost and time consumption. By encoding the input multi-modal information through a corresponding encoder and using the powerful understanding and perception ability of the attribute prediction model to mine point-of-interest attributes, the accuracy of point-of-interest attribute mining is further improved.
[0087] Figure 4 It is a schematic flowchart of another method for generating point-of-interest attributes provided in the embodiment of this application. This embodiment is a further refinement of steps S201 and S203 based on the Figure 2 embodiment shown. In this embodiment, the spatial location information includes the location information of point elements, line elements, and surface elements. As shown in Figure 4 the method for generating point-of-interest attributes provided in this embodiment can specifically include the following steps:
[0088] Step S401, obtain multiple initial images of the target point of interest taken.
[0089] The electronic map system can issue a shooting task for the target interest, and send the shooting task to the terminal of the collector or user, so as to obtain the images uploaded by the collector or user when executing the shooting task. The images may include the signs, buildings, etc. of the target point of interest, and the images uploaded when executing the shooting task of the target interest are used as the initial images of the target point of interest or part of the initial images.
[0090] In the electronic map software, the shooting task can be displayed to the user. After the user accepts the shooting task, the user can upload the images of the target point of interest taken through the corresponding interface.
[0091] Exemplarily, Figure 5 It is a schematic diagram of the interface of the shooting task in the electronic map software provided in the embodiment of this application. As shown in Figure 5As shown in the figure, the user can view the task list on the order receiving interface of the electronic map software. The tasks in the task list displayed on the interface include Task 51 to Task 53, and more tasks can be displayed by swiping down. If the user wants to view the details of a task in the task list, they can select one task, such as Task 52, and open the details page of this task. The details page displays the information of the point of interest corresponding to Task 52 (such as XX Nail Salon (XX Square Store)) and some information about the task, and also includes a "Do Task" button. After the user arrives near the location of the point of interest corresponding to Task 52, they click the "Do Task" button to open the image upload interface, and the user can upload the photos taken on this image upload interface.
[0092] After the user discovers that a point of interest at a certain location has been updated, such as a new store opened, the user can also actively upload the image of this point of interest through the corresponding interface as part of the initial images.
[0093] Step S402, based on the recognition results of the multiple initial images, screen at least one image of the target point of interest from the multiple initial images.
[0094] At least one image of the target point of interest is a set of associated images screened from multiple initial images of the target point of interest.
[0095] After obtaining the initial image of the target point of interest, identify information such as text, roads, road signs, and buildings in the initial image. Based on the recognition results, determine whether the initial image is an associated image of the target point of interest.
[0096] Specifically, based on the information such as text, buildings, roads, and road signs in the recognized initial image, determine whether the building in the initial image is the building of the target point of interest; if so, determine that the initial image is an associated image of the target interest.
[0097] If the initial image includes a signboard, the text in the signboard can be extracted, and the extracted text is compared with the name of the target point of interest stored or the signboard text of the target point of interest stored. Based on the comparison results, determine whether the initial image is an associated image of the target point of interest.
[0098] In some embodiments, the viewpoint coordinates of the initial image can be determined based on the shooting location and camera parameters of the initial image. Based on the viewpoint coordinates of the initial image and the recognition results of the initial image, determine the point of interest corresponding to the initial image, and regard the initial image as an associated image of its corresponding point of interest.
[0099] Step S403, based on the attribute information of each waybill in the waybill library, determine the waybills associated with the target point of interest to obtain at least one text of the target point of interest.
[0100] Among them, the attribute information of the waybill includes one or more of the waybill address, the name of the sender, and the contact information.
[0101] The waybill library is used to store the collected waybills. For the newly added waybills in the waybill library, based on the attribute information of the waybill, associated points of interest are assigned to the waybill, so as to obtain the associated waybills of the points of interest (i.e., the waybills associated with the points of interest), that is, the waybill is determined as the associated waybill of the associated point of interest.
[0102] Exemplarily, it is possible to determine whether the waybill is an associated waybill of the target point of interest based on the similarity between the waybill address and the address of the stored target point of interest.
[0103] It is possible to determine whether the waybill is an associated waybill of the target point of interest based on the similarity or the judgment result of whether they are the same between the values of each item in the attribute information of the waybill and the values of each attribute of the stored target point of interest.
[0104] Exemplarily, if the contact information of the waybill is the same as the contact information of the stored target point of interest, and the name of the sender has a high similarity with the name of the stored target point of interest, then it is determined that the waybill is an associated waybill of the target point of interest. If the contact information of the waybill is not the same as the contact information of the stored target point of interest, but the name of the sender or the waybill address has a high similarity with the name or address of the stored target point of interest, then it is determined that the waybill is an associated waybill of the target point of interest.
[0105] Step S404, extract the coordinates carried by the at least one image and the coordinates carried by the at least one text to obtain the position information of the point elements.
[0106] The coordinates carried by the image can be the coordinates of the camera when the image is taken, or the coordinates of the viewpoint of the image. The coordinates carried by the text are the coordinates corresponding to the address described in the text. Taking the waybill as an example, the coordinates carried by the waybill are the coordinates corresponding to the waybill address in the waybill, specifically the coordinates corresponding to the waybill address with a high similarity to the address of the target point of interest (the shipper address or the consignee address).
[0107] The coordinates mentioned in this application are usually geographical coordinates, and can also be coordinates in other world coordinate systems describing the real world, such as coordinates in the northeast sky coordinate system. This application does not limit this.
[0108] Extract the coordinates carried by each associated image of the target point of interest and the coordinates carried by each associated text to obtain the coordinates of multiple point elements, that is, obtain the position information of the point elements in the spatial position information.
[0109] When the multimedia information only includes associated images, it is possible to obtain the position information of the point elements in the spatial position information only based on the coordinates carried by the associated images.
[0110] When the multimedia information only includes associated text, the position information of the point elements in the spatial position information can be obtained based on the coordinates carried by the associated text; for the position information of the line elements and surface elements in the spatial position information, it can be omitted, or based on the position information of the point elements in the spatial position information, the position information of the line elements and surface elements associated with the point elements can be extracted from the road network data, so as to obtain the spatial position information of the target point of interest including the position information of the point, line, and surface elements. The surface element associated with the point element is the surface element containing the point element, and the line element associated with the point element is the line element around the surface element associated with the point element.
[0111] Step S405, based on the roads and buildings included in the at least one image, obtain the position information of the line elements and the position information of the surface elements respectively.
[0112] Based on the information of the roads included in the associated image of the identified target point of interest, including information such as road names and road sections, the coordinates of multiple points on the roads included in the associated image can be read from the pre-stored mapping relationship between the road network and coordinates, so as to obtain the position information of a line element.
[0113] Specifically, in the stored road network data, the road section corresponding to the road included in the associated image can be obtained, and the coordinates of multiple points on the road section can be read from the road network data, so as to obtain the position information of a line element.
[0114] Based on parameters such as the shape, texture, and color of the buildings included in the identified associated image, the buildings that match the buildings included in the associated image can be determined from the pre-stored buildings within the preset range of the target point of interest, and the coordinates of multiple points describing the outline of the matching buildings stored can be regarded as the position information of a surface element.
[0115] Optionally, based on the roads and buildings included in the at least one image, obtaining the position information of the line elements and the position information of the surface elements respectively includes:
[0116] Identifying the identifiers of the roads and buildings included in the at least one image; based on the identified identifiers of the roads and buildings, reading the geographical coordinates of multiple points on the roads included in the at least one image from the first target dataset to obtain the position information of the line elements, and reading the geographical coordinates of each vertex of the buildings included in the at least one image from the second target dataset to obtain the position information of the surface elements; wherein, the first target dataset includes the identifiers of the roads in the first target geographical area and the geographical coordinates of multiple points thereon, the second target dataset includes the identifiers of the buildings in the second target geographical area and the geographical coordinates of each vertex thereof, and the target point of interest is located within the first target geographical area and the second target geographical area.
[0117] Exemplarily, the first target data set may be electronic map data of the first target geographical area, which is used to store relevant data of map elements within the first target geographical area, including the coordinates of points on the map elements. The map elements include roads, buildings, road signs, etc. The second target data set may be map rendering data of the second target geographical area, which contains model data of each building to be rendered within the second target geographical area, including the attributes of the vertices of the building. The attributes of the vertices include geographical coordinates, and also include colors, textures, etc. required for rendering.
[0118] The identifier of a road may be the name of the road, the number of the road, etc., and the identifier of a building may be the name of the building, the number, etc.
[0119] After identifying the identifiers of roads and buildings in the associated image based on the object detection algorithm, based on the identified identifier of the road, as well as parameters such as the buildings near the road and the shooting position of the associated image, determine the road section corresponding to the road identified in the associated image from the road network stored in the first target data set, read the geographical coordinates of multiple points on this road section stored in the first target data set, and obtain the position information of this road section or the road identified in the associated image, which is a linear element; at the same time, based on parameters such as the identifier of the building identified in the associated image and the area where it is located, determine the building that matches the building identified in the associated image from the buildings within this area stored in the second target data set, read the geographical coordinates of multiple points on the contour polygon of this matching building stored in the second target data set, and obtain the position information of this matching building or the building identified in the associated image, which is a planar element.
[0120] The linear element can be a polyline, which is obtained by connecting adjacent points among the multiple points it contains using line segments; the planar element is a polygon, which is a closed polygon composed of multiple points it includes.
[0121] By combining the recognition results of the image and the way of storing the point sets of the corresponding map elements, the acquisition of the position information of the linear element and the planar element is realized with high accuracy. The position information of multi-dimensional elements is used to characterize the positions of the point-of-interest building, nearby buildings, and nearby roads, providing a rich data basis for the model to predict the position of the point of interest and improving the accuracy of position prediction.
[0122] Step S406: Encode the spatial position information based on a pre-trained position encoder to obtain position encoding features.
[0123] Step S407: Input the at least one image into a pre-trained visual encoder to obtain image encoding features of the at least one image.
[0124] Step S408: Perform modality alignment on the position encoding features and the image encoding features via the modality alignment layer.
[0125] The modality alignment layer can perform modality alignment based on the attention mechanism, such as the self-attention mechanism, the spatial attention mechanism, etc.
[0126] Specifically, a cross-attention layer can be used to perform modality alignment on the position encoding features and the image encoding features, and convert them to the space where the output of the embedding layer is located.
[0127] The cross-attention layer can be an attention layer based on a learnable query, that is, the relationship between the query vector Q and the input is not fixed and is obtained through learning.
[0128] Step S409: Input the position encoding features and the image encoding features after modality alignment, and the at least one piece of text into a pre-trained large language model layer to obtain the attributes of the target point of interest.
[0129] Specifically, the associated text of the target point of interest is input into the embedding layer of the large language model to obtain an embedding vector; the embedding vector, and the position encoding features and the image encoding features after modality alignment are used as the input of the encoder of the large language model, and the encoder and decoder of the large language model are used to predict the attributes of the target point of interest.
[0130] In this embodiment, information in three modalities, namely associated text, associated images, and spatial location information, is used as the original data for mining the attributes of points of interest. The information of multiple modalities is fully utilized, improving the accuracy of mining the attributes of points of interest; the associated images and associated text are screened from the initially captured images and the waybill library using corresponding screening strategies, with high accuracy, improving the degree of association between the associated text and associated images and the points of interest, reducing the number of associated waybills and associated images to be processed, and improving the efficiency of attribute mining; at the same time, the spatial location information in the original data is automatically extracted from the associated text and associated images, improving the efficiency of obtaining spatial location information; by including multi-dimensional spatial location information of points, lines, and surfaces, the prediction of the location of points of interest is improved, and the accuracy of prediction is improved; by encoding the associated images and location information through a visual encoder and a position encoder respectively, after modality alignment processing, they are input into the large language model together with the associated text, and the large language model is used to analyze the data in the three modalities to realize the mining of the attributes of points of interest, improving the accuracy of attribute mining.
[0131] The method for generating the attributes of points of interest provided in any of the foregoing embodiments of the present application can be executed periodically or irregularly to realize the generation of new attributes of points of interest and the update of existing attributes of points of interest.
[0132] After obtaining the attributes of the newly added point of interest or the attributes of the existing point of interest after update, the newly added point of interest can be marked on the electronic map, and based on the attributes of the existing point of interest after update, the marked existing point of interest on the electronic map can be adjusted.
[0133] Optionally, the attributes of the target point of interest include coordinates, address, and name; the method further includes: marking the target point of interest on the electronic map based on the coordinates, address, and name of the target point of interest.
[0134] Marking a point of interest on the electronic map can specifically be adding a pin at the position corresponding to the coordinates of the point of interest on the electronic map, connecting the pin and a display box through a connection line, and displaying the marking information of the point of interest, such as attributes like name, icon, etc., through the display box.
[0135] For the point of interest marked on the electronic map, if it is detected that the attributes of the point of interest change, then based on the new attributes, the marking information or the marking position of the point of interest, that is, the position of the pin, is updated.
[0136] Figure 6 It is a schematic diagram of the update process of the point of interest marked on the electronic map provided by the embodiment of the present application. As Figure 6 shown, at time t0, there are 2 points of interest marked in a certain area of the electronic map, POI61 (XX gas station) and POI62 (XX restaurant), and their respective positions and marking information are as Figure 6 shown; through the mining of the attributes of the points of interest, it is detected that the address of POI61 changes, and the other attributes remain unchanged. At the same time, there is a newly added point of interest POI63 (XX scenic area) in this area. Then, the points of interest marked in this area of the electronic map can be updated based on these changes to obtain the electronic map at time t1 (t1 > t0).
[0137] Figure 7 It is a schematic flowchart of an encoder training method provided by the embodiment of the present application. This encoder training method is used to train the position encoder in the foregoing embodiment and can be executed by a dedicated training device or by an attribute mining device. As Figure 7 shown, the encoder training method provided in this embodiment includes the following steps:
[0138] Step S701, obtain a plurality of spatial position information samples and a plurality of multimedia information samples.
[0139] Step S702, based on the plurality of spatial position information samples and the plurality of multimedia information samples, train the pre-trained position encoder connected to the large language model.
[0140] The spatial location information sample is a sample of information in the modality of spatial location information, and the multimedia information sample is a sample of multimedia information, which may include an image sample and / or a text sample.
[0141] The position encoder can be pre-trained in advance to obtain a pre-trained position encoder, and then the pre-trained position encoder is connected to the large language model. If the multiple multimedia information samples include an image sample, the large language model can also be connected to a visual encoder to obtain a training framework. If the multimedia information sample does not include an image sample, that is, the multimedia information sample only includes a text sample, the large language model in the training framework does not need to be connected to the visual encoder.
[0142] The training framework is trained with multiple spatial location information samples and multiple multimedia information samples. Based on the loss value between the output of the large language model and the ground truth, the parameters of each part in the training framework, such as the position encoder, the large language model, and the visual encoder, are fine-tuned until the training end condition is met.
[0143] Among the multiple multimedia information samples, there may be multimedia information samples corresponding to the same interest point sample as some of the spatial location information samples. The multimedia information samples and the spatial location information samples corresponding to the same interest point sample form a multi-modal information sample of the interest point sample. Thus, the multi-modal information sample of the interest point sample is used as the input of the training framework, and then, based on the loss value between the position of the interest point predicted by the large language model and the actual position of the interest point, at least the position encoder in the training framework is adjusted until the training end condition is met.
[0144] In some embodiments, the multiple spatial location information samples and the multiple multimedia information samples can respectively form a single-modal training set and a multi-modal training set; the training samples in the single-modal training set are information in one modality, that is, an image sample, a text sample, or a spatial location information sample; the training samples in the multi-modal training set are multi-modal, including information in multiple modalities, that is, including at least two of the image sample, the text sample, and the spatial location information sample. The pre-trained position encoder connected to the large language model is trained with the training samples in the single-modal training set and the multi-modal training set to obtain a pre-trained position encoder and a pre-trained large language model.
[0145] Training the training framework with both single-modal and multi-modal training samples enables the overall model obtained after training to have the ability to predict the interest point attributes in both single-modal input and multi-modal input cases, improving the application scope of the model.
[0146] Since the large language model connected with a visual encoder has good understanding ability for images and texts, especially texts, but has far from enough understanding of spatial position information. In order to improve the large language model's understanding ability of spatial position information, the embodiment of the present application also provides a pre-training method for a position encoder for spatial position information. Before training the encoder, pre-training can be performed based on the encoder pre-training method provided by the embodiment of the present application.
[0147] Optionally, the encoder training method further includes:
[0148] Obtain a plurality of pre-training samples, where the plurality of pre-training samples include a plurality of spatial position information samples; based on the plurality of pre-training samples, pre-train a position encoder connected with a head network; the head network is used to output a prediction result under a corresponding task based on the position encoding features output by the position encoder.
[0149] The head network and the position encoder as the backbone network form a pre-training framework. Use a plurality of pre-training samples to pre-train the position encoder in the pre-training framework. Adjust the parameters of the position encoder through the loss value between the prediction result output by the head network and the ground truth (GT) of the pre-training samples until the pre-training end condition is met, and obtain the pre-trained position encoder.
[0150] The pre-training samples can include samples corresponding to elements such as points, lines, and planes. The pre-training samples can include samples corresponding to various understanding tasks, such as classification tasks, regression tasks, etc. The classification task can be used to determine the type of the input spatial position information, including three types: points, lines, and planes. The regression task is used to calculate the parameters of the input spatial position information, such as area, center point, etc., or determine the relationship between multiple spatial position information.
[0151] The head network can be an MLP (Multilayer Perceptron), a CNN (Convolutional Neural Network), etc.
[0152] Exemplarily, Figure 8 is a schematic structural diagram of a pre-training framework provided by an embodiment of the present application, as Figure 8As shown, in the pre-training framework, the position encoder is connected to an MLP. The pre-training samples can be divided into three categories: points, lines, and planes. During pre-training, according to the corresponding tasks, a batch of pre-training samples corresponding to the tasks are input. By comparing with the classification results or regression results output by the MLP through the corresponding ground truth, such as attributes, relationships, etc., and using the loss function, the loss value is calculated. Through the backpropagation of the loss value, the parameters of the position encoder are adjusted to complete one round of pre-training. And so on, until the number of pre-training rounds or time reaches the corresponding upper limit value, or the obtained loss value meets certain conditions, then the pre-training is terminated, and the pre-trained position encoder is obtained.
[0153] Figure 9 FIG. is a schematic flowchart of an encoder pre-training method provided by an embodiment of the present application. This encoder pre-training method is used to pre-train the position encoder in the foregoing embodiment, and can be executed by a dedicated pre-training device, or by a training device or an attribute mining device. As Figure 9 shown, the encoder pre-training method provided in this embodiment includes the following steps:
[0154] Step S901, obtain a plurality of pre-training samples.
[0155] Among them, the plurality of pre-training samples include a pre-training sample set corresponding to the attribute detection task and a pre-training sample set corresponding to the relationship detection task. The pre-training sample set corresponding to the attribute detection task includes a plurality of first spatial position information samples, and the pre-training sample set corresponding to the relationship detection task includes a plurality of second spatial position information samples.
[0156] The attribute detection task is used to detect the attributes of an input single spatial position information sample, such as type, center point, area, etc.; the relationship detection task is used to detect the relationship between a plurality of input spatial position information samples, including distance, spatial relationship, position prediction of points of interest, etc.
[0157] Optionally, the pre-training samples corresponding to the attribute detection task include a first sample subset corresponding to the type recognition task, a second sample subset corresponding to the area calculation task, and a third sample subset corresponding to the center point calculation task; the type recognition task is used to recognize the type of the first spatial position information sample, including three types: points, lines, and planes; the area calculation task is used to calculate the area of the first spatial position information sample of the plane type; the center point calculation task is used to determine the center point coordinates of the first spatial position information sample of the plane type.
[0158] Optionally, the pre-training samples corresponding to the relationship detection task include: a fourth sample subset corresponding to the spatial relationship detection task, a fifth sample subset corresponding to the distance calculation task, and a sixth sample subset corresponding to the coordinate prediction task; the spatial relationship detection task is used to detect the spatial relationship between multiple input second spatial position information samples; the distance calculation task is used to calculate the distance between multiple input second spatial position information samples; the coordinate prediction task is used to predict the coordinates of an interest point based on multiple second spatial position information samples of the same interest point.
[0159] The spatial relationships between line and surface type spatial position information samples include, but are not limited to: parallel, intersecting, tangent, etc., and intersection can be further divided into perpendicular and skew intersection.
[0160] The spatial relationships between point type and surface type spatial position information samples include, but are not limited to: a point is inside a surface, a point is on the contour line of a surface, a point is on one side of a surface, etc.
[0161] The spatial relationships between point type and line type spatial position information samples include, but are not limited to: a point is on a line, a point is on one side of a line, etc.
[0162] In the stage of pre-training sample collection, for the same interest point, multiple spatial position information samples of different types of the interest point can be collected, and labels of the corresponding interest points are added to the spatial position information samples, so that when the coordinate prediction task is executed, multiple second spatial position information samples of the same interest point can be screened out through the labels of the corresponding interest points.
[0163] Step S902, based on the multiple first spatial position information samples and the multiple second spatial position information samples, pre-train the position encoder connected with the head network, so as to adjust the parameters of the position encoder based on the loss value between the predicted attribute of the first spatial position information sample output by the head network and the attribute true value, and the loss value between the predicted relationship of the multiple second spatial position information samples output by the head network and the relationship true value.
[0164] Pre-training can be first performed based on multiple first spatial position information samples (or multiple second spatial position information samples), then pre-training can be performed based on multiple second spatial position information samples (or multiple first spatial position information samples), or pre-training can be performed alternately based on the spatial position information samples of the two tasks, that is, first pre-training is performed based on a batch of first spatial position information samples (or second spatial position information samples), then pre-training is performed based on a batch of second spatial position information samples (or first spatial position information samples), and pre-training is performed based on the next batch of first spatial position information samples (or second spatial position information samples), and so on.
[0165] Through the pre-training paradigm provided by steps S901 and S902, the understanding ability of the large language model for spatial position vectors is strengthened, and some basic spatial operations can be realized. On this basis, it is also necessary to connect the position encoder to the large language model, and through tasks related to points of interest, strengthen the large language model's understanding of the generation principle of point-of-interest attributes. Specifically, the second pre-training can be carried out by the methods provided in steps S903 and S904.
[0166] This embodiment is described by taking the case where the large language model is connected to a visual encoder as an example. If the multi-modal information does not involve images, the visual encoder and related steps can be omitted.
[0167] Step S903, after the pre-training of the position encoder is completed, connect the pre-trained position encoder to the large language model connected to the visual encoder to obtain a pre-training framework.
[0168] Step S904, based on at least one of the seventh sample subset corresponding to the map understanding task, the eighth sample subset corresponding to the text address detection task, the ninth sample subset corresponding to the image position detection task, and the tenth sample subset corresponding to the modality alignment task, pre-train the pre-training framework to calculate the loss value based on the deviation between the output of the large language model and the ground truth, and adjust the parameters of the position encoder and the visual encoder based on the loss value to obtain the pre-trained position encoder and visual encoder.
[0169] Among them, the text sample is specifically a waybill sample, and the image sample is an image of the point of interest taken.
[0170] The samples in the seventh sample subset corresponding to the map understanding task are masked image samples or text samples, and the loss value of the loss function is calculated based on the deviation between the content of the masked part predicted by the large language model and the actual content.
[0171] For the image samples in the seventh sample subset, part of the content of the signboard of the point of interest in the image sample can be masked by a mask, or part of the image of the building of the point of interest in the image sample can be masked. After the image sample is encoded by the visual encoder, it is input into the large language model, and the large language model infers the content of the masked part to realize the prediction of the name, address, etc. of the point of interest.
[0172] For the text samples in the seventh sample subset, part of the information describing the address, name, etc. in the text sample can be masked, and the large language model infers the content of the masked part.
[0173] The samples in the eighth sample subset corresponding to the text address detection task are text samples describing points of interest, and the loss value of the loss function is calculated based on the deviation between the coordinates of the point of interest corresponding to the text sample predicted by the large language model and the actual coordinates of the point of interest.
[0174] The samples in the ninth sample subset corresponding to the image position detection task are image samples containing at least part of the point of interest or image samples collected near the point of interest, and the loss value of the loss function is calculated based on the deviation between the coordinates of the point of interest corresponding to the image sample predicted by the large language model and the actual coordinates of the point of interest.
[0175] The samples in the tenth sample subset corresponding to the modality alignment task include multiple multimodal sample pairs. Each multimodal sample pair corresponds to a point of interest and includes at least two of the spatial position information sample, image sample, and text sample corresponding to the point of interest. The loss value of the loss function is calculated based on the deviation between the attributes of the point of interest corresponding to the multiple multimodal sample pairs predicted by the large language model and the actual attributes of the point of interest.
[0176] During the second pre-training, the large language model can be frozen, that is, the parameters of the large language model are fixed. The pre-training of the pre-training framework is carried out through the sample subsets under various provided tasks. The loss value between the true value and the output of the large language model is calculated using the loss function, and the parameters of the visual encoder and the position encoder are adjusted through the backpropagation of the loss value.
[0177] Through the pre-training paradigms of the above multiple tasks, the large language model is enabled to have the ability to understand spatial position information and the ability to understand the process of generating the position of the point of interest, so as to predict the attributes of the point of interest through multimodal information containing spatial position information.
[0178] After outputting the pre-trained position encoder and visual encoder, it enters the training or fine-tuning stage. Multiple training samples in the training set, that is, multiple spatial position information samples and multiple text-image information samples, are used to train the training framework. The large language model in this training framework is connected to the pre-trained position encoder and the pre-trained visual encoder. Through the training stage, the parameters of the large language model, the pre-trained position encoder, and the pre-trained visual encoder are fine-tuned to obtain the pre-trained large language model, position encoder, and visual encoder.
[0179] Corresponding to the foregoing method embodiments, an apparatus for generating point-of-interest attributes provided by the embodiments of the present application includes: a multi-modal information obtaining module, configured to obtain multi-modal information of a target point of interest; wherein the multi-modal information includes spatial location information and multimedia information; a location encoding module, configured to encode the spatial location information based on a pre-trained location encoder to obtain location encoding features; an attribute prediction module, configured to input the location encoding features and the multimedia information into a pre-trained attribute prediction model to obtain the attributes of the target point of interest, and the attributes of the target point of interest include location.
[0180] Optionally, the graphic and text information includes at least one image and at least one text, and the multi-modal information obtaining module includes: an associated graphic and text obtaining unit, configured to obtain at least one image and at least one text of the target point of interest; a location information obtaining unit, configured to obtain the spatial location information based on the at least one image and the at least one text.
[0181] Optionally, the spatial location information includes location information of point elements, line elements, and surface elements, and the location information obtaining unit includes: a point element extraction sub-unit, configured to extract coordinates carried by the at least one image and coordinates carried by the at least one text to obtain the location information of the point elements; a line and surface element extraction sub-unit, configured to obtain the location information of the line elements and the location information of the surface elements respectively based on roads and buildings included in the at least one image.
[0182] Optionally, the line and surface element extraction sub-unit is specifically configured to: identify identifiers of roads and buildings included in the at least one image; based on the identified identifiers of the roads and buildings, read geographical coordinates of multiple points on the roads included in the at least one image from a first target data set to obtain the location information of the line elements, and read geographical coordinates of vertices of the buildings included in the at least one image from a second target data set to obtain the location information of the surface elements; wherein the first target data set includes identifiers of roads in a first target geographical area and geographical coordinates of multiple points thereon, and the second target data set includes identifiers of buildings in a second target geographical area and geographical coordinates of their vertices, and the target point of interest is located in the first target geographical area and the second target geographical area.
[0183] Optionally, the associated graphic and text acquisition unit is specifically configured to: obtain multiple initial images of the target point of interest captured; based on the recognition results of the multiple initial images, screen and obtain the at least one associated image from the multiple initial images; based on the attribute information of each waybill in the waybill library, determine the waybill associated with the target point of interest to obtain the at least one text; wherein, the attribute information of the waybill includes one or more of the waybill address, sender name, and contact information.
[0184] Optionally, the multimedia information includes at least one image and at least one text; the attribute prediction model includes a visual encoder, a modality alignment layer, and a large language model layer; the attribute prediction module is specifically configured to: input the at least one image into a pre-trained visual encoder to obtain the image encoding features of the at least one image; via the modality alignment layer, perform modality alignment on the position encoding features and the image encoding features; input the position encoding features and the image encoding features after modality alignment, and the at least one text, into a pre-trained large language model layer to obtain the attributes of the target point of interest.
[0185] Optionally, the attributes of the target point of interest include coordinates, address, and name; the device for generating the attributes of the point of interest further includes: a point of interest annotation module, configured to annotate the target point of interest on an electronic map based on the coordinates, address, and name of the target point of interest.
[0186] The device for generating the attributes of the point of interest provided in the embodiments of the present application can be used to execute the technical solutions of the method for generating the attributes of the point of interest provided in any of the above embodiments of the present application. The implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.
[0187] The embodiments of the present application further provide an encoder training device for training the position encoder provided in any embodiment of the present application. The encoder training device includes:
[0188] A sample acquisition module, configured to obtain multiple spatial position information samples and multiple graphic and text information samples; an encoder training module, configured to train a pre-trained position encoder connected to a large language model based on the multiple spatial position information samples and the multiple graphic and text information samples.
[0189] Optionally, the encoder training device further includes a first pre-training module, configured to:
[0190] Obtain multiple pre-training samples, where the multiple pre-training samples include multiple spatial position information samples; based on the multiple pre-training samples, pre-train a position encoder connected to a head network; the head network is configured to output a prediction result under a corresponding task based on the position encoding features output by the position encoder.
[0191] Optionally, the multiple pre-training samples include a pre-training sample set corresponding to an attribute detection task and a pre-training sample set corresponding to a relationship detection task. The pre-training sample set corresponding to the attribute detection task includes multiple first spatial position information samples, and the pre-training sample set corresponding to the relationship detection task includes multiple second spatial position information samples. The first pre-training module is specifically configured to: based on the multiple first spatial position information samples and the multiple second spatial position information samples, pre-train the position encoder connected to the head network, so as to adjust the parameters of the position encoder based on the loss value between the predicted attribute of the first spatial position information sample output by the head network and the attribute ground truth, and the loss value between the predicted relationship of the multiple second spatial position information samples output by the head network and the relationship ground truth.
[0192] Optionally, the encoder training device further includes a second pre-training module, configured to: after the pre-training of the position encoder is completed, connect the pre-trained position encoder to a large language model connected to a vision encoder; based on at least one of a seventh sample subset corresponding to a map understanding task, an eighth sample subset corresponding to a text address detection task, a ninth sample subset corresponding to an image position detection task, and a tenth sample subset corresponding to a modality alignment task, pre-train the position encoder and the vision encoder, so as to calculate a loss value based on the deviation between the output of the large language model and the ground truth, and adjust the parameters of the position encoder and the vision encoder based on the loss value; the samples in the seventh sample subset corresponding to the map understanding task are masked image samples or text samples, so as to calculate the loss value of the loss function based on the deviation between the content of the masked part predicted by the large language model and the actual content; the samples in the eighth sample subset corresponding to the text address detection task are text samples describing points of interest, so as to calculate the loss value of the loss function based on the deviation between the coordinates of the point of interest predicted by the large language model for the text sample and the actual coordinates of the point of interest; the samples in the ninth sample subset corresponding to the image position detection task are image samples including at least part of a point of interest or image samples collected near the point of interest, so as to calculate the loss value of the loss function based on the deviation between the coordinates of the point of interest predicted by the large language model for the image sample and the actual coordinates of the point of interest; the samples in the tenth sample subset corresponding to the modality alignment task include multiple multimodal sample pairs, each multimodal sample pair corresponding to a point of interest and including at least two of the spatial position information sample, image sample, and text sample corresponding to the point of interest, so as to calculate the loss value of the loss function based on the deviation between the attribute of the point of interest predicted by the large language model for the multiple multimodal sample pairs and the actual attribute of the point of interest.
[0193] The encoder training device provided by the embodiments of the present application can be used to implement the technical solutions of the encoder training method provided by any of the above embodiments of the present application. The implementation principles and technical effects are similar, and will not be elaborated herein.
[0194] Figure 10 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 10 shown, the electronic device of this embodiment may include: at least one processor 1001; and a memory 1002 communicatively connected to the at least one processor; wherein, the memory 1002 stores instructions executable by the at least one processor 1001, and when the instructions are executed by the at least one processor 1001, the electronic device executes the method described in any of the above embodiments.
[0195] Optionally, the memory 1002 can be either independent or integrated with the processor 1001.
[0196] The implementation principles and technical effects of the electronic device provided by this embodiment can be referred to the foregoing embodiments, and will not be elaborated herein.
[0197] The embodiments of the present application also provide a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the method described in any of the foregoing embodiments is implemented.
[0198] The embodiments of the present application also provide a computer program product, including a computer program, which when executed by a processor implements the method described in any of the foregoing embodiments.
[0199] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0200] The integrated modules implemented in the form of software function modules as described above can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in the embodiments of the present application.
[0201] It should be understood that the above-mentioned processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by the execution of the hardware processor, or by the combination of the hardware and software modules in the processor. The memory may include RAM (Random Access Memory), and may also include NVM (Non-Volatile Memory), such as at least one disk memory, and can also be a USB flash drive, a portable hard drive, a read-only memory, a magnetic disk, or an optical disc, etc.
[0202] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0203] An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master control device.
[0204] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element.
[0205] The serial numbers of the embodiments of the present application above are for description only and do not represent the superiority or inferiority of the embodiments.
[0206] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0207] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall equally be included in the patent protection scope of the present application.
Claims
1. A method for generating point-of-interest attributes, characterized in that Including: Obtaining multimodal information of a target point of interest; wherein, the multimodal information includes spatial location information and multimedia information; Encoding the spatial location information based on a pre-trained position encoder to obtain position encoding features; Inputting the position encoding features and the multimedia information into a pre-trained attribute prediction model to obtain the attributes of the target point of interest, where the attributes of the target point of interest include location.
2. The method according to claim 1, wherein The multimedia information includes at least one image and at least one text; the obtaining of the multimodal information of the target point of interest includes: Obtaining at least one image and at least one text of the target point of interest; Based on the at least one image and the at least one text, obtaining the spatial location information.
3. The method according to claim 2, wherein The spatial location information includes the location information of point elements, line elements, and surface elements; the obtaining of the spatial location information based on the at least one image and the at least one text includes: Extracting the coordinates carried by the at least one image and the coordinates carried by the at least one text to obtain the location information of the point elements; Based on the roads and buildings included in the at least one image, obtaining the location information of the line elements and the location information of the surface elements respectively.
4. The method according to claim 3, wherein The obtaining of the location information of the line elements and the location information of the surface elements respectively based on the roads and buildings included in the at least one image includes: Identifying the identifiers of the roads and buildings included in the at least one image; Based on the identified identifiers of the roads and buildings, reading the geographical coordinates of multiple points on the roads included in the at least one image from a first target dataset to obtain the location information of the line elements, and reading the geographical coordinates of the vertices of the buildings included in the at least one image from a second target dataset to obtain the location information of the surface elements; Wherein, the first target dataset includes the identifiers of the roads in a first target geographical area and the geographical coordinates of multiple points thereon, and the second target dataset includes the identifiers of the buildings in a second target geographical area and the geographical coordinates of their vertices, and the target point of interest is located in the first target geographical area and the second target geographical area.
5. The method according to claim 2, wherein The obtaining of at least one image and at least one text of the target point of interest includes: Obtaining multiple initial images of the target point of interest taken; Based on the recognition results of the multiple initial images, screening out the at least one image from the multiple initial images; Based on the attribute information of each waybill in the waybill library, determining the waybills associated with the target point of interest to obtain the at least one text; Wherein, the attribute information of the waybill includes one or more of the waybill address, the name of the sender, and the contact information.
6. The method according to any one of claims 2-5, characterized in that, The attribute prediction model includes a visual encoder, a modality alignment layer, and a large language model layer; The inputting of the position encoding features and the multimedia information into a pre-trained attribute prediction model to obtain the attributes of the target point of interest includes: Inputting the at least one image into the pre-trained visual encoder to obtain image encoding features of the at least one image; Perform modality alignment on the position encoding features and the image encoding features via the modality alignment layer; Input the position encoding features and the image encoding features after modality alignment, as well as the at least one piece of text, into the pre-trained large language model layer to obtain the attributes of the target point of interest.
7. A method for training an encoder, characterized in that, For training the position encoder provided in any one of claims 1-6, the method includes: Obtain a plurality of spatial position information samples and a plurality of multimedia information samples; Based on the plurality of spatial position information samples and the plurality of multimedia information samples, train the pre-trained position encoder connected with an attribute prediction model.
8. The method according to claim 7, characterized in that, The method further includes: Obtain a plurality of pre-training samples; wherein, the plurality of pre-training samples include a pre-training sample set corresponding to an attribute detection task and a pre-training sample set corresponding to a relationship detection task, the pre-training sample set corresponding to the attribute detection task includes a plurality of first spatial position information samples, and the pre-training sample set corresponding to the relationship detection task includes a plurality of second spatial position information samples; Based on the plurality of first spatial position information samples and the plurality of second spatial position information samples, pre-train the position encoder connected with a head network, and adjust the parameters of the position encoder based on the loss value between the predicted attribute of the first spatial position information samples output by the head network and the attribute ground truth, and the loss value between the predicted relationships of the plurality of second spatial position information samples output by the head network and the relationship ground truth.
9. An apparatus for generating point-of-interest attributes, characterized in that Includes: A multi-modal information acquisition module, configured to acquire multi-modal information of a target point of interest; wherein, the multi-modal information includes spatial position information and multimedia information; A position encoding module, configured to encode the spatial position information based on a pre-trained position encoder to obtain position encoding features; An attribute prediction module, configured to obtain the attributes of the target point of interest through a pre-trained large language model, as well as the position encoding features and the multimedia information after modality alignment, and the attributes of the target point of interest include position.
10. A computer program product, characterized in that, Includes a computer program, which when executed by a processor implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Method and device for associating point of interest with waybill, and electronic equipment
CN114443975A
POI (Point of Interest) name generation method, model training method, device, equipment and medium
CN118536476A