Information processing method and device, electronic equipment and storage medium

By generating multimodal features based on points of interest and image data, the problem of inaccurate feature representation in urban areas is solved, achieving more accurate feature representation and higher input data quality for downstream tasks.

CN117131223BActive Publication Date: 2025-12-30BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210529869.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-16
Publication Date
2025-12-30
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

How to use multiple types of urban data to create more accurate feature representations of regions, thereby providing more accurate input data to support downstream tasks.

Method used

POI features and image features are generated based on region-specific point of interest (POI) data and image data. Multimodal features are generated through the adjacency relationship of regions. Feature fusion is performed using an image processing model and a graph neural network with mutual attention mechanism.

Benefits of technology

This improves the accuracy of regional features, providing more accurate input data for downstream processing and thus enhancing the accuracy of subsequent processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117131223B_ABST
    Figure CN117131223B_ABST
Patent Text Reader

Abstract

The present disclosure provides an information processing method and device, electronic equipment and storage medium, and relates to the technical field of artificial intelligence. The specific implementation scheme is: based on point of interest (POI) data and image data corresponding to N regions respectively, generating POI features and image features corresponding to the N regions respectively; N is a positive integer; based on the adjacent relationship of the N regions and the POI features and the image features corresponding to the N regions respectively, generating multi-modal features corresponding to the N regions respectively. The embodiments of the present disclosure can ensure that the multi-modal features of each region obtained finally are more accurate, provide more accurate input data for downstream processing, so that the subsequent downstream processing can also obtain more accurate results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more particularly to the field of artificial intelligence technology. Background Technology

[0002] As global urbanization gradually increases, large-scale, multimodal urban data is becoming increasingly abundant, such as Point of Interest (POI) data on maps and satellite imagery data. This data provides a foundation for multi-faceted representation of urban areas. However, how to use various types of data to more accurately represent the features of urban areas, thereby providing more accurate input data for downstream tasks, has become a problem that needs to be solved. Summary of the Invention

[0003] This disclosure provides an information processing method, apparatus, electronic device, and storage medium.

[0004] According to a first aspect of this disclosure, an information processing method is provided, comprising:

[0005] Based on the Points of Interest (POI) data and image data corresponding to N regions, generate POI features and image features corresponding to the N regions respectively; N is a positive integer;

[0006] Based on the adjacency relationships of the N regions, and the POI features and image features corresponding to the N regions respectively, multimodal features corresponding to the N regions are generated.

[0007] According to a second aspect of this disclosure, an information processing apparatus is provided, comprising:

[0008] The region feature generation module is used to generate POI features and image features corresponding to the N regions based on the POI data and image data corresponding to the N regions respectively; N is a positive integer;

[0009] The first processing module is used to generate multimodal features corresponding to the N regions based on the adjacency relationship of the N regions, the POI features and the image features corresponding to the N regions respectively.

[0010] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0011] At least one processor; and

[0012] The memory is communicatively connected to the at least one processor; wherein,

[0013] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.

[0014] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the aforementioned method.

[0015] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the aforementioned method.

[0016] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description.

[0017] The solution provided in this embodiment can obtain the multimodal features corresponding to the N regions through the adjacency relationships of the N regions, the POI features corresponding to the N regions, and image features. In this way, cross-modal related information can be introduced when characterizing the features of each region. Compared with the method of using only a single modality to characterize the features of each region, this ensures that the final multimodal features of each region are more accurate. Furthermore, it provides more accurate input data for downstream processing, thereby enabling subsequent downstream processing to obtain more accurate results. Attached Figure Description

[0018] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0019] Figure 1 This is a flowchart illustrating an information processing method according to an embodiment of the present disclosure. Figure 1 ;

[0020] Figure 2 This is a flowchart illustrating an information processing method according to an embodiment of the present disclosure. Figure 2 ;

[0021] Figure 3 This is a schematic diagram of the processing flow of the first preset sub-model in an information processing method according to an embodiment of the present disclosure;

[0022] Figure 4 This is a schematic diagram of the composition structure of an information processing apparatus according to an embodiment of the present disclosure;

[0023] Figure 5 This is a schematic diagram of another component structure of an information processing apparatus according to another embodiment of the present disclosure;

[0024] Figure 6This is a block diagram of an electronic device used to implement the information processing method of the embodiments of this disclosure. Detailed Implementation

[0025] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0026] The first aspect of this disclosure provides an information processing method, such as... Figure 1 As shown, it includes:

[0027] S101: Based on the Point of Interest (POI) data and image data corresponding to N regions, generate POI features and image features corresponding to the N regions respectively; N is a positive integer;

[0028] S102: Based on the adjacency relationship of the N regions, and the POI features and image features corresponding to the N regions respectively, generate multimodal features corresponding to the N regions respectively.

[0029] The information processing method provided in this embodiment can be applied to electronic devices, specifically servers or terminal devices.

[0030] The N regions can be N regions within the target range. Determining the N regions from the target range can be done by dividing the target range using a target granularity to obtain the N regions contained within the target range. The target range can be selected based on actual conditions; for example, it can be the area covered by a city, or the area covered by a province, or other target range selection methods, which are not exhaustively listed here. The unit of the target granularity can also be set according to actual conditions; for example, the unit of the target granularity can be meters, kilometers, centimeters, etc., which are not exhaustively listed here. Correspondingly, the target granularity can be determined according to actual conditions; for example, it can be 10 x 10, or 20 x 10, or 50 x 50, etc., again, which are not exhaustively listed. In this embodiment, the number of regions that can be contained within the target range is represented by N. It should be understood that the aforementioned N only represents the number of regions obtained from any single division, and does not mean that the number of regions obtained from different target ranges and / or different target granularities is the same. In other words, the number N of regions obtained from different target ranges and / or different target granularities can be different.

[0031] The POI data of any one of the N regions can include the POI data of one or more POIs within that region. That is, the number of POI data in any one region can be one or more, and this embodiment does not limit it.

[0032] The image data of any one of the N regions can be pre-acquired; for example, the image data of any one region can be satellite image data of that region.

[0033] As can be seen, by adopting the above scheme, the multimodal features corresponding to the N regions can be obtained from the adjacency relationships of the N regions, the POI features corresponding to the N regions, and the image features. In this way, cross-modal related information can be introduced when representing the features of each region. Compared with the method of using only a single modality to represent the features of each region, this ensures that the final multimodal features of each region are more accurate. Furthermore, this provides more accurate input data for downstream processing, thereby enabling subsequent downstream processing to obtain more accurate results.

[0034] In one implementation, generating POI features and image features corresponding to the N regions based on POI data and image data corresponding to the N regions respectively includes:

[0035] Based on the POI data corresponding to the N regions, the distribution features of the i-th POI, the distance features of the i-th POI, and the distribution features of the i-th facility in the i-th region are obtained. Based on the distribution features of the i-th POI, the distance features of the i-th POI, and the distribution features of the i-th facility, the i-th POI feature of the i-th region is generated. Feature extraction is performed on the i-th image data of the i-th region to obtain the i-th image feature of the i-th region. Wherein, i is a positive integer less than or equal to N.

[0036] The aforementioned i-th region can be any one of the N regions. Since the processing method for each of the N regions is the same, it will not be described in detail. It should also be understood that in actual processing, each of the N regions can be treated as the i-th region, and the above processing can be performed in parallel to finally obtain the POI features and image features of each region. For ease of understanding, this embodiment will use the i-th region for specific explanation below.

[0037] It should be noted that in the aforementioned process of generating POI features and image features corresponding to the N regions based on the POI data and image data corresponding to the N regions respectively, the processes of generating POI features corresponding to each of the N regions and generating image features corresponding to each of the N regions can be executed simultaneously or sequentially. This embodiment does not limit their execution order.

[0038] The POI data corresponding to the N regions can include one or more POI data points from all regions within the N regions. Further, each of the N regions can contain one or more POI data points, and the number of POI data points contained in different regions can be the same or different. Any POI data point can include: the POI's name, the POI's location, and the POI's type. The POI's location can be represented by geographic coordinates, such as the latitude and longitude information of the POI's center point. The POI's type can be any one of several preset POI types. These preset POI types can be set according to actual conditions. For example, the preset POI types may include at least one of the following: shopping malls, schools, tourist attractions, hospitals, gas stations, etc. This is only an example; in actual processing, the preset types can include many more types beyond those described above, but this is not exhaustive.

[0039] A processing method for obtaining the distribution characteristics of the i-th POI in the i-th region among the N regions based on the POI data corresponding to the i-th region may include: obtaining the POI types corresponding to one or more POIs contained in the i-th region from the i-th group of POI data in the N regions; determining the POI type distribution characteristics of the i-th region based on the POI types corresponding to the one or more POIs; and using the POI type distribution characteristics of the i-th region as the i-th POI distribution characteristics.

[0040] Specifically, the aforementioned determination of the POI type distribution characteristics of the i-th region based on the POI types corresponding to the one or more POIs can refer to: determining the number of POIs corresponding to each of the multiple preset POI types based on the POI types corresponding to the one or more POIs; dividing the number of POIs corresponding to each preset POI type by the total number of POIs in the i-th region to obtain the proportion of each preset POI type in the i-th region; and generating the POI type distribution characteristics of the i-th region based on the proportion of each preset POI type in the i-th region.

[0041] The step of generating the POI type distribution feature of the i-th region based on the proportion of each preset POI type in the i-th region can be achieved by: concatenating the proportions of each preset POI type in the i-th region to obtain the POI type distribution feature of the i-th region; or by converting the proportions of each preset POI type in the i-th region to a preset base and then concatenating them to obtain the POI type distribution feature of the i-th region. The preset base can be any of binary, octal, or hexadecimal, and is not limited here. It should be understood that the concatenation order can be preset. For example, if it is known in advance that there are three preset POI types, namely preset POI type 1, preset POI type 2, and preset POI type 3, then the concatenation order can be preset to the proportion of preset POI type 1, the proportion of preset POI type 2, and the proportion of preset POI type 3. This is only an illustrative example; in actual processing, the number of preset POI types can be more, and the actual concatenation order is not limited to the above example, but will not be exhaustively listed here.

[0042] For example, the various preset POI types include: shopping malls, schools, hospitals, and tourist attractions. The i-th region contains 10 POIs, of which 8 are shopping malls, 1 is a hospital, and 1 is a tourist attraction. The percentage of each preset POI type in the i-th region is obtained by dividing the number of POIs corresponding to each preset POI type by the total number of POIs in the i-th region. For example, dividing the number of shopping mall POIs (8) by the total number of POIs in the i-th region (10) gives a percentage of 0.8 for shopping malls. Similarly, the percentage of hospitals is 0.1, tourist attractions are 0.1, and schools are 0%. The above-mentioned use of the proportion of each preset POI type in the i-th region as the POI type distribution feature of the i-th region can refer to: concatenating the proportions of each preset POI type in the i-th region to obtain the POI type distribution feature of the i-th region.

[0043] Another processing method for obtaining the distribution characteristics of the i-th POI in the i-th region among the N regions, based on the POI data corresponding to the N regions respectively, may include:

[0044] From the POI data of the ith group in the ith region of the N regions, obtain the POI types corresponding to one or more POIs contained in the ith region; based on the POI types corresponding to the one or more POIs, determine the POI type distribution characteristics of the ith region;

[0045] From K sets of POI data in K reference regions containing the i-th region, obtain the POI types of one or more reference POIs contained in the K reference regions; based on the POI types of the one or more reference POIs, determine the distribution characteristics of the reference POI types in the K reference regions; K is an integer greater than or equal to 1;

[0046] The POI type distribution characteristics of the i-th region and the reference POI type distribution characteristics of the K reference regions are concatenated to obtain the i-th POI distribution characteristics.

[0047] In this processing method, the aforementioned processing of determining the POI type distribution characteristics of the i-th region and the processing of determining the reference POI type distribution characteristics of the K reference regions can be processed in parallel or executed sequentially; the execution order is not limited here.

[0048] The K reference regions encompassing the i-th region refer to the i-th region itself and the K-1 reference regions spatially adjacent to it. The number of K can be fixed; for example, K can be equal to 9, or K can be equal to 25, etc. Not all possible values ​​are exhaustively listed here. For instance, if K equals 9, the 8 reference regions spatially adjacent to the i-th region can be located to the north, east, south, west, northeast, southeast, southwest, and northwest of the i-th region, respectively.

[0049] The process of obtaining the POI types corresponding to one or more POIs contained in the i-th region from the i-th group of POI data in the N regions, and determining the POI type distribution characteristics of the i-th region based on the POI types corresponding to the one or more POIs, is the same as the processing method in the previous embodiment and will not be repeated.

[0050] The step of obtaining the POI type of one or more reference POIs contained in the K reference regions from the K sets of POI data containing the i-th region can refer to obtaining the POI type of each POI among all the reference POIs contained in the K reference regions from the K sets of POI data containing the i-th region; or it can refer to obtaining the POI type of one or more reference POIs contained in each of the K reference regions from the K sets of POI data containing the i-th region.

[0051] Accordingly, determining the reference POI type distribution characteristics of the K reference regions based on the POI types of the one or more reference POIs may include: determining the number of reference POIs corresponding to each of the multiple preset POI types based on the POI types of each POI among all the reference POIs included in the K reference regions; dividing the number of reference POIs corresponding to each preset POI type by the total number of all reference POIs to obtain the proportion of each preset POI type in the K reference regions; and generating the reference POI type distribution characteristics of the K reference regions based on the proportion of each preset POI type in the K reference regions.

[0052] For example, the various preset POI types include: shopping malls, schools, hospitals, and tourist attractions. In the K reference areas encompassing the i-th region, there are a total of 20 reference POIs. Of these, 10 are shopping malls, 4 are hospitals, and 6 are tourist attractions. Dividing the number of shopping malls (10) by the total number of reference POIs in the K reference areas (20) yields a percentage of 0.5. Similarly, the percentage of hospitals is 0.2%, tourist attractions are 0.3%, and schools are 0%. Concatenating the percentages of each preset POI type in the K reference areas yields the distribution characteristics of the reference POI types in those K reference areas.

[0053] The step of generating the reference POI type distribution feature of the K reference regions based on the proportion of each preset POI type in the K reference regions can be: concatenating the proportions of each preset POI type in the K reference regions to obtain the reference POI type distribution feature of the K reference regions; or, converting the proportions of each preset POI type in the K reference regions to a preset base and then concatenating them to obtain the reference POI type distribution feature of the K reference regions.

[0054] The step of concatenating the POI type distribution features of the i-th region and the reference POI type distribution features of the K reference regions to obtain the i-th POI distribution feature can refer to concatenating the POI type distribution features of the i-th region and the reference POI type distribution features of the K reference regions according to a preset concatenation order. The preset concatenation order can be set according to actual conditions and is not limited thereto.

[0055] The processing method for obtaining the distance feature of the i-th POI in the i-th region among the N regions based on the POI data corresponding to the N regions may include: determining the location and POI type corresponding to the one or more POIs contained in the N regions; determining the shortest distance between the i-th region and the POIs of each of the multiple preset POI types based on the center location of the i-th region and the location of the POIs in each of the multiple preset POI types; and generating the distance feature of the i-th POI in the i-th region based on the shortest distance between the i-th region and the POIs of each of the multiple preset POI types.

[0056] The "one or more POI data" within the N regions can refer to all POI data contained within all regions of the N regions. As previously explained, any POI data can include the POI's location, name, and POI type. The multiple preset POI types have also been described above and will not be repeated here.

[0057] The center position of the i-th region can be the coordinates of the center point of the i-th region, specifically represented by latitude and longitude; however, all possible representations are not exhaustively listed here. Correspondingly, determining the shortest distance between the i-th region and the POIs of each of the multiple preset POI types based on the center position of the i-th region and the positions of POIs in each of the multiple preset POI types can include: determining the shortest distance between the center position of the i-th region and one or more POIs in the k-th preset POI type based on the center position of the i-th region and the positions of one or more POIs in the k-th preset POI type; k is a positive integer. The k-th preset POI type can be any one of the multiple preset POI types. The processing for different preset POI types among the multiple preset POI types is the same as the processing for the k-th preset POI type, and therefore will not be elaborated upon.

[0058] For example, suppose there are three preset POI types: Preset POI Type 1, Preset POI Type 2, and Preset POI Type 3. Among all the POI data contained in the N regions, there are 30 POIs of Preset POI Type 1. Based on the center position of the i-th region and the positions of the 30 POIs of Preset POI Type 1, the distance between the center position of the i-th region and the 30 POIs of Preset POI Type 1 is calculated. The shortest distance among these distances is selected as the shortest distance between the i-th region and the POIs of Preset POI Type 1. The processing methods for other preset POI types 2 and 3 are the same as in the previous example, and therefore will not be repeated.

[0059] The step of generating the i-th POI distance feature of the i-th region based on the nearest distance between the i-th region and POIs of each of the multiple preset POI types can be achieved by concatenating the nearest distances between the i-th region and POIs of each of the multiple preset POI types. The concatenation can be performed based on a preset concatenation order, which can be set according to actual conditions. For example, if there are three preset POI types, the preset concatenation order can be set based on the multiple preset POI types, such as setting the preset concatenation order to the order of preset POI type 1, preset POI type 2, and preset POI type 3.

[0060] The processing method for obtaining the distribution characteristics of the i-th facility in the i-th region among the N regions based on the POI data corresponding to the N regions may include:

[0061] Based on the POI data contained in the N regions, determine the POI type of each POI in one or more POIs contained within a specified range; wherein, the specified range includes the i-th region;

[0062] Based on the POI type of each POI in one or more POIs included in the specified range, determine whether the specified range contains all the specified POI types;

[0063] If it is determined that the specified range contains all the specified POI types, the i-th facility distribution characteristic of the i-th region is determined to be a first value; if it is determined that the specified range does not contain all the specified POI types, the i-th facility distribution characteristic of the i-th region is determined to be a second value.

[0064] The specified range can be: containing only the i-th region. Alternatively, the specified range can be: the entire area covered by extending outwards from the center of the i-th region by a specified distance. The specified distance can be set according to actual conditions, such as 1 kilometer, longer, or shorter; it is not limited here. The entire area covered by extending outwards from the center of the i-th region by a specified distance can be: the entire area covered by extending outwards from the edge of the i-th region by a specified distance. For example, assuming the i-th region is an original square region of 128 meters by 128 meters centered at longitude a1 and latitude b1, then by extending outwards (i.e., in the opposite direction to the center point) by a specified distance from each of the four sides of the i-th region as starting points, a new square region can be obtained, and this new square region is the aforementioned specified range.

[0065] Wherein, the first value and the second value are different. For example, the first value can be 1 and the second value can be 0; or the first value can be 0 and the second value can be 1; or the first value can be 01 and the second value can be 10, or vice versa. It should be understood that the foregoing is only an illustrative example, and all differences between the first value and the second value are within the scope of protection of this embodiment.

[0066] The specified POI types can be determined based on actual circumstances. These specified POI types can be at least one of the aforementioned multiple preset POI types. For example, the specified POI types can be all of the aforementioned multiple preset POI types. Or, for example, if the aforementioned preset POI types include preset POI type 1, preset POI type 2, and preset POI type 3, the specified POI types can include preset POI type 1 and preset POI type 2. It should be understood that this is merely an illustrative example; in actual processing, as long as the specified POI types are at least one of the multiple preset POI types, it is within the scope of protection of this embodiment, and no exhaustive list is provided.

[0067] Through the above processing, based on the POI data corresponding to the N regions respectively, the distribution features of the i-th POI, the distance features of the i-th POI, and the distribution features of the i-th facility in the i-th region can be obtained. Based on this, generating the i-th POI feature of the i-th region based on the i-th POI distribution features, the i-th POI distance features, and the i-th facility distribution features can include: concatenating the i-th POI distribution features, the i-th POI distance features, and the i-th facility distribution features to obtain the i-th POI feature of the i-th region. This concatenation process can be implemented based on a preset concatenation order of the POI features. This preset concatenation order of the POI features can be set according to actual conditions, for example, setting the preset concatenation order of the POI features to the order of the i-th POI distribution features, the i-th POI distance features, and the i-th facility distribution features; or, setting the preset concatenation order of the POI features to the order of the i-th POI distance features, the i-th POI distribution features, and the i-th facility distribution features. This section does not exhaustively list all possible cases of the preset splicing order for this POI feature.

[0068] The i-th image data of the aforementioned i-th region can specifically be the i-th satellite image data of that i-th region. The i-th satellite image data can be a 256×256 (unit can be pixels) RGB image. Regarding the method of obtaining satellite image data of any region, it can be pre-acquired and saved, or it can be extracted from the database of the satellite ground station when needed; this embodiment does not limit this.

[0069] The step of extracting features from the i-th image data of the i-th region to obtain the i-th image feature of the i-th region can be specifically performed by inputting the i-th image data of the i-th region into an image processing model for feature extraction, and obtaining the i-th image feature of the i-th region output by the image processing model.

[0070] The i-th image feature can be a multi-dimensional vector, and the specific number of dimensions of the vector can be related to the image processing model. For example, the i-th image feature can be a 4096-dimensional vector.

[0071] The image processing model can be derived from an original image processing model. For example, the image processing model can be obtained by removing the last one or more connected layers from the original image processing model. The number of connected layers removed can be related to the actual original image processing model. For instance, the original image processing model can be the Visual Geometry Group (VGG) 16 model. The image processing model used in this embodiment is obtained by removing the last two fully connected layers of the VGG16 model. This embodiment uses an image processing model to process image data to extract image features. This avoids the problem of overfitting that may occur when directly using pixel-level image data for subsequent target model training due to its high dimensionality.

[0072] As explained above, the region map corresponding to the target range can be represented as G = (V, E, A, X), where X represents the region feature matrix. Through the aforementioned processing, the POI features and image features of each region can be obtained. Correspondingly, the region features of each node in the region feature matrix of the region map corresponding to the target range can be constructed based on the POI features and image features of each region. For example, the POI features and image features of the i-th region can be set at the position of the region feature matrix corresponding to the i-th region.

[0073] As can be seen, by adopting the above scheme, the POI data corresponding to the N regions can be processed to obtain the POI features of the i-th region among the N regions, and feature extraction can be performed based on the image data of the i-th region among the N regions to obtain the image features of the i-th region. In this way, more accurate multimodal features can be obtained when using the POI features and image features of each region to generate multimodal features.

[0074] In one embodiment, the method may further include:

[0075] Based on the spatial characteristics of the N regions, spatially adjacent regions corresponding to the N regions are determined, and connected adjacent regions corresponding to the N regions are determined based on preset road network connectivity data.

[0076] The spatially adjacent regions corresponding to the N regions and the connected adjacent regions corresponding to the N regions are respectively regarded as the adjacent regions corresponding to the N regions.

[0077] Based on the adjacent regions corresponding to the N regions, the adjacency relationships of the N regions are generated.

[0078] The spatial characteristics of the N regions specifically refer to the geographical location of each region within the N regions. For example, the geographical location of each region within the N regions can be represented by the latitude and longitude of its center.

[0079] The step of determining the spatially adjacent regions corresponding to each of the N regions based on their spatial characteristics may include: determining M spatially adjacent regions that are spatially adjacent to the i-th region based on the geographical location of each of the N regions and the geographical location of the i-th region; where i is a positive integer less than or equal to N, and M can be a positive integer. It should be understood that the i-th region can be any one of the N regions, and the above method can be used to obtain M spatially adjacent regions for any one of the N regions; however, it will not be elaborated upon in detail.

[0080] In one example, M can be equal to 8, that is, the M spatially adjacent regions that are adjacent to the i-th region can be the 8 spatially adjacent regions that are adjacent to the i-th region. Specifically, the i-th region can be used as the intermediate region to obtain 3×3 regions (i.e., 9 regions). Among these 9 regions, excluding the i-th region, the remaining 8 regions are the 8 regions that are spatially adjacent to the i-th region.

[0081] The step of determining the connected adjacent regions corresponding to the N regions based on preset road network connectivity data may specifically include: determining the target road network node corresponding to the i-th region among the N regions based on the position of each preset road network node in the preset road network connectivity data; determining one or more candidate road network nodes that the target road network node can connect to through one or more preset road network edges; determining one or more candidate road network nodes within a preset connectivity range from the one or more candidate road network nodes, starting from the target road network node; and taking the one or more regions corresponding to the one or more candidate road network nodes within the preset connectivity range as one or more connected adjacent regions of the i-th region.

[0082] The preset road network connectivity data can also be referred to as road network data, which may include multiple preset road network nodes and multiple preset road network edges; each of the multiple preset road network edges is used to connect two preset road network nodes. The preset connectivity range can be set according to actual conditions, for example, it can be set to 5 preset road network edges, or more or fewer, which is not limited here. That is to say, the aforementioned determination of one or more candidate road network nodes within the preset connectivity range from the one or more candidate road network nodes, starting from the target road network node, can include: determining one or more candidate road network nodes within 5 preset road network edges from the one or more candidate road network nodes, starting from the target road network node.

[0083] The aforementioned processing establishes interconnected adjacent areas by introducing preset road network connectivity data. This is because preset road network connectivity data, as the most core part of the urban transportation system, plays a crucial role in capturing cross-regional functional correlations. While the multiple areas divided using this embodiment may not be spatially adjacent, they can be connected through preset road network connectivity data, thus allowing for a more comprehensive acquisition of the correlations between various areas.

[0084] The phrase "taking the spatially adjacent regions and the connected adjacent regions corresponding to the N regions as the adjacent regions of the N regions" can mean that both the spatially adjacent region and the connected adjacent region of the i-th region are taken as the adjacent regions of the i-th region. Since the i-th region is any one of the N regions, the processing of each region within these N regions is the same as described above, and therefore will not be elaborated upon further.

[0085] The step of generating the adjacency relationship of the N regions based on their respective adjacent regions can be achieved by using the adjacent regions of the N regions as elements in the region adjacency matrix to obtain the region adjacency matrix representing the adjacency relationship of the N regions.

[0086] Based on the solutions provided in the foregoing embodiments, the target area can be divided into N regions, which means the target area can be gridded. If the target area is a city, then the city can be gridded. After gridding the target area, this embodiment can construct a region map (or city region map) corresponding to the target area. This region map can be used to describe the characteristics of the regions and the relationships between them. In this region map, each of the N regions can be considered a node. Let's assume the region map is represented as G = (V, E, A, X), where V represents the set of nodes composed of the regions, X represents the region feature matrix, E represents edges, and A represents the region adjacency matrix, i.e., the adjacency relationships between regions.

[0087] Specifically, in the region graph corresponding to the target range, V can be a set of identifiers of nodes corresponding to each region. X in the region graph corresponding to the target range can be a matrix composed of features corresponding to each region. These features can include POI features and image features corresponding to each region. For example, the features corresponding to V1 are {POI feature 1, image feature 1}. The features corresponding to V1 can be an element of X (region feature matrix), and so on, with the features corresponding to each region forming the aforementioned region feature matrix. E in the region graph corresponding to the target range can represent the edges that may exist among the nodes corresponding to each region. For example, there may be an edge between V1 and V2. A in the region graph corresponding to the target range can refer to the region adjacency matrix representing the adjacency relationships of the N regions, constructed through the aforementioned embodiments; for example, in this region adjacency matrix, A... 12 =1 indicates that regions V1 and V2 are adjacent, A 13 =0 indicates that regions V1 and V3 are not adjacent. It should be noted that the adjacency relationship between any two regions can include spatial adjacency and connected adjacency. That is, if two regions are spatially adjacent, then one region is a spatially adjacent region of the other region; if two regions are connectedly adjacent, then one region is a connected adjacent region of the other region. It should be understood that in the subsequent schemes of this embodiment, both the connected adjacent regions and the spatially adjacent regions of a region are referred to as the adjacent regions of that region. It should also be understood that any region can have one or more adjacent regions, and this embodiment does not limit the number of adjacent regions that a region can have.

[0088] As can be seen, by adopting the above scheme, the spatial adjacent regions and connected adjacent regions of a region can be determined from the spatial characteristics of the region and the preset road network connectivity data, so as to more accurately represent the adjacent relationship of each region, and the subsequent processing can extract more comprehensive and accurate adjacent regions, thereby ensuring that the generated multimodal features are more accurate.

[0089] In one implementation, generating multimodal features corresponding to the N regions based on the adjacency relationships of the N regions and the POI features and image features corresponding to the N regions respectively includes: inputting the adjacency relationships of the N regions and the POI features and image features corresponding to the N regions respectively into a first preset sub-model to obtain the multimodal features corresponding to the N regions output by the first preset sub-model.

[0090] Here, the first preset sub-model can be a graph neural network based on mutual attention mechanism, or it can be other types of preset networks. We will not exhaustively list all possible architectures here.

[0091] In the scheme provided in the foregoing embodiments, a region map corresponding to the target range can be obtained. This region map can include the adjacency relationships of the N regions, and the POI features and image features corresponding to each of the N regions. Correspondingly, inputting the adjacency relationships of the N regions, and the POI features and image features corresponding to each of the N regions into a first preset sub-model to obtain the multimodal features corresponding to each of the N regions output by the first preset sub-model can be: inputting the region map corresponding to the target range into a first preset sub-model to obtain the multimodal features corresponding to each of the N regions output by the first preset sub-model. In the first preset sub-model, processing can be performed based on the adjacency relationships of the N regions included in the region map corresponding to the target range, and the POI features and image features corresponding to each of the N regions, to obtain the multimodal features corresponding to each of the N regions.

[0092] As can be seen, by adopting the above scheme, the adjacency relationship of the aforementioned N regions, as well as the POI features and image features corresponding to the N regions, can be processed by the first preset sub-model to obtain multimodal features, thereby ensuring efficient and accurate input information for downstream processing and improving overall processing efficiency and accuracy.

[0093] In one embodiment, the method may further include: inputting the multimodal features corresponding to the N regions into a second preset sub-model to obtain the prediction information output by the second preset sub-model;

[0094] Based on the prediction information, a loss function is determined, and the first preset sub-model and the second preset sub-model are updated by backpropagation based on the loss function;

[0095] If the first preset sub-model and the second preset sub-model have completed training, the trained first target sub-model and the second target sub-model are obtained.

[0096] Here, the second preset sub-model can be related to the training objective to be obtained in this training. For example, if the training aims to obtain a multi-class classification result, meaning the training is for a multi-class classification task, then the second preset sub-model can be a regression model. Correspondingly, the type of loss function that can be used varies depending on the selected second preset sub-model; taking the aforementioned multi-class classification task as an example, the type of loss function used can be the cross-entropy loss function. It should be understood that the specific training objective can be determined according to the actual situation. Different second preset sub-models can be selected for different training objectives, and correspondingly, the type of loss function used can also vary depending on the type of second preset sub-model selected. This embodiment exhaustively lists all possibilities.

[0097] The aforementioned inputting the multimodal features corresponding to the N regions into the second preset sub-model to obtain the prediction information output by the second preset sub-model can refer to inputting the multimodal features corresponding to the N regions into the second preset sub-model, processing the multimodal features corresponding to the N regions in the second preset sub-model, and obtaining the prediction information output by the second preset sub-model.

[0098] The specific type of the aforementioned loss function can be related to the type of the second preset sub-model selected in this study. The loss function can be one or more, such as the first loss function and the second loss function being weighted and summed to form the loss function, or there can be only one loss function. Here, we will not exhaust all possible types.

[0099] Determining how the first and second preset sub-models complete training may include: the number of training iterations reaching a preset threshold; the loss function no longer changing or falling below a specified value, etc. The preset threshold can be set according to actual conditions, such as 100 iterations, 50 iterations, etc. This embodiment does not exhaustively list all possible convergence conditions for determining whether the first and second preset sub-models have completed training.

[0100] It is important to understand that after obtaining the aforementioned first target sub-model and second target sub-model, the aforementioned first target sub-model and second target sub-model can be used for actual prediction processing. The specific processing is related to the training objective determined during the training process. Taking the multi-class classification task as an example, the actual application can be as follows: output the adjacency relationship of multiple regions, the POI features and image features corresponding to the multiple regions respectively to the first target sub-model to obtain the multimodal features corresponding to the multiple regions output by the first target sub-model; input the multimodal features corresponding to the multiple regions respectively into the second target sub-model to obtain the classification result output by the second target sub-model.

[0101] Combination Figure 2 The aforementioned solution provided in this embodiment will be illustrated by example:

[0102] S201: Obtain POI data and image data corresponding to N regions respectively, and obtain preset road network connectivity data;

[0103] S202: Based on the POI data and image data corresponding to the N regions respectively, generate the POI features and image features corresponding to the N regions respectively;

[0104] S203: Based on the preset road network connectivity data and the spatial characteristics of the N regions, generate the adjacency relationship of the N regions.

[0105] Specifically, this step may include: determining the spatially adjacent regions corresponding to the N regions based on the spatial characteristics of the N regions, and determining the connected adjacent regions corresponding to the N regions based on preset road network connectivity data; taking the spatially adjacent regions and the connected adjacent regions corresponding to the N regions as the adjacent regions corresponding to the N regions; and generating the adjacency relationship of the N regions based on the adjacent regions corresponding to the N regions.

[0106] The processing order of the aforementioned S202 and S203 can be parallel, or S202 can be executed first and then S203, or S203 can be executed first and then S202. This embodiment does not limit it.

[0107] In addition, it should be noted that after completing the processing of S202 and S203, a region map corresponding to the target range can be obtained. The detailed description of the region map corresponding to the target range has been described in the foregoing embodiments and will not be repeated here.

[0108] S204: Input the adjacency relationship of the N regions, the POI features and the image features corresponding to the N regions respectively into the first preset sub-model to obtain the multimodal features corresponding to the N regions respectively output by the first preset sub-model.

[0109] S205: Input the multimodal features corresponding to the N regions into the second preset sub-model to obtain the prediction information output by the second preset sub-model.

[0110] S206: Determine the loss function based on the prediction information, and update the first preset sub-model and the second preset sub-model based on the backpropagation of the loss function.

[0111] After completing the aforementioned S206, the method may further include: if it is determined that the first preset sub-model and the second preset sub-model have completed training, obtaining the trained first target sub-model and the second target sub-model.

[0112] As can be seen, by adopting the above scheme, the second preset sub-model used in the downstream task after the first preset sub-model has completed processing can be determined. This is achieved by further processing the multimodal features obtained from the first preset sub-model until the trained first and second target sub-models are obtained. Since the first preset sub-model provides more accurate multimodal features as input to the second preset sub-model, the second preset sub-model can obtain more accurate outputs. Therefore, after training the first and second target sub-models, they can be used in practical applications to obtain more accurate output information.

[0113] In one implementation, the step of inputting the adjacency relationships of the N regions, and the POI features and image features corresponding to the N regions respectively into a first preset sub-model to obtain the multimodal features corresponding to the N regions output by the first preset sub-model includes:

[0114] The adjacency relationships of the N regions, as well as the POI features and image features corresponding to the N regions, are input into the first preset sub-model. Based on the adjacency relationships of the N regions, as well as the POI features and image features corresponding to the N regions, the first cross-modal vector and the second cross-modal vector corresponding to the N regions are generated in the first preset sub-model.

[0115] In the first preset sub-model, the first cross-modal vector and the second cross-modal vector corresponding to the N regions are processed to obtain the multimodal features corresponding to the N regions output by the first preset sub-model.

[0116] Here, the step of generating the first cross-modal vector and the second cross-modal vector corresponding to the N regions in the first preset sub-model based on the adjacency relationship of the N regions and the POI features and image features corresponding to the N regions respectively can specifically include:

[0117] In the first preset sub-model, the adjacent regions corresponding to the N regions are determined based on the adjacency relationship of the N regions; based on the POI features of the adjacent regions corresponding to the N regions and the image features corresponding to the N regions, the first cross-modal vector corresponding to the N regions is generated; and based on the image features of the adjacent regions corresponding to the N regions and the POI features corresponding to the N regions, the second cross-modal vector corresponding to the N regions is generated.

[0118] The step of processing the first cross-modal vectors and the second cross-modal vectors corresponding to the N regions in the first preset sub-model to obtain the multimodal features corresponding to the N regions output by the first preset sub-model may include: concatenating the first cross-modal vectors and the second cross-modal vectors corresponding to the N regions in the first preset sub-model to obtain the multimodal features corresponding to the N regions output by the first preset sub-model.

[0119] As can be seen, by adopting the above scheme, the adjacent regions corresponding to the N regions, as well as the POI features and image characteristics corresponding to the N regions, can be fused in the first preset model to obtain multiple cross-modal vectors. Finally, multimodal features are generated based on these multiple cross-modal vectors. In this way, cross-modal vectors can be used to represent each region, ensuring that when training downstream models using multimodal features, the trained first target sub-model and second target sub-model can perform more accurate prediction processing.

[0120] In one implementation, generating a first cross-modal vector and a second cross-modal vector corresponding to each of the N regions in the first preset sub-model based on the adjacency relationships of the N regions and the POI features and image features corresponding to the N regions respectively includes:

[0121] In the first preset sub-model, the neighboring regions of the i-th region are obtained from the adjacency relationships of the N regions; in the first preset sub-model, the first cross-modal vector of the i-th region is generated based on the POI features of the neighboring regions of the i-th region and the i-th image features of the i-th region; and in the first preset sub-model, the second cross-modal vector corresponding to the i-th region is generated based on the image features of the neighboring regions of the i-th region and the i-th POI features of the i-th region.

[0122] As described in the foregoing embodiments, the adjacency relationships of the N regions include the adjacent regions of each of the N regions. Accordingly, obtaining the adjacent region of the i-th region from the adjacency relationships of the N regions in the first preset sub-model can mean directly extracting the adjacent region of the i-th region from the adjacency relationships of the N regions in the first preset sub-model. As also described above, for each of the N regions, corresponding POI features and image features are obtained. Accordingly, the POI features and image features corresponding to the adjacent regions of the i-th region can be obtained in the first preset sub-model.

[0123] It should be understood that the number of neighboring regions of the i-th region can be one or more. Correspondingly, the step of generating the first cross-modal vector of the i-th region based on the POI features of the neighboring regions of the i-th region and the i-th image feature of the i-th region in the first preset sub-model, and generating the second cross-modal vector corresponding to the i-th region based on the image features of the neighboring regions of the i-th region and the i-th POI feature of the i-th region, can include: generating the first cross-modal vector of the i-th region based on the POI features of each neighboring region in one or more neighboring regions of the i-th region and the i-th image feature of the i-th region in the first preset sub-model, and generating the second cross-modal vector corresponding to the i-th region based on the image features of each neighboring region in one or more neighboring regions of the i-th region and the i-th POI feature of the i-th region in the first preset sub-model. Furthermore, since the i-th region can be any one of the N regions, the above method can be used for each of the N regions to obtain the first cross-modal vector and the second cross-modal vector corresponding to each region, but it will not be elaborated one by one.

[0124] As can be seen, by adopting the above scheme, the first cross-modal vector of each region can be obtained based on the POI features of its neighboring regions and the image features of the region itself. Similarly, the second cross-modal vector of each region can be obtained based on the image features of its neighboring regions and the POI features of the region itself. This allows each region to fuse features from its neighboring regions and represent each region with cross-modal vectors, ensuring that the multimodal features generated subsequently based on these cross-modal vectors are more accurate. Therefore, using these multimodal features to train downstream models enables the trained first and second target sub-models to perform more accurate predictions.

[0125] In one implementation, generating the first cross-modal vector of the i-th region based on the POI features of the neighboring regions of the i-th region and the i-th image features of the i-th region in the first preset sub-model includes: obtaining a first cross-modal weight value of the i-th region based on the POI features of the neighboring regions of the i-th region and the i-th image features of the i-th region in the first preset sub-model; and obtaining the first cross-modal vector of the i-th region based on the first cross-modal weight value of the i-th region and the POI features of the neighboring regions of the i-th region.

[0126] The step of generating the second cross-modal vector corresponding to the i-th region in the first preset sub-model based on the image features of the neighboring regions of the i-th region and the i-th POI feature of the i-th region includes: obtaining the second cross-modal weight value of the i-th region based on the image features of the neighboring regions of the i-th region and the i-th POI feature of the i-th region in the first preset sub-model; and obtaining the second cross-modal vector of the i-th region based on the second cross-modal weight value of the i-th region and the image features of the neighboring regions of the i-th region.

[0127] Specifically, obtaining the first cross-modal weight value of the i-th region in the first preset sub-model based on the POI features of the adjacent regions of the i-th region and the i-th image feature of the i-th region may include: obtaining the first cross-modal weight value between the i-th region and the j-th adjacent region based on the j-th POI feature of the j-th adjacent region of the i-th region and the i-th image feature of the i-th region; where j is a positive integer less than or equal to N.

[0128] The i-th region may include one or more neighboring regions, and the j-th neighboring region may be any one of the one or more neighboring regions of the i-th region. It should be understood that the above processing can be applied to each neighboring region of the i-th region to obtain the first cross-modal weight value between the i-th region and each neighboring region, but it will not be elaborated on one by one.

[0129] Specifically, in the first preset sub-model, based on the j-th POI feature of the j-th adjacent region of the i-th region and the i-th image feature of the i-th region, the first cross-modal weight value between the i-th region and the j-th adjacent region is obtained. This can include: obtaining a transformed j-th POI feature based on a first POI feature transformation matrix and the j-th POI feature of the j-th adjacent region in the first preset sub-model; obtaining a transformed i-th image feature based on a first image transformation matrix and the i-th image feature of the i-th region in the first preset sub-model; concatenating the transformed j-th POI feature and the transformed i-th image feature in the first preset sub-model to obtain a first vector; and multiplying the first vector by a preset first parameter vector and the first vector, and then using a target activation function to calculate the first cross-modal weight value between the i-th region and the j-th adjacent region.

[0130] The parameters in the first image feature transformation matrix can be adjusted through backpropagation of the loss function during subsequent training with the second preset sub-model. Similarly, the parameters in the first POI feature transformation matrix can also be adjusted through backpropagation of the loss function during subsequent training with the second preset sub-model. The preset first parameter vector can also be adjusted through backpropagation of the loss function during subsequent training with the second preset sub-model.

[0131] The step of obtaining the first cross-modal vector of the i-th region based on the first cross-modal weight value of the i-th region and the POI features of the adjacent regions of the i-th region in the first preset sub-model includes: determining the normalized first cross-modal weight value between the i-th region and each of the adjacent regions based on the first cross-modal weight value between the i-th region and each of the adjacent regions in the first preset sub-model; and determining the first cross-modal vector of the i-th region based on the normalized first cross-modal weight value between the i-th region and each of the adjacent regions in the first preset sub-model.

[0132] The step of determining the normalized first cross-modal weight value between the i-th region and each of the adjacent regions based on the first cross-modal weight value between the i-th region and each of the adjacent regions in the first preset sub-model may include: dividing the first cross-modal weight value between the i-th region and the j-th adjacent region by the sum of the first cross-modal weight values ​​of each of the adjacent regions in the first preset sub-model to obtain the normalized first cross-modal weight value between the i-th region and the j-th adjacent region.

[0133] The step of determining the first cross-modal vector of the i-th region in the first preset sub-model based on the normalized first cross-modal weight values ​​between the i-th region and each region, and the POI features of each adjacent region, may include: calculating each first cross-modal sub-vector of the i-th region in the first preset sub-model based on the normalized first cross-modal weight values ​​between the i-th region and each region, and the POI features of each adjacent region; summing the each first cross-modal sub-vector of the i-th region, and then calculating the first cross-modal vector of the i-th region based on the target activation function.

[0134] The step of determining the first cross-modal vector of the i-th region in the first preset sub-model based on the normalized first cross-modal weight value between the i-th region and each region, and the POI features of each adjacent region, may include: multiplying the normalized first cross-modal weight value between the i-th region and the j-th adjacent region, the POI features of the j-th adjacent region of the i-th region, and the first POI feature transformation matrix in the first preset sub-model to obtain the j-th first cross-modal sub-vector of the i-th region.

[0135] The aforementioned method of obtaining the second cross-modal weight value of the i-th region based on the image features of the neighboring regions of the i-th region and the i-th POI feature of the i-th region in the first preset sub-model may include: obtaining the second cross-modal weight value between the i-th region and the j-th neighboring region based on the j-th image feature of the j-th neighboring region of the i-th region and the i-th POI feature of the i-th region in the first preset sub-model; where j is a positive integer. It should be understood that the above processing can be applied to each neighboring region of the i-th region to obtain the second cross-modal weight value between the i-th region and each neighboring region, but it will not be elaborated upon individually.

[0136] Specifically, in the first preset sub-model, based on the j-th image feature of the j-th neighboring region of the i-th region and the i-th POI feature of the i-th region, the second cross-modal weight value between the i-th region and the j-th neighboring region is obtained. This can include: obtaining a transformed j-th image feature based on a second image feature transformation matrix and the j-th image feature of the j-th neighboring region in the first preset sub-model; obtaining a transformed i-th POI feature based on a second POI feature transformation matrix and the i-th POI feature of the i-th region; concatenating the transformed j-th image feature and the transformed i-th POI feature to obtain a second vector; and multiplying the second vector by a preset second parameter vector and the second vector, then using a target activation function to calculate the second cross-modal weight value between the i-th region and the j-th neighboring region.

[0137] The parameters in the second image feature transformation matrix can be adjusted through backpropagation of the loss function during subsequent training with the second preset sub-model. Similarly, the parameters in the second POI feature transformation matrix can also be adjusted through backpropagation of the loss function during subsequent training with the second preset sub-model. The preset second parameter vector can also be adjusted through backpropagation of the loss function during subsequent training with the second preset sub-model.

[0138] The step of obtaining the second cross-modal vector of the i-th region based on the second cross-modal weight value of the i-th region and the image features of the adjacent regions of the i-th region in the first preset sub-model includes: determining the normalized second cross-modal weight value between the i-th region and each of the adjacent regions based on the second cross-modal weight value between the i-th region and each of the adjacent regions in the first preset sub-model; and determining the second cross-modal vector of the i-th region based on the normalized second cross-modal weight value between the i-th region and each of the adjacent regions in the first preset sub-model.

[0139] The step of determining the normalized second cross-modal weight value between the i-th region and each of the adjacent regions based on the second cross-modal weight value between the i-th region and each of the adjacent regions in the first preset sub-model may include: dividing the second cross-modal weight value between the i-th region and the j-th adjacent region by the sum of the second cross-modal weight values ​​of each of the adjacent regions in the first preset sub-model to obtain the normalized second cross-modal weight value between the i-th region and the j-th adjacent region.

[0140] The step of determining the second cross-modal vector of the i-th region in the first preset sub-model based on the normalized second cross-modal weight values ​​between the i-th region and each other region, and the image features of each adjacent region, may include: calculating each second cross-modal sub-vector of the i-th region in the first preset sub-model based on the normalized second cross-modal weight values ​​between the i-th region and each other region, and the image features of each adjacent region; and summing the each second cross-modal sub-vector of the i-th region in the first preset sub-model, and then calculating the second cross-modal vector of the i-th region based on the target activation function.

[0141] The step of calculating each second cross-modal sub-vector of the i-th region based on the normalized second cross-modal weight value between the i-th region and each other region, and the image features of each adjacent region in the first preset sub-model may include: multiplying the normalized second cross-modal weight value between the i-th region and the j-th adjacent region, the image features of the j-th adjacent region of the i-th region, and the first image feature transformation matrix in the first preset sub-model to obtain the j-th second cross-modal sub-vector of the i-th region.

[0142] As can be seen, by adopting the above scheme, multiple cross-modal vectors can be generated for different types of features in any adjacent region of each region. This makes the multimodal features generated by the cross-modal vectors of each region more accurate. Therefore, by using multimodal features to train the downstream model, the trained target model can perform more accurate prediction processing.

[0143] In one implementation, generating multimodal features corresponding to the N regions in the first preset sub-model based on the first cross-modal vector and the second cross-modal vector corresponding to the N regions respectively includes:

[0144] In the first preset sub-model, based on the first cross-modal vector of the i-th region and the first homomodal vector of the i-th region in the N regions, a first multimodal vector of the i-th region is obtained; and in the first preset sub-model, based on the second cross-modal vector of the i-th region and the second homomodal vector of the i-th region in the N regions, a second multimodal vector of the i-th region is obtained; based on the first multimodal vector of the i-th region and the second multimodal vector of the i-th region, a multimodal feature corresponding to the i-th region is generated.

[0145] In the first preset sub-model, obtaining the first multimodal vector of the i-th region based on the first cross-modal vector and the first homomodal vector of the i-th region in the N regions can include: aggregating the first cross-modal vector and the first homomodal vector of the i-th region in the N regions using an aggregation function in the first preset sub-model to obtain the first multimodal vector of the i-th region. Similarly, obtaining the second multimodal vector of the i-th region based on the second cross-modal vector and the second homomodal vector of the i-th region in the first preset sub-model can include: aggregating the second cross-modal vector and the second homomodal vector of the i-th region in the N regions using an aggregation function in the first preset sub-model to obtain the second multimodal vector of the i-th region.

[0146] The aggregation function can be set according to the actual situation. The aggregation function can be any one of the following methods: concatenation, summation, or attention mechanism. We will not exhaustively list them here.

[0147] Generating multimodal features corresponding to the i-th region based on the first multimodal vector and the second multimodal vector of the i-th region can refer to concatenating the first multimodal vector and the second multimodal vector of the i-th region to obtain the multimodal features corresponding to the i-th region.

[0148] As can be seen, by adopting the above scheme, in the final processing of multimodal features, not only are the features of different modes in adjacent regions fused, but also the features of the same modes in adjacent regions are fused, thereby making the representation of each region more accurate and providing more accurate reference features for downstream tasks.

[0149] In one embodiment, the method further includes:

[0150] In the first preset sub-model, based on the image features of the neighboring regions of the i-th region and the i-th image features of the i-th region, a first isomodal weight of the i-th region is generated; based on the i-th image features of the i-th region and the first isomodal weight of the i-th region, a first isomodal vector of the i-th region is generated.

[0151] And in the first preset sub-model, based on the POI features of the adjacent regions of the i-th region and the i-th POI features of the i-th region, a second isomodal weight of the i-th region is generated; based on the i-th POI features of the i-th region and the second isomodal weight of the i-th region, a second isomodal vector of the i-th region is generated.

[0152] Specifically, generating the first isomodal weight of the i-th region in the first preset sub-model based on the image features of the neighboring regions of the i-th region and the i-th image feature of the i-th region may include: obtaining the first isomodal weight value between the i-th region and the j-th neighboring region based on the j-th image feature of the j-th neighboring region of the i-th region and the i-th image feature of the i-th region in the first preset sub-model; where j is a positive integer.

[0153] As explained above, the i-th region may include one or more neighboring regions, and the j-th neighboring region may be any one of the one or more neighboring regions of the i-th region. It should be understood that the above processing can be applied to each neighboring region of the i-th region to obtain the first cross-modal weight value between the i-th region and each neighboring region, but it will not be elaborated on in detail.

[0154] Specifically, in the first preset sub-model, based on the j-th image feature of the j-th neighboring region of the i-th region and the i-th image feature of the i-th region, the first isomodal weight value between the i-th region and the j-th neighboring region is obtained. This can include: obtaining a transformed j-th image feature based on a third image transformation matrix and the j-th image feature of the j-th neighboring region in the first preset sub-model; obtaining a transformed i-th image feature based on the third image transformation matrix and the i-th image feature of the i-th region; concatenating the transformed j-th image feature and the transformed i-th image feature to obtain a third vector; and multiplying the preset third parameter vector and the fourth vector, and then using a target activation function to calculate the first isomodal weight value between the i-th region and the j-th neighboring region.

[0155] The parameters in the third image transformation matrix can be adjusted via backpropagation of the loss function during subsequent model training. Similarly, the preset third parameter vector can also be adjusted via backpropagation of the loss function during subsequent model training.

[0156] The step of generating the first isomodal vector of the i-th region based on the i-th image feature of the i-th region and the first isomodal weight of the i-th region includes: determining the normalized first isomodal weight value between the i-th region and each of the adjacent regions based on the first isomodal weight value between the i-th region and each of the adjacent regions in the first preset sub-model; and determining the first isomodal vector of the i-th region based on the normalized first isomodal weight value between the i-th region and each of the adjacent regions and the i-th image feature of the i-th region in the first preset sub-model.

[0157] The step of determining the normalized first modal weight value between the i-th region and each of the adjacent regions based on the first modal weight value between the i-th region and each of the adjacent regions in the first preset sub-model may include: dividing the first modal weight value between the i-th region and the j-th adjacent region by the sum of the first modal weight values ​​of each of the adjacent regions in the first preset sub-model to obtain the normalized first modal weight value between the i-th region and the j-th adjacent region.

[0158] The step of determining the first isomodal vector of the i-th region in the first preset sub-model based on the normalized first isomodal weight values ​​between the i-th region and other regions, and the i-th image feature of the i-th region, may include: calculating each first isomodal sub-vector of the i-th region based on the normalized first isomodal weight values ​​between the i-th region and other regions, and the i-th image feature of the i-th region in the first preset sub-model; adding the first isomodal sub-vectors of the i-th region together, and then calculating the first isomodal vector of the i-th region based on the target activation function.

[0159] The step of determining the first isomodal vector of the i-th region in the first preset sub-model based on the normalized first isomodal weight value between the i-th region and each region, and the i-th image feature of the i-th region, may include: multiplying the normalized first cross-modal weight value between the i-th region and the j-th adjacent region, the i-th image feature of the i-th region, and the third image feature transformation matrix in the first preset sub-model to obtain the j-th first isomodal sub-vector of the i-th region.

[0160] In the first preset sub-model, based on the POI features of the neighboring regions of the i-th region and the i-th POI feature of the i-th region, a second isomodal weight of the i-th region is generated. This may include: in the first preset sub-model, based on the j-th POI feature of the j-th neighboring region of the i-th region and the i-th POI feature of the i-th region, obtaining a second isomodal weight value between the i-th region and the j-th neighboring region; where j is a positive integer.

[0161] It should be understood that the above processing can be applied to each adjacent region of the i-th region to obtain the first cross-modal weight value between the i-th region and each adjacent region, but it will not be elaborated on one by one.

[0162] Specifically, in the first preset sub-model, based on the j-th POI feature of the j-th adjacent region of the i-th region and the i-th POI feature of the i-th region, the second isomodal weight value between the i-th region and the j-th adjacent region is obtained. This can include: obtaining a transformed j-th POI feature based on a third POI transformation matrix and the j-th POI feature of the j-th adjacent region in the first preset sub-model; obtaining a transformed i-th POI feature based on the third POI transformation matrix and the i-th POI feature of the i-th region; concatenating the transformed j-th POI feature and the transformed i-th POI feature to obtain a fourth vector; and multiplying the fourth vector by a preset fourth parameter vector and the fourth vector, then using a target activation function to calculate the second isomodal weight value between the i-th region and the j-th adjacent region.

[0163] The parameters in the third POI transformation matrix can be adjusted via backpropagation of the loss function during subsequent model training. Similarly, the preset fourth parameter vector can also be adjusted via backpropagation of the loss function during subsequent model training.

[0164] The step of generating the second isomodal vector of the i-th region based on the i-th POI feature of the i-th region and the second isomodal weight of the i-th region includes: determining the normalized second isomodal weight value between the i-th region and each of the adjacent regions based on the second isomodal weight value between the i-th region and each of the adjacent regions in the first preset sub-model; and determining the second isomodal vector of the i-th region based on the normalized second isomodal weight value between the i-th region and each of the adjacent regions and the i-th POI feature of the i-th region in the first preset sub-model.

[0165] The step of determining the normalized second modal weight value between the i-th region and each of the adjacent regions based on the second modal weight value between the i-th region and each of the adjacent regions in the first preset sub-model may include: dividing the normalized second modal weight value between the i-th region and the j-th adjacent region by the sum of the second modal weight values ​​of each of the adjacent regions in the first preset sub-model.

[0166] The step of determining the second isomodal vector of the i-th region in the first preset sub-model based on the normalized second isomodal weight values ​​between the i-th region and each region, and the i-th POI feature of the i-th region, may include: calculating each second isomodal sub-vector of the i-th region based on the normalized first isomodal weight values ​​between the i-th region and each region, and the i-th POI feature of the i-th region in the first preset sub-model; adding the each second isomodal sub-vector of the i-th region together, and then calculating the second isomodal vector of the i-th region based on the target activation function.

[0167] The step of determining the second isomodal vector of the i-th region in the first preset sub-model based on the normalized second isomodal weight value between the i-th region and each region, and the i-th POI feature of the i-th region, may include: multiplying the normalized second cross-modal weight value between the i-th region and the j-th adjacent region, the i-th POI feature of the i-th region, and the third POI feature transformation matrix in the first preset sub-model to obtain the j-th second isomodal sub-vector of the i-th region.

[0168] Combination Figure 3 The processing in the aforementioned first preset sub-model is illustrated by example:

[0169] In the first preset sub-model, the image features of the neighboring regions of the i-th region are used ( Figure 3 The Chinese character is represented as The i-th image feature of the i-th region ( Figure 3 The Chinese character is represented as ), generate the first isomodal weights of the i-th region ( Figure 3 The Chinese character is represented as Based on the i-th image feature of the i-th region, the first same-modal weight of the i-th region, and the preset third parameter vector ( Figure 3 The middle is represented as a P←P ), generating the first isomodal vector of the i-th region ( Figure 3 The Chinese character is represented as In the first preset sub-model, based on the POI features of the neighboring regions of the i-th region ( Figure 3 The Chinese character is represented as The i-th POI feature of the i-th region ( Figure 3 The Chinese character is represented as ), preset fourth parameter vector ( Figure 3 The middle is represented as a I←I ), generate the second isomodal weights for the i-th region ( Figure 3 The Chinese character is represented as Based on the i-th POI feature of the i-th region, the second isomodal weight of the i-th region, and the preset fourth parameter vector, the second isomodal vector of the i-th region is generated. Figure 3 The Chinese character is represented as ).

[0170] While performing the above processing, the following processing can also be performed: Based on the POI features of the neighboring regions of the i-th region in the first preset sub-model ( Figure 3 The Chinese character is represented as The i-th image feature of the i-th region ( Figure 3 The Chinese character is represented as ), preset first parameter vector ( Figure 3 The middle is represented as a I←P ), to obtain the first cross-modal weight value of the i-th region ( Figure 3 The Chinese character is represented as Based on the first cross-modal weight value of the i-th region and the POI features of the adjacent regions of the i-th region, the first cross-modal vector of the i-th region is obtained. Figure 3 The Chinese character is represented as ); and image features of neighboring regions of the i-th region in the first preset sub-model ( Figure 3 The Chinese character is represented as The i-th POI feature of the i-th region ( Figure 3 The Chinese character is represented as ), preset second parameter vector ( Figure 3 The middle is represented as a P←I ), to obtain the second cross-modal weight value of the i-th region ( Figure 3 The Chinese character is represented as Based on the second cross-modal weight value of the i-th region and the image features of the neighboring regions of the i-th region, the second cross-modal vector of the i-th region is obtained. Figure 3 The Chinese character is represented as ).

[0171] After completing the aforementioned processing, the following processing is performed: In the first preset sub-model, based on the first cross-modal vector of the i-th region and the first homomodal vector of the i-th region in the N regions, the first multimodal vector of the i-th region is obtained; and based on the second cross-modal vector of the i-th region and the second homomodal vector of the i-th region in the N regions, the second multimodal vector of the i-th region is obtained; finally, based on the first multimodal vector of the i-th region and the second multimodal vector of the i-th region, the multimodal feature corresponding to the i-th region is generated.

[0172] As can be seen, by adopting the above scheme, features of the same modality in adjacent regions can be fused. Thus, when generating multimodal features based on the features of the same modality in adjacent regions and the cross-modal features in adjacent regions, the multimodal features can more accurately represent each region, providing more accurate input data for downstream tasks.

[0173] A second aspect of this disclosure provides an information processing apparatus, such as... Figure 4 As shown, it includes:

[0174] The region feature generation module 401 is used to generate POI features and image features corresponding to the N regions based on the POI data and image data corresponding to the N regions respectively; N is a positive integer;

[0175] The first processing module 402 is used to generate multimodal features corresponding to the N regions based on the adjacency relationship of the N regions, the POI features and the image features corresponding to the N regions respectively.

[0176] The region feature generation module 401 is used to obtain the distribution feature of the i-th POI, the distance feature of the i-th POI, and the distribution feature of the i-th facility in the i-th region of the N regions based on the POI data corresponding to the i-th regions respectively; generate the i-th POI feature of the i-th region based on the distribution feature of the i-th POI, the distance feature of the i-th POI, and the distribution feature of the i-th facility; and extract features from the i-th image data of the i-th region to obtain the i-th image feature of the i-th region.

[0177] exist Figure 4 On the basis of, such as Figure 5 As shown, the device further includes:

[0178] The adjacency relationship generation module 501 is used to determine the spatial adjacent regions corresponding to the N regions based on the spatial characteristics of the N regions, and to determine the connected adjacent regions corresponding to the N regions based on preset road network connectivity data; to use the spatial adjacent regions and the connected adjacent regions corresponding to the N regions as the adjacent regions corresponding to the N regions; and to generate the adjacency relationship of the N regions based on the adjacent regions corresponding to the N regions.

[0179] The first processing module 402 is used to input the adjacency relationship of the N regions, as well as the POI features and image features corresponding to the N regions respectively, into a first preset sub-model to obtain the multimodal features corresponding to the N regions respectively output by the first preset sub-model.

[0180] The device further includes: a second processing module 502, used to input the multimodal features corresponding to the N regions into a second preset sub-model to obtain the prediction information output by the second preset sub-model; determine a loss function based on the prediction information; update the first preset sub-model and the second preset sub-model based on the loss function through backpropagation; and obtain the trained first target sub-model and the second target sub-model when it is determined that the first preset sub-model and the second preset sub-model have completed training.

[0181] The first processing module 402 is used to input the adjacency relationship of the N regions, and the POI features and image features corresponding to the N regions respectively into the first preset sub-model. In the first preset sub-model, based on the adjacency relationship of the N regions, and the POI features and image features corresponding to the N regions respectively, a first cross-modal vector and a second cross-modal vector corresponding to the N regions are generated respectively. In the first preset sub-model, the first cross-modal vector and the second cross-modal vector corresponding to the N regions are processed to obtain the multimodal features corresponding to the N regions respectively output by the first preset sub-model.

[0182] The first processing module 402 is configured to: obtain the neighboring regions of the i-th region from the adjacency relationships of the N regions in the first preset sub-model; generate the first cross-modal vector of the i-th region based on the POI features of the neighboring regions of the i-th region and the i-th image features of the i-th region in the first preset sub-model; and generate the second cross-modal vector corresponding to the i-th region based on the image features of the neighboring regions of the i-th region and the i-th POI features of the i-th region in the first preset sub-model; where i is a positive integer less than or equal to N.

[0183] The first processing module 402 is configured to: obtain a first cross-modal weight value for the i-th region based on the POI features of the adjacent regions of the i-th region and the i-th image features of the i-th region in the first preset sub-model; obtain a first cross-modal vector for the i-th region based on the first cross-modal weight value of the i-th region and the POI features of the adjacent regions of the i-th region; obtain a second cross-modal weight value for the i-th region based on the image features of the adjacent regions of the i-th region and the i-th POI features of the i-th region in the first preset sub-model; and obtain a second cross-modal vector for the i-th region based on the second cross-modal weight value of the i-th region and the image features of the adjacent regions of the i-th region.

[0184] The first processing module 402 is configured to obtain a first multimodal vector of the i-th region based on the first cross-modal vector and the first homomodal vector of the i-th region in the N regions in the first preset sub-model; and obtain a second multimodal vector of the i-th region based on the second cross-modal vector and the second homomodal vector of the i-th region in the N regions; and generate multimodal features corresponding to the i-th region based on the first multimodal vector and the second multimodal vector of the i-th region.

[0185] The first processing module 402 is configured to generate a first isomodal weight of the i-th region based on the image features of the neighboring regions of the i-th region and the i-th image features of the i-th region in the first preset sub-model; and generate a first isomodal vector of the i-th region based on the i-th image features of the i-th region and the first isomodal weight of the i-th region.

[0186] And in the first preset sub-model, based on the POI features of the adjacent regions of the i-th region and the i-th POI features of the i-th region, a second isomodal weight of the i-th region is generated; based on the i-th POI features of the i-th region and the second isomodal weight of the i-th region, a second isomodal vector of the i-th region is generated.

[0187] In this embodiment, the information processing device can be specifically installed in an electronic device, such as a server. Alternatively, different modules of the aforementioned information processing device can be installed in different electronic devices. Or, at least some modules of the aforementioned information processing device can be installed in the same electronic device, while the remaining modules can be installed in another electronic device. This embodiment does not exhaustively list all possibilities.

[0188] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0189] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0190] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0191] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0192] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0193] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above. For example, information processing methods, in some embodiments, the information processing methods described above can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the information processing methods described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the information processing methods described above by any other suitable means (e.g., by means of firmware).

[0194] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0195] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable information processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0196] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0197] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0198] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0199] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0200] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0201] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An information processing method, comprising: generating POI features and image features corresponding to N regions respectively based on POI data and image data corresponding to the N regions respectively, including: splicing POI category distribution features of an i-th region in the N regions and reference POI category distribution features of K reference regions containing the i-th region to obtain i-th POI distribution features of the i-th region, generating i-th POI distance features of the i-th region based on the nearest distances between the i-th region and POIs of each of a plurality of preset POI categories, determining that i-th facility distribution features of the i-th region take a first value in a case where a specified range containing the i-th region contains all specified POI categories, and generating i-th POI features of the i-th region based on the i-th POI distribution features, the i-th POI distance features, and the i-th facility distribution features; N is a positive integer, and i is a positive integer less than or equal to N; generating multi-modal features corresponding to the N regions respectively based on the adjacent relationships of the N regions and the POI features and the image features corresponding to the N regions respectively.

2. The method of claim 1, wherein, The generating of the POI features and the image features corresponding to the N regions respectively based on the POI data and the image data corresponding to the N regions respectively further includes: performing feature extraction on i-th image data of the i-th region to obtain i-th image features of the i-th region.

3. The method of claim 1, further comprising: determining spatial adjacent regions corresponding to the N regions respectively based on spatial features of the N regions, and determining connected adjacent regions corresponding to the N regions respectively based on preset road network connectivity data; taking the spatial adjacent regions corresponding to the N regions respectively and the connected adjacent regions corresponding to the N regions respectively as adjacent regions corresponding to the N regions respectively; generating the adjacent relationships of the N regions based on the adjacent regions corresponding to the N regions respectively.

4. The method according to any one of claims 1 to 3, wherein, The generating of the multi-modal features corresponding to the N regions respectively based on the adjacent relationships of the N regions and the POI features and the image features corresponding to the N regions respectively includes: inputting the adjacent relationships of the N regions and the POI features and the image features corresponding to the N regions respectively into a first preset sub-model to obtain multi-modal features corresponding to the N regions respectively output by the first preset sub-model.

5. The method of claim 4, further comprising: inputting the multi-modal features corresponding to the N regions respectively into a second preset sub-model to obtain prediction information output by the second preset sub-model; determining a loss function based on the prediction information, updating the first preset sub-model and the second preset sub-model based on the loss function; in a case where the first preset sub-model and the second preset sub-model complete training, obtaining a first target sub-model and a second target sub-model after training.

6. The method of claim 4, wherein, The adjacent relationship of the N regions, and the POI features and the image features corresponding to the N regions are input into a first preset sub-model, and a plurality of modal features corresponding to the N regions output by the first preset sub-model are obtained. The adjacent relationship of the N regions, and the POI features and the image features corresponding to the N regions are input into the first preset sub-model, and the first cross-modal vector and the second cross-modal vector corresponding to the N regions are generated in the first preset sub-model based on the adjacent relationship of the N regions, and the POI features and the image features corresponding to the N regions. The first cross-modal vector and the second cross-modal vector corresponding to the N regions are processed in the first preset sub-model, and the plurality of modal features corresponding to the N regions output by the first preset sub-model are obtained.

7. The method of claim 6, wherein, The first cross-modal vector and the second cross-modal vector corresponding to the N regions are generated in the first preset sub-model based on the adjacent relationship of the N regions, and the POI features and the image features corresponding to the N regions, including: In the first preset sub-model, the adjacent regions of an i-th region in the N regions are obtained from the adjacent relationship of the N regions, and the first cross-modal vector of the i-th region is generated in the first preset sub-model based on the POI features of the adjacent regions of the i-th region and the i-th image feature of the i-th region, and the second cross-modal vector corresponding to the i-th region is generated in the first preset sub-model based on the image features of the adjacent regions of the i-th region and the i-th POI feature of the i-th region.

8. The method of claim 7, wherein, The first cross-modal vector of the i-th region is generated in the first preset sub-model based on the POI features of the adjacent regions of the i-th region and the i-th image feature of the i-th region, including: first cross-modal weight values of the i-th region are obtained in the first preset sub-model based on the POI features of the adjacent regions of the i-th region and the i-th image feature of the i-th region; and the first cross-modal vector of the i-th region is obtained based on the first cross-modal weight values of the i-th region and the POI features of the adjacent regions of the i-th region. The second cross-modal vector corresponding to the i-th region is generated in the first preset sub-model based on the image features of the adjacent regions of the i-th region and the i-th POI feature of the i-th region, including: second cross-modal weight values of the i-th region are obtained in the first preset sub-model based on the image features of the adjacent regions of the i-th region and the i-th POI feature of the i-th region; and the second cross-modal vector of the i-th region is obtained based on the second cross-modal weight values of the i-th region and the image features of the adjacent regions of the i-th region.

9. The method of claim 8, wherein, The generating, in the first preset sub-model, of the multi-modal feature corresponding to each of the N regions based on the first cross-modal vector and the second cross-modal vector corresponding to each of the N regions comprises: In the first preset sub-model, a first multi-modal vector of the ith region is obtained based on the first cross-modal vector of the ith region of the N regions and a first intra-modal vector of the ith region, and a second multi-modal vector of the ith region is obtained based on the second cross-modal vector of the ith region of the N regions and a second intra-modal vector of the ith region. The first multi-modal vector of the ith region and the second multi-modal vector of the ith region are used to generate a multi-modal feature corresponding to the ith region.

10. The method of claim 9, further comprising: In the first preset sub-model, a first intra-modal weight of the ith region is generated based on image features of adjacent regions of the ith region and ith image features of the ith region. The first intra-modal weight of the ith region and the ith image features of the ith region are used to generate a first intra-modal vector of the ith region. In the first preset sub-model, a second intra-modal weight of the ith region is generated based on POI features of adjacent regions of the ith region and ith POI features of the ith region. The second intra-modal weight of the ith region and the ith POI features of the ith region are used to generate a second intra-modal vector of the ith region.

11. An information processing apparatus, comprising: A region feature generation module configured to generate POI features and image features corresponding to N regions based on POI data and image data corresponding to the N regions, including: splicing POI category distribution features of an ith region of the N regions and reference POI category distribution features of K reference regions containing the ith region to obtain ith POI distribution features of the ith region, generating ith POI distance features of the ith region based on the nearest distances between the ith region and POIs of each of a plurality of preset POI categories, determining that an ith facility distribution feature of the ith region takes a first value in a case where a specified range containing the ith region contains all specified POI categories, and generating ith POI features of the ith region based on the ith POI distribution features, the ith POI distance features, and the ith facility distribution features; N is a positive integer, and i is a positive integer less than or equal to N. A first processing module configured to generate multi-modal features corresponding to the N regions based on adjacent relationships of the N regions and the POI features and the image features corresponding to the N regions.

12. The apparatus of claim 11, wherein, The region feature generation module is configured to perform feature extraction on ith image data of the ith region to obtain ith image features of the ith region.

13. The apparatus of claim 11, further comprising: an adjacent relationship generation module configured to determine, based on spatial features of the N regions, spatial adjacent regions corresponding to the N regions respectively, and determine, based on preset road network connectivity data, connectivity adjacent regions corresponding to the N regions respectively; use the spatial adjacent regions corresponding to the N regions respectively and the connectivity adjacent regions corresponding to the N regions respectively as adjacent regions corresponding to the N regions respectively; and generate adjacent relationships of the N regions based on the adjacent regions corresponding to the N regions respectively.

14. The apparatus of any one of claims 11-13, wherein, The first processing module is configured to input the adjacent relationships of the N regions, and the POI features and the image features corresponding to the N regions respectively into a first preset sub-model to obtain multi-modal features corresponding to the N regions respectively output by the first preset sub-model.

15. The apparatus of claim 14, further comprising: a second processing module configured to input the multi-modal features corresponding to the N regions respectively into a second preset sub-model to obtain prediction information output by the second preset sub-model; determine a loss function based on the prediction information, update the first preset sub-model and the second preset sub-model based on the loss function, and obtain a first target sub-model and a second target sub-model trained under the condition that the first preset sub-model and the second preset sub-model complete training.

16. The apparatus of claim 14, wherein, The first processing module is configured to input the adjacent relationships of the N regions, and the POI features and the image features corresponding to the N regions respectively into the first preset sub-model, generate first cross-modal vectors and second cross-modal vectors corresponding to the N regions respectively based on the adjacent relationships of the N regions, and the POI features and the image features corresponding to the N regions respectively in the first preset sub-model, and process the first cross-modal vectors and the second cross-modal vectors corresponding to the N regions respectively in the first preset sub-model to obtain the multi-modal features corresponding to the N regions respectively output by the first preset sub-model.

17. The apparatus of claim 16, wherein, The first processing module is configured to obtain adjacent regions of an i-th region of the N regions from the adjacent relationships of the N regions in the first preset sub-model. generate the first cross-modal vector of the i-th region based on the POI features of the adjacent regions of the i-th region and the i-th image features of the i-th region in the first preset sub-model, and generate the second cross-modal vector corresponding to the i-th region based on the image features of the adjacent regions of the i-th region and the i-th POI features of the i-th region in the first preset sub-model.

18. The apparatus of claim 17, wherein, The first processing module is configured to obtain, in the first preset sub-model, a first cross-modal weight value of the ith region based on the POI features of the adjacent regions of the ith region and the ith image features of the ith region; obtain the first cross-modal vector of the ith region based on the first cross-modal weight value of the ith region and the POI features of the adjacent regions of the ith region; and obtain, in the first preset sub-model, a second cross-modal weight value of the ith region based on the image features of the adjacent regions of the ith region and the ith POI features of the ith region; and obtain the second cross-modal vector of the ith region based on the second cross-modal weight value of the ith region and the image features of the adjacent regions of the ith region.

19. The apparatus of claim 18, wherein, The first processing module is configured to obtain, in the first preset sub-model, a first multi-modal vector of the ith region based on the first cross-modal vector of the ith region and a first same-modal vector of the ith region among the N regions; and obtain a second multi-modal vector of the ith region based on the second cross-modal vector of the ith region and a second same-modal vector of the ith region among the N regions; and generate the multi-modal feature corresponding to the ith region based on the first multi-modal vector of the ith region and the second multi-modal vector of the ith region.

20. The apparatus of claim 19, wherein, The first processing module is configured to generate, in the first preset sub-model, a first same-modal weight of the ith region based on the image features of the adjacent regions of the ith region and the ith image features of the ith region; generate a first same-modal vector of the ith region based on the ith image features of the ith region and the first same-modal weight of the ith region; generate a second same-modal weight of the ith region based on the POI features of the adjacent regions of the ith region and the ith POI features of the ith region in the first preset sub-model; generate a second same-modal vector of the ith region based on the ith POI features of the ith region and the second same-modal weight of the ith region. 21.An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10.

22. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-10. 23.A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-10.

Citation Information

Patent Citations

  • Data fusion method and data fusion equipment

    CN109993184A

  • Multi-modal POI feature extraction method and device

    CN113032672A