A method, apparatus, device, and storage medium for processing multimedia location information.

By acquiring the feature vectors of the target multimedia and historical multimedia, and utilizing feature extraction and location prediction models, the problem of inaccurate location information in similar street scenes or indoor scenes is solved, achieving more efficient and accurate location information determination.

CN114186085BActive Publication Date: 2026-03-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2026-03-10

Smart Images

  • Figure CN114186085B_ABST
    Figure CN114186085B_ABST
Patent Text Reader

Abstract

This disclosure relates to a method, apparatus, device, and storage medium for processing multimedia location information. It involves acquiring target multimedia information and publishing object information of a target multimedia, determining the historical multimedia associated with the publishing object information, as well as the historical multimedia information and historical location information of the historical multimedia, and using a feature extraction model to perform feature extraction processing on the target multimedia information, historical multimedia information, and the location information of the historical multimedia to obtain feature vectors for the target multimedia and historical multimedia, thereby improving the efficiency and quality of feature vector acquisition. Furthermore, the target location information of the target multimedia is predicted based on the feature vectors of the historical multimedia associated with the publishing object information and the feature vector of the target multimedia. By mining the relationship between the target multimedia and the historical multimedia associated with the publishing object, the target location information of the target multimedia can be determined, thereby improving the accuracy of the target location information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, device and storage medium for processing multimedia location information. Background Technology

[0002] With the continuous development of technology, short video platforms are increasingly appearing in people's lives. Users can create and publish short videos on these platforms, and platforms can also recommend short videos to users or deliver videos to users based on their searches. For short video recommendations and searches, determining the location information of the short video is particularly important. One related technology involves registering the image content information of the target video with the image content information of a known location to find the closest matching location, which is then used as the target video's location. However, for similar street scenes or indoor scenes, using image registration to determine the target video's location can be inaccurate due to the low accuracy of image feature recognition. Summary of the Invention

[0003] This disclosure provides a method, apparatus, device, and storage medium for processing multimedia location information, to at least solve the problem in related technologies where accurate location information cannot be guaranteed for similar street scenes or indoor scenes. The technical solution of this disclosure is as follows:

[0004] According to a first aspect of the present disclosure, a method for processing multimedia location information is provided, comprising:

[0005] Acquire target multimedia information and publishing object information for the target multimedia;

[0006] The historical multimedia associated with the published object information, as well as the historical multimedia information and historical location information of the historical multimedia, are determined; the historical location information is the geographical location information of the media content in the historical multimedia.

[0007] The target multimedia information, the historical multimedia information, and the location information of the historical multimedia are input into the feature extraction model for feature extraction processing to obtain the feature vectors of the target multimedia and the historical multimedia.

[0008] The feature vectors of the target multimedia and the historical multimedia are input into the location prediction model to obtain the target location information of the target multimedia, which is the geographical location information of the media content in the target multimedia.

[0009] In one possible implementation, the method further includes:

[0010] Obtain first sample display location information, first sample feature vector and node information corresponding to multiple first sample multimedia, wherein the first sample display location information is the display location information of the multiple first sample multimedia, and the node information is the adjacency information of the first node in the first sample display location information and the second node in the first sample multimedia;

[0011] The first sample display location information and the node information are input into a preset feature extraction model to perform feature extraction and obtain a predicted feature vector.

[0012] Based on the first sample feature vector and the predicted feature vector, the preset feature extraction model is trained to obtain the feature extraction model.

[0013] In one possible implementation, the step of inputting the feature vector of the target multimedia and the feature vector of the historical multimedia into a location prediction model to obtain the target location information of the target multimedia includes:

[0014] The feature vectors of the target multimedia and the historical multimedia are subjected to feature merging processing to obtain the semantic features corresponding to the target multimedia and the historical multimedia respectively;

[0015] The semantic features of the target multimedia and the historical multimedia are fitted to obtain the target location information of the target multimedia.

[0016] In one possible implementation, the target location information of the target multimedia includes identification information of the target location;

[0017] After the step of inputting the feature vector of the target multimedia and the feature vector of the historical multimedia into the location prediction model to obtain the target location information of the target multimedia, the method further includes:

[0018] Obtain a location mapping table, which represents the correspondence between identification information and location text information;

[0019] The target location text information corresponding to the identifier information of the target location is determined from the location mapping table.

[0020] In one possible implementation, after the step of inputting the target multimedia information, the historical multimedia information, and the location information of the historical multimedia into a feature extraction model for feature extraction processing to obtain the feature vectors of the target multimedia and the historical multimedia, the method further includes:

[0021] After the target multimedia is displayed, the target display location information of the target multimedia is obtained, and the target display location information is the display location information of the target multimedia;

[0022] Based on the feature vector of the target multimedia information and the target display location information, the first sample feature vector and the first sample display location information are updated respectively to obtain the updated first sample feature vector and the first sample display location information.

[0023] In one possible implementation, obtaining the first sample display location information corresponding to multiple first sample multimedias includes:

[0024] Determine multiple display location information corresponding to the multiple first sample multimedia;

[0025] Based on the multiple display position information corresponding to the multiple first sample multimedia, determine the reference position information corresponding to each first sample multimedia.

[0026] Display position information with a distance less than a preset distance threshold from the reference position information is selected to obtain the first sample display position information corresponding to each first sample multimedia.

[0027] In one possible implementation, the method further includes:

[0028] Obtain second sample data and corresponding data tags. The second sample data includes sample multimedia information of multiple second sample multimedia, associated sample multimedia information related to the publishing object information of the second sample multimedia, and associated sample location information of each associated sample multimedia. The data tags are the geographical location information of the media content in the second sample multimedia. The associated sample location information of each associated sample multimedia is the geographical location information of the media content of each associated sample multimedia.

[0029] The sample multimedia information of the second sample multimedia, the associated sample multimedia information, and the associated sample location information are input into the feature extraction model for feature extraction processing to obtain the feature vectors of the multiple second sample multimedias and the feature vectors of the associated multimedias.

[0030] The feature vectors of the multiple second sample multimedias and the feature vectors of the associated sample multimedias are input into a preset machine learning model for location prediction processing to obtain the location prediction information of the second sample multimedias.

[0031] Based on the location prediction information and the corresponding data labels, the preset machine learning model is trained to obtain the location prediction model.

[0032] In one possible implementation, after the step of inputting the feature vector of the target multimedia and the feature vector of the historical multimedia into a location prediction model to obtain the target location information of the target multimedia, the method further includes:

[0033] The sample multimedia information of the second sample multimedia is updated based on the target multimedia information to obtain the updated sample multimedia information of the second sample multimedia.

[0034] The data tag is updated based on the target location information to obtain the updated data tag.

[0035] According to a second aspect of the present disclosure, a multimedia location information processing apparatus is provided, comprising:

[0036] The target information acquisition module is configured to acquire the target multimedia information and the publishing object information of the target multimedia.

[0037] The module for determining the information to be extracted is configured to determine the historical multimedia associated with the published object information, as well as the historical multimedia information and the historical location information of the historical multimedia; the historical location information is the geographical location information of the media content in the historical multimedia.

[0038] The target feature extraction module is configured to input the target multimedia information, the historical multimedia information, and the location information of the historical multimedia into the feature extraction model, perform feature extraction processing, and obtain the feature vector of the target multimedia and the feature vector of the historical multimedia.

[0039] The location information prediction module is configured to input the feature vector of the target multimedia and the feature vector of the historical multimedia into the location prediction model to obtain the target location information of the target multimedia, wherein the target location information is the geographical location information of the media content in the target multimedia.

[0040] In one possible implementation, the device further includes:

[0041] The first sample information acquisition module is configured to acquire first sample display location information, first sample feature vector and node information corresponding to multiple first sample multimedia, wherein the first sample display location information is the display location information of the multiple first sample multimedia, and the node information is the adjacency information of the first node in the first sample display location information and the second node in the first sample multimedia.

[0042] The sample feature extraction module is configured to input the first sample display location information and the node information into a preset feature extraction model to perform feature extraction and obtain a predicted feature vector.

[0043] The first training module is configured to train the preset feature extraction model based on the first sample feature vector and the predicted feature vector to obtain the feature extraction model.

[0044] In one possible implementation, the location information prediction module includes:

[0045] The feature merging unit is configured to perform feature merging processing on the feature vector of the target multimedia and the feature vector of the historical multimedia to obtain the semantic features corresponding to the target multimedia and the historical multimedia respectively.

[0046] The fitting unit is configured to perform fitting processing on the semantic features of the target multimedia and the historical multimedia in the multilayer perceptron of the location association model to obtain the target location information of the target multimedia.

[0047] In one possible implementation, the device further includes:

[0048] The location mapping table acquisition module is configured to acquire a location mapping table, which represents the correspondence between identification information and location text information.

[0049] The location text information determination module is configured to determine the target location text information corresponding to the identifier information of the target location from the location mapping table.

[0050] In one possible implementation, the device further includes:

[0051] The display location information acquisition module is configured to acquire the target display location information of the target multimedia after the target multimedia is displayed, wherein the target display location information is the display location information of the target multimedia;

[0052] The first update module is configured to update the first sample feature vector and the first sample display location information based on the feature vector of the target multimedia information and the target display location information, respectively, to obtain the updated first sample feature vector and the first sample display location information.

[0053] In one possible implementation, the sample information acquisition module includes:

[0054] The display location information acquisition unit is configured to determine multiple display location information corresponding to the multiple first sample multimedia;

[0055] The reference position information acquisition unit is configured to determine the reference position information corresponding to each first sample multimedia based on the multiple display position information corresponding to the multiple first sample multimedia.

[0056] The sample information acquisition unit is configured to filter out display position information whose distance from the reference position information is less than a preset distance threshold, thereby obtaining the first sample display position information corresponding to each sample multimedia.

[0057] In one possible implementation, the device further includes:

[0058] The second sample information acquisition module is configured to acquire second sample data and corresponding data tags. The second sample data includes sample multimedia information of multiple second sample multimedia, associated sample multimedia information related to the publishing object information of the second sample multimedia, and associated sample location information of each associated sample multimedia. The data tags are the geographical location information of the media content in the second sample multimedia. The associated sample location information of each associated sample multimedia is the geographical location information of the media content of each associated sample multimedia.

[0059] The sample feature determination module is configured to input the sample multimedia information of the second sample multimedia, the associated sample multimedia information, and the associated sample location information into the feature extraction model, perform feature extraction processing, and obtain the feature vectors of the multiple second sample multimedias and the feature vectors of the associated multimedias.

[0060] The prediction information determination module is configured to input the feature vectors of the plurality of second sample multimedia and the feature vectors of the associated sample multimedia into a preset machine learning model for position prediction processing to obtain the position prediction information of the second sample multimedia.

[0061] The second training module is configured to train the preset machine learning model based on the location prediction information and the corresponding data labels to obtain the location prediction model.

[0062] In one possible implementation, the device further includes:

[0063] The second update module is configured to update the sample multimedia information of the second sample multimedia based on the target multimedia information, so as to obtain the updated sample multimedia information of the second sample multimedia.

[0064] The third update module is configured to update the data tag based on the target location information to obtain the updated data tag.

[0065] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method as described in any one of the first aspects above.

[0066] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided such that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the methods described in the first aspect of the present disclosure.

[0067] According to a fifth aspect of the present disclosure, a computer program product is provided, including computer instructions that, when executed by a processor, cause a computer to perform the method described in any one of the first aspects of the present disclosure. The technical solutions provided by the embodiments of the present disclosure offer at least the following beneficial effects:

[0068] By acquiring target multimedia information and publishing object information, we determine the historical multimedia associated with the publishing object information, as well as the historical multimedia information and historical location information of the historical multimedia. Using a feature extraction model, we perform feature extraction processing on the target multimedia information, historical multimedia information, and the location information of the historical multimedia to obtain feature vectors for the target multimedia and historical multimedia, thus improving the efficiency and quality of feature vector acquisition. Furthermore, the target location information of the target multimedia is predicted based on the feature vectors of the historical multimedia associated with the publishing object information and the feature vector of the target multimedia. By mining the relationship between the target multimedia and the historical multimedia associated with the publishing object, we can determine the target location information of the target multimedia, thereby improving the accuracy of the target location information. Attached Figure Description

[0069] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0070] Figure 1 This is a schematic diagram illustrating an implementation environment according to an exemplary embodiment.

[0071] Figure 2 This is a flowchart illustrating a multimedia location information processing method according to an exemplary embodiment.

[0072] Figure 3 This is a flowchart illustrating a multimedia location information processing method according to an exemplary embodiment.

[0073] Figure 4 This is a flowchart illustrating a method for predicting target location information according to an exemplary embodiment.

[0074] Figure 5 This is a flowchart illustrating a multimedia location information processing method according to an exemplary embodiment.

[0075] Figure 6 This is a flowchart illustrating a multimedia location information processing method according to an exemplary embodiment.

[0076] Figure 7 This is a flowchart illustrating an exemplary embodiment of obtaining the display location information of a first sample corresponding to multiple first sample multimedia.

[0077] Figure 8 This is a schematic diagram illustrating the display of location information of a first sample according to an exemplary embodiment.

[0078] Figure 9 This is a flowchart illustrating a multimedia location information processing method according to an exemplary embodiment.

[0079] Figure 10 This is a flowchart illustrating a multimedia location information processing method according to an exemplary embodiment.

[0080] Figure 11 This is a block diagram of a multimedia location information processing apparatus according to an exemplary embodiment.

[0081] Figure 12 This is a block diagram illustrating an electronic device for a multimedia location information processing method according to an exemplary embodiment. Detailed Implementation

[0082] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0083] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0084] Please see Figure 1 It illustrates a schematic diagram of an implementation environment provided by an embodiment of this disclosure, which may include:

[0085] At least one terminal 01, at least one server 02, and at least one terminal 03. The at least one terminal 01 and at least one terminal 03 can communicate with the at least one server 02 via a network.

[0086] In an optional embodiment, terminal 01 can be a client that captures and / or creates target multimedia, providing the target multimedia to server 02. Terminal 01 can be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. The operating system running on terminal 01 can be, but is not limited to, Android, iOS, Linux, Windows, and Unix.

[0087] In an optional embodiment, server 02 can be a server that processes multimedia location information based on the target multimedia provided by terminal 01 to obtain the target location information of the target multimedia. Optionally, server 02 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0088] In an optional embodiment, terminal 03 can be a client that provides server 02 with the display location information of sample multimedia. Server 02 can train a preset feature extraction model based on the display location information of the sample multimedia. Terminal 03 can be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. The operating system running on terminal 03 can be, but is not limited to, Android, iOS, Linux, Windows, and Unix.

[0089] It should be noted that the following diagram illustrates one possible sequence of steps, and it is not strictly required to follow this order. Some steps can be performed in parallel without interdependence. The user information (including but not limited to user device information, user personal information, user behavior information, etc.) and data (including but not limited to data used for display, training data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0090] Figure 2This is a flowchart illustrating a multimedia location information processing method according to an exemplary embodiment. This multimedia location information processing method can be applied to server 02, such as... Figure 2 As shown, the method for processing multimedia location information includes the following steps:

[0091] In step S21, the target multimedia information and the publishing object information of the target multimedia are obtained.

[0092] In the embodiments of this specification, "target multimedia" refers to multimedia captured and / or produced by the publishing object. The target multimedia information may include image content information, interactive information, and associated information of the publishing object. The image content information may be image-text information obtained by converting the audio / video file of the target multimedia into text. Interactive information may include likes, comments, etc. The associated information of the publishing object may include basic information about the publishing object, such as its age, gender, and city. The publishing object information may be the identity identifier of the publishing object, which is used to identify the publishing object of the target multimedia.

[0093] In practical applications, the target multimedia information and the publishing object information can be obtained in real time from the terminal of the publishing object. Alternatively, a pre-built multimedia library can be constructed. After the publishing object uploads the target multimedia to the platform, the pre-built multimedia library stores the target multimedia. The pre-built multimedia library can store at least one multimedia corresponding to different publishing objects; the platform can be an e-commerce platform, a multimedia resource platform, etc. The server can then retrieve the target multimedia information and the publishing object information of the target multimedia from the pre-built multimedia library.

[0094] In step S22, the historical multimedia associated with the published object information, as well as the historical multimedia information and historical location information of the historical multimedia, are determined.

[0095] In the embodiments of this specification, the historical multimedia associated with the publishing object can refer to historical multimedia published by the publishing object before publishing the target multimedia. The number of historical multimedia can be at least one, and this disclosure does not limit this. The historical multimedia information refers to the image content information, interactive information, and association information of the publishing object, etc., of the historical multimedia. The historical location information of the historical multimedia refers to the geographical location information of the media content in the historical multimedia. The specific historical location information of the historical multimedia can be manually annotated or predicted when the historical multimedia is published.

[0096] In practical applications, to accurately determine the correspondence between publishing object information and historical multimedia content, and to improve the efficiency of determining this correspondence, a mapping table between publishing object information and historical multimedia content can be pre-set. This mapping table includes the correspondence between the identity identifier of the publishing object and the identifier of the historical multimedia content. Based on this mapping table, the identifier of the historical multimedia content corresponding to the publishing object information can be determined. Based on the identifier of the historical multimedia content, the historical multimedia information and historical location information of the historical multimedia content can be determined from a pre-set multimedia database.

[0097] In step S23, the target multimedia information, historical multimedia information, and the location information of historical multimedia are input into the feature extraction model for feature extraction processing to obtain the feature vectors of the target multimedia and the historical multimedia.

[0098] In the embodiments of this specification, the feature extraction model can be obtained by training a preset feature extraction model based on a training sample set. Inputting the target multimedia information into the feature extraction model yields the feature vector of the target multimedia; inputting historical multimedia information and its location information yields the feature vector of the historical multimedia. Specifically, the target multimedia information and historical multimedia information can be discrete variables. Inputting these into the feature extraction model allows the model to obtain a floating-point vector with fewer dimensions (compared to the number of words representing the semantics of the multimedia).

[0099] In step S24, the feature vectors of the target multimedia and the feature vectors of historical multimedia are input into the location prediction model to obtain the target location information of the target multimedia.

[0100] In the embodiments of this specification, the feature vectors of the target multimedia and the historical multimedia are used together as inputs to the location prediction model. The location prediction model can determine the similarity between the target multimedia and the historical multimedia based on their feature vectors, and determine the weight corresponding to the historical multimedia. Based on the weight corresponding to the historical multimedia and the similarity between the historical multimedia and the target multimedia, the historical location information of the historical multimedia is weighted to obtain the target location information of the target multimedia. Specifically, the target location information of the target multimedia can refer to the geographical location information of the media content in the target multimedia.

[0101] By acquiring the target multimedia information and the publishing object information of the target multimedia, determining the historical multimedia associated with the publishing object information, as well as the historical multimedia information and historical location information of the historical multimedia, and using a feature extraction model to directly extract features from the target multimedia information, historical multimedia information, and the location information of the historical multimedia, the feature vectors of the target multimedia and the historical multimedia can be obtained, which can improve the efficiency and quality of feature vector acquisition. Furthermore, the target location information of the target multimedia is predicted based on the feature vectors of the historical multimedia associated with the publishing object information and the feature vector of the target multimedia. By mining the relationship between the target multimedia and the historical multimedia associated with the publishing object, the target location information of the target multimedia can be determined, which can improve the accuracy of the target location information.

[0102] Figure 3 This is a flowchart illustrating a multimedia location information processing method according to an exemplary embodiment, used to pre-generate a feature extraction model, such as... Figure 3 As shown, the method may further include the following steps:

[0103] In step S31, the display location information of the first sample corresponding to the first sample multimedia, the feature vector of the first sample, and the node information are obtained.

[0104] In the embodiments of this specification, the first sample multimedia can refer to historical multimedia published by the publishing object; the first sample display location information corresponding to the first sample multimedia can refer to the display location information of the first sample multimedia. Node information can be the adjacency information of the first node in the first sample display location information and the second node in the first sample multimedia. Specifically, a bipartite graph structure of the first sample display location information and the first sample multimedia can be pre-constructed, which can represent the correspondence between the first sample display location information and the first sample multimedia. For example, when user 1's terminal displays sample multimedia 1, an undirected edge is established between the display location information of the sample multimedia 1 displayed by user 1's terminal and the sample multimedia 1; for each multimedia in the first sample multimedia, an undirected edge is established between the multimedia and each display location information corresponding to the multimedia, thereby obtaining the bipartite graph structure of the first sample display location information and the first sample multimedia. Adjacency information refers to the adjacency information of the first sample display location and the first sample multimedia in a bipartite graph structure. For example, when user 1's terminal displays sample multimedia 1, the display location information of user 1's terminal displaying sample multimedia 1 can be used as the first node, sample multimedia 1 can be used as the second node, and the first node and the second node are adjacent nodes to each other.

[0105] In step S32, the location information and node information of the first sample are input into a preset feature extraction model to extract features and obtain a predicted feature vector.

[0106] In the embodiments of this specification, the preset feature extraction model can be a VGG (Visual Geometry Group) convolutional neural network, and this disclosure does not limit it.

[0107] In step S33, a preset feature extraction model is trained based on the first sample feature vector and the predicted feature vector to obtain the feature extraction model.

[0108] In the embodiments of this specification, a loss function for a preset feature extraction model can be determined based on the first sample feature vector and the predicted feature vector. For example, the difference between the first sample feature vector and the predicted feature vector can be taken, and the absolute value of the difference can be used as the loss function, or the square of the difference can be used as the loss function. The display location information of the first sample can be used as training data, and the first sample feature vector can be used as the label corresponding to the training data. The preset feature extraction model can be trained based on the training data and the corresponding labels until the value of the loss function no longer changes or the value of the loss function is less than a threshold, thus obtaining the feature extraction model.

[0109] Optionally, during the training of the preset feature extraction model, the relationship between the first sample display location information, the first sample feature vector, and the node information can be expressed by the following formulas (1) and (2):

[0110]

[0111]

[0112] in, This indicates the display location information of the first sample after the (k+1)th iteration. This represents the feature vector of the first sample after the k-th iteration. This represents the feature vector of the first sample after the (k+1)th iteration.

[0113] N represents the display location information of the first sample after the k-th iteration. u N represents the adjacent nodes in the bipartite graph structure where the first sample displays its location information. i Let represent the adjacent nodes of the first sample multimedia in the bipartite graph structure; after iteration, the display location information of the first sample and the feature vector of the first sample can be expressed as the following formulas (3) and (4):

[0114]

[0115]

[0116] Where, α k =1 / (K+1), the α k These are the model weights obtained after training.

[0117] By acquiring the first sample display location information, first sample feature vector, and node information corresponding to multiple first sample multimedia, and inputting the first sample display location information and node information into a preset feature extraction model for feature extraction, a predicted feature vector is obtained. Then, based on the first sample feature vector and the predicted feature vector, the preset feature extraction model is trained to obtain the feature extraction model. This feature extraction model can be used to extract the feature vector of multimedia. Since this feature extraction model fully explores the relationship between the first sample display location information and the first sample feature vector during training, when applying this feature extraction model, the accuracy of the target multimedia feature vector can be improved by the relationship between the display location information and the target multimedia feature vector, thereby providing an accurate input basis for the subsequent location prediction model.

[0118] In one exemplary implementation, such as Figure 4 As shown, the steps of inputting the feature vectors of the target multimedia and the feature vectors of historical multimedia into the location prediction model to obtain the target location information of the target multimedia may include:

[0119] In step S41, in the connection layer of the location prediction model, the feature vectors of the target multimedia and the historical multimedia are merged to obtain the semantic features corresponding to the target multimedia and the historical multimedia respectively.

[0120] In the embodiments of this specification, the connection layer in the location prediction model can concatenate the feature vectors of the target multimedia and the feature vectors of the historical multimedia respectively to obtain the semantic features corresponding to the target multimedia and the historical multimedia respectively.

[0121] In step S42, the semantic features of the target multimedia and historical multimedia are fitted in the multilayer perceptron of the location association model to obtain the target location information of the target multimedia.

[0122] In the embodiments of this specification, the multilayer perceptron in the location prediction model can fit the semantic features of the target multimedia and historical multimedia. Specifically, it can calculate the similarity between the semantic features of the target multimedia and the semantic features of historical multimedia (the similarity here may include, but is not limited to, cosine distance, Euclidean distance, Manhattan distance, etc.), and combine the weights corresponding to each historical multimedia to output the association value of the location information of the target multimedia and historical multimedia. Based on this association value, the target location information of the target multimedia is determined from the location information of the historical multimedia. Specifically, the target location information represents the geographical location information of the media content in the target multimedia.

[0123] In addition, it should be noted that, Figure 4 This is just one example of a location prediction model predicting target location information. In practical applications, location prediction models can include more layers, for example, three connection layers.

[0124] By merging the feature vectors of the target multimedia and historical multimedia in the connection layer of the location prediction model, the semantic features corresponding to the target multimedia and historical multimedia are obtained. In the multi-layer perception layer of the location association model, the semantic features of the target multimedia and historical multimedia are fitted to obtain the target location information of the target multimedia. When using the location prediction model to predict the target location information, the similarity between the target multimedia and historical multimedia can be combined with the feature vector merging process, which greatly improves the representation ability of the semantic features on the corresponding target multimedia and historical multimedia, as well as the accuracy of similarity determination, thereby improving the accuracy of subsequent target location information.

[0125] In one exemplary embodiment, the target location information of the target multimedia may include identification information of the target location. For example... Figure 5 As shown, after the step of inputting the feature vectors of the target multimedia and the feature vectors of historical multimedia into the location prediction model to obtain the target location information of the target multimedia, the method may further include:

[0126] In step S51, the location mapping table is obtained.

[0127] In the embodiments of this specification, the location mapping table can represent the correspondence between identification information and location text information, that is, the location mapping table can be indexed to the location text information based on the identification information.

[0128] In step S52, the target location text information corresponding to the identifier information of the target location is determined from the location mapping table.

[0129] In practical applications, location text information can include multi-level locations, such as "China-Beijing-Chaoyang North Road Roast Duck Restaurant," and the corresponding identifier information can also include multi-level identifiers, such as XX1-XX2-XX3. The identifier information in the location mapping table can be hierarchically indexed with the location text information. Accordingly, after determining the identifier information of the target location, the corresponding target location text information can be determined based on the hierarchical index in the location mapping table. For example, first determine the country in the location mapping table based on XX1, then determine the city under that country index based on XX2, and finally determine the specific store under that city based on XX3, thus quickly and accurately determining the complete location text information corresponding to XX1-XX2-XX3.

[0130] Predicting the location identifier of media content in a target multimedia using a location prediction model can improve the processing efficiency of the location prediction model in predicting target location information. Obtaining a location mapping table and determining the target location text information corresponding to the target location identifier from the location mapping table can quickly and accurately determine the target location text information.

[0131] In one exemplary embodiment, the target multimedia information, historical multimedia information, and the location information of the historical multimedia are input into a feature extraction model for feature extraction processing to obtain the feature vectors of the target multimedia and the historical multimedia. Following this step, as follows... Figure 6 As shown, the method may further include:

[0132] In step S61, the target display location information of the target multimedia is obtained.

[0133] In the embodiments of this specification, the target display location information of the target multimedia can refer to the geographical location information of the target multimedia. For example, after the target multimedia is displayed on the terminals of user A and user B, the geographical location information of user A's terminal and user B's terminal can be obtained, and the geographical location information of user A's terminal and user B's terminal is the target display location information.

[0134] In step S62, based on the feature vector of the target multimedia information and the target display position information, the feature vector of the first sample and the display position information of the first sample are updated respectively to obtain the updated feature vector of the first sample and the display position information of the first sample.

[0135] In the embodiments of this specification, the feature vector of the target multimedia information can be added to the first sample feature vector, and the target display position information can be added to the first sample display position information. For example, before the update, the first sample feature vector is the feature vector of N sample multimedias, and the first sample display position information is the display position information corresponding to the N sample multimedias; after the update, the first sample feature vector is the feature vector of N+1 sample multimedias, and the first sample display position information is the display position information corresponding to the N+1 sample multimedias, that is, the target multimedia is treated as one of the N+1 sample multimedias. Alternatively, the feature vector of the target multimedia information can be added to the first sample feature vector, and one sample feature vector in the first sample feature vector can be deleted. The deleted sample feature vector can be the sample feature vector that is furthest from the generation time of the feature vector of the target multimedia information, and this application does not limit this.

[0136] After the target multimedia is displayed, the target display location information of the target multimedia is obtained. Based on the feature vector of the target multimedia information and the target display location information, the feature vector of the first sample and the display location information of the first sample are updated respectively to obtain the updated feature vector of the first sample and the display location information of the first sample. The updated feature vector of the first sample and the display location information of the first sample can be used to update the feature extraction model, efficiently obtain new sample data, greatly save the cost of obtaining sample data during the training of the feature extraction model, and can also continuously improve the accuracy of the parameters in the feature extraction model.

[0137] In one exemplary implementation, such as Figure 7 As shown, obtaining the display location information of the first sample corresponding to multiple first sample multimedia can include:

[0138] In step S71, multiple display location information corresponding to multiple first sample multimedia is determined.

[0139] In the embodiments of this specification, the multiple display location information corresponding to multiple first sample multimedia refers to the full display location information corresponding to each first sample multimedia. In practical applications, users can view sample multimedia through search operations. For example, a user located in Beijing can view videos related to Peking duck restaurants in Beijing, or a user in another country can view videos related to Peking duck restaurants in Beijing from abroad; or the platform can deliver sample multimedia to a user's terminal, and the display location information is the terminal location of the user who searches for and views the sample multimedia, or the terminal location of the user who views the sample multimedia delivered by the platform.

[0140] In step S72, the reference position information corresponding to each first sample multimedia is determined based on the multiple display position information corresponding to the multiple first sample multimedia.

[0141] In the embodiments of this specification, the average display position information of each sample multimedia can be determined based on multiple display position information corresponding to each sample multimedia; or, based on multiple display position information corresponding to each sample multimedia, an area with a relatively dense distribution of display position information can be determined as the target area, and the average display position information in the target area can be further determined based on the display position information in the target area. The average display position information of each sample multimedia or the average display position information in the target area is used as the reference position information corresponding to each sample multimedia.

[0142] In step S73, display position information with a distance from the reference position information less than a preset distance threshold is selected to obtain the first sample display position information corresponding to each sample multimedia.

[0143] In the embodiments of this specification, a preset distance threshold can be set, for example, the preset distance threshold can be 100km, and this application does not limit it. Display location information whose distance from the reference location information is less than the preset distance threshold is selected from multiple display location information corresponding to each first sample multimedia to obtain the first sample display location information corresponding to each sample multimedia; for example... Figure 8 As shown, Figure 8 All points in the region represent the full display location information corresponding to sample multimedia 1. Point 1 represents the reference location information. Region 2 represents the target region in step S72. The points in region 2 are the display location information with a distance of less than 100km from point 1. The points in region 2 represent the display location information of the first sample.

[0144] By determining multiple display position information corresponding to multiple first sample multimedia, and based on the multiple display position information corresponding to multiple first sample multimedia, determining the reference position information corresponding to each first sample multimedia, and filtering the display position information whose distance between the reference position information is less than a preset distance threshold, the first sample display position information corresponding to each first sample multimedia can be obtained. The method of filtering the display position information can remove unstable factors, improve the stability of the first sample display position information, and thus ensure the stability of the input samples during the training of the feature extraction model, thereby improving the accuracy of the feature extraction model.

[0145] In one exemplary implementation, a location prediction model can be obtained by training sample data, such as... Figure 9 As shown, the method may further include:

[0146] In step S91, the second sample data and the corresponding data label are obtained.

[0147] In the embodiments of this specification, the second sample data may include sample multimedia information of multiple second sample multimedia, associated sample multimedia information related to the publishing object information of the second sample multimedia, and associated sample location information of each associated sample multimedia. Data tags may be the geographical location information of the media content in the second sample multimedia. The associated sample location information of each associated sample multimedia may be the geographical location information of the media content of each associated sample multimedia.

[0148] In step S92, the sample multimedia information of the second sample multimedia, the associated sample multimedia information, and the associated sample location information are input into the feature extraction model for feature extraction processing to obtain feature vectors of multiple second sample multimedias and feature vectors of associated multimedias.

[0149] In the embodiments of this specification, the associated sample multimedia of the second sample multimedia may be the same as or different from the first sample multimedia, and this application does not limit this. Specifically, the process of inputting the sample multimedia information of the second sample multimedia, the associated sample multimedia information, and the associated sample location information into the feature extraction model for feature extraction processing can be referred to the process of inputting the target multimedia information, historical multimedia information, and the location information of historical multimedia into the feature extraction model for feature extraction processing, and will not be repeated here.

[0150] In step S93, the feature vectors of multiple second sample multimedias and the feature vectors of associated sample multimedias are input into a preset machine learning model for location prediction processing to obtain the location prediction information of the second sample multimedias.

[0151] In this embodiment, feature vectors of multiple second sample multimedias and feature vectors of associated sample multimedias can be input into a preset machine learning model. After processing by the connection layer, semantic features of the second sample multimedias and semantic features of the associated sample multimedias can be obtained. Then, a multilayer perceptron is used to perform similarity fitting based on the semantic features of the second sample multimedias and the semantic features of the associated sample multimedias. Optionally, the preset machine learning model may also include multiple activation functions, which can be used to improve the data computation speed in the multilayer perceptron and improve sparse activation. The loss function of the preset machine learning model can be the cross-entropy loss function.

[0152] In step S94, a preset machine learning model is trained based on the location prediction information and the corresponding data labels to obtain the location prediction model.

[0153] In the embodiments of this specification, the location prediction information and corresponding data labels can be substituted into the loss function to calculate the loss value, and the preset machine learning model can be trained until the loss value no longer changes or the loss value is less than the threshold to obtain the location prediction model.

[0154] By acquiring second sample data, including sample multimedia information of the second sample multimedia, associated sample multimedia information related to the publishing object information of the second sample multimedia, and associated sample location information of each associated sample multimedia, and inputting the acquired information into a feature extraction model, training data for a preset machine learning model is obtained. This training data is then used to train the preset machine learning model to obtain a location prediction model. This fully explores the positional relationship between the target multimedia and the second sample data, greatly improving the accuracy of the location prediction model in predicting location information, thereby ensuring the accuracy of the target location information of the subsequent target multimedia.

[0155] In one exemplary implementation, such as Figure 10 After the step shown above, where the feature vectors of the target multimedia and historical multimedia are input into the location prediction model to obtain the target location information of the target multimedia, the method may further include:

[0156] In step S101, the sample multimedia information of the second sample multimedia is updated based on the target multimedia information to obtain the updated sample multimedia information of the second sample multimedia.

[0157] In the embodiments of this specification, the target multimedia information can be added to the sample multimedia information of the second sample multimedia to obtain the updated sample multimedia information of the second sample multimedia; alternatively, the target multimedia information can be used to replace one of the sample multimedia information in the second sample multimedia. This application does not limit this.

[0158] In step S102, the data tag is updated based on the target location information to obtain the updated data tag.

[0159] In this embodiment of the specification, the same update method as in step S101 can be used to update the data tag using the target location information. That is, when the target multimedia information is added to the sample multimedia information of the second sample multimedia, the target location information can be added to the data tag to obtain the updated data tag; when the target multimedia information is used to replace one of the sample multimedia information in the second sample multimedia, the target location information can be used to replace the corresponding data tag, and the data tag is the data tag corresponding to the replaced sample multimedia information.

[0160] By updating the sample multimedia information of the second sample multimedia based on the target multimedia information, the updated sample multimedia information of the second sample multimedia is obtained. The data labels are updated based on the target location information to obtain the updated data labels. This method can efficiently obtain new sample data, greatly saving the cost of obtaining sample data during the training of the location prediction model, and can also continuously improve the accuracy of the parameters in the location prediction model.

[0161] Figure 11 This is a block diagram of a multimedia location information processing apparatus according to an exemplary embodiment. (Reference) Figure 11 The device may include:

[0162] The target information acquisition module 111 is configured to acquire the target multimedia information and the publishing object information of the target multimedia.

[0163] The module 112 for determining the information to be extracted is configured to determine the historical multimedia associated with the publishing object information, as well as the historical multimedia information and the historical location information of the historical multimedia; the historical location information is the geographical location information of the media content in the historical multimedia.

[0164] The target feature extraction module 113 is configured to input the target multimedia information, historical multimedia information, and the location information of historical multimedia into the feature extraction model, perform feature extraction processing, and obtain the feature vector of the target multimedia and the feature vector of the historical multimedia.

[0165] The location information prediction module 114 is configured to input the feature vector of the target multimedia and the feature vector of the historical multimedia into the location prediction model to obtain the target location information of the target multimedia, wherein the target location information is the geographical location information of the media content in the target multimedia.

[0166] By acquiring the target multimedia information and the publishing object information of the target multimedia, determining the historical multimedia associated with the publishing object information, as well as the historical multimedia information and historical location information of the historical multimedia, and using a feature extraction model to directly extract features from the target multimedia information, historical multimedia information, and the location information of the historical multimedia, the feature vectors of the target multimedia and the historical multimedia can be obtained, which can improve the efficiency and quality of feature vector acquisition. Furthermore, the target location information of the target multimedia is predicted based on the feature vectors of the historical multimedia associated with the publishing object information and the feature vector of the target multimedia. By mining the relationship between the target multimedia and the historical multimedia associated with the publishing object, the target location information of the target multimedia can be determined, which can improve the accuracy of the target location information.

[0167] In one exemplary embodiment, the device may further include:

[0168] The first sample information acquisition module is configured to acquire the first sample display location information, first sample feature vector and node information corresponding to multiple first sample multimedia. The first sample display location information is the display location information of multiple first sample multimedia, and the node information is the adjacency information of the first node in the first sample display location information and the second node in the first sample multimedia.

[0169] The sample feature extraction module is configured to input the display location information and node information of the first sample into a preset feature extraction model to extract features and obtain a predicted feature vector.

[0170] The first training module is configured to train a preset feature extraction model based on the first sample feature vector and the predicted feature vector, thereby obtaining the feature extraction model.

[0171] In one exemplary embodiment, the location information prediction module may include:

[0172] The feature merging unit is configured to perform feature merging processing on the feature vectors of the target multimedia and the feature vectors of the historical multimedia to obtain the semantic features corresponding to the target multimedia and the historical multimedia respectively.

[0173] The fitting unit is configured to fit the semantic features of the target multimedia and historical multimedia to obtain the target location information of the target multimedia.

[0174] In one exemplary embodiment, the device may further include:

[0175] The location mapping table acquisition module is configured to acquire a location mapping table, which represents the correspondence between identification information and location text information.

[0176] The location text information determination module is configured to determine the target location text information corresponding to the identifier information of the target location from the location mapping table.

[0177] In one exemplary embodiment, the device may further include:

[0178] The display location information acquisition module is configured to acquire the target display location information of the target multimedia after the target multimedia is displayed. The target display location information is the display location information of the target multimedia.

[0179] The first update module is configured to update the first sample feature vector and the first sample display location information based on the feature vector of the target multimedia information and the target display location information, respectively, to obtain the updated first sample feature vector and the first sample display location information.

[0180] In one exemplary embodiment, the sample information acquisition module may include:

[0181] The display location information acquisition unit is configured to determine multiple display location information corresponding to multiple first sample multimedia;

[0182] The reference position information acquisition unit is configured to determine the reference position information corresponding to each first sample multimedia based on the multiple display position information corresponding to the multiple first sample multimedia.

[0183] The sample information acquisition unit is configured to filter out display position information whose distance from the reference position information is less than a preset distance threshold, and obtain the first sample display position information corresponding to each sample multimedia.

[0184] In one exemplary embodiment, the device may further include:

[0185] The second sample information acquisition module is configured to acquire second sample data and corresponding data tags. The second sample data includes sample multimedia information of multiple second sample multimedia, associated sample multimedia information related to the publishing object information of the second sample multimedia, and associated sample location information of each associated sample multimedia. The data tags are the geographical location information of the media content in the second sample multimedia. The associated sample location information of each associated sample multimedia is the geographical location information of the media content of each associated sample multimedia.

[0186] The sample feature determination module is configured to input the sample multimedia information of the second sample multimedia, the associated sample multimedia information, and the associated sample location information into the feature extraction model, perform feature extraction processing, and obtain feature vectors of multiple second sample multimedia and feature vectors of associated multimedia.

[0187] The prediction information determination module is configured to input the feature vectors of multiple second sample multimedias and the feature vectors of associated sample multimedias into a preset machine learning model for position prediction processing to obtain the position prediction information of the second sample multimedias.

[0188] The second training module is configured to train a preset machine learning model based on location prediction information and corresponding data labels to obtain a location prediction model.

[0189] In one exemplary embodiment, the device may further include:

[0190] The second update module is configured to update the sample multimedia information of the second sample multimedia based on the target multimedia information, so as to obtain the updated sample multimedia information of the second sample multimedia.

[0191] The third update module is configured to update the data tags based on the target location information to obtain the updated data tags.

[0192] Figure 12 This is a block diagram illustrating an electronic device for a multimedia location information processing method according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 12 As shown, the electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for constructing virtual objects.

[0193] Those skilled in the art will understand that Figure 12 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0194] In an exemplary embodiment, an electronic device is also provided, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the multimedia location information processing method as described in the embodiments of this disclosure.

[0195] In an exemplary embodiment, a computer-readable storage medium is also provided, which, when executed by a processor of an electronic device, enables the electronic device to perform the virtual object construction method of the present disclosure embodiments. The computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.

[0196] In an exemplary embodiment, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the multimedia location information processing method of the present disclosure embodiments.

[0197] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0198] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0199] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method of processing multimedia location information, characterized by, The method comprises the following steps: obtaining target multimedia information of a target multimedia and publishing object information; determining historical multimedia associated with the publishing object information, historical multimedia information of the historical multimedia, and historical location information of the historical multimedia; the historical location information is geographical location information of media content in the historical multimedia; inputting the target multimedia information, the historical multimedia information, and the location information of the historical multimedia into a feature extraction model to perform feature extraction processing, to obtain a feature vector of the target multimedia and a feature vector of the historical multimedia; inputting the feature vector of the target multimedia and the feature vector of the historical multimedia into a location prediction model to obtain target location information of the target multimedia; the target location information is geographical location information of media content in the target multimedia; a training method of the feature extraction model comprises the following steps: obtaining first sample display location information, first sample feature vectors, and node information corresponding to a plurality of first sample multimedia; the node information is adjacency information of a bipartite graph structure composed of first nodes and second nodes; the first nodes are the first sample display location information, and the second nodes are the first sample multimedia; inputting the first sample display location information and the node information into a preset feature extraction model to perform feature extraction, to obtain a predicted feature vector; training the preset feature extraction model based on the first sample feature vectors and the predicted feature vector, to obtain the feature extraction model.

2. The method of claim 1, wherein, The step of inputting the feature vector of the target multimedia and the feature vector of the historical multimedia into the location prediction model to obtain the target location information of the target multimedia comprises the following steps: in a connection layer of the location prediction model, performing feature merging processing on the feature vector of the target multimedia and the feature vector of the historical multimedia to obtain respective semantic features of the target multimedia and the historical multimedia; in a multilayer perceptron of the location prediction model, performing fitting processing on the semantic features of the target multimedia and the historical multimedia to obtain the target location information of the target multimedia.

3. The method of claim 1, wherein, The target location information of the target multimedia comprises identification information of a target location; after the step of inputting the feature vector of the target multimedia and the feature vector of the historical multimedia into the location prediction model to obtain the target location information of the target multimedia, the method further comprises the following steps: obtaining a location mapping table, the location mapping table representing a corresponding relationship between identification information and location text information; determining target location text information corresponding to the identification information of the target location from the location mapping table.

4. The method of claim 1, wherein, after the step of inputting the target multimedia information, the historical multimedia information, and the location information of the historical multimedia into the feature extraction model to perform feature extraction processing to obtain the feature vector of the target multimedia and the feature vector of the historical multimedia, the method further comprises the following steps: obtaining target display location information of the target multimedia; the target display location information is geographical location information of displaying the target multimedia; The first sample feature vector and the first sample display position information are updated respectively based on the feature vector of the target multimedia information and the target display position information, to obtain an updated first sample feature vector and an updated first sample display position information.

5. The method of claim 1, wherein, The obtaining of the first sample display position information corresponding to the plurality of first sample multimedia comprises: determining a plurality of display position information corresponding to the plurality of first sample multimedia; determining reference position information corresponding to each first sample multimedia according to the plurality of display position information corresponding to the plurality of first sample multimedia; screening display position information with a distance less than a preset distance threshold from the reference position information to obtain first sample display position information corresponding to each first sample multimedia.

6. The method of claim 1, wherein, The method further comprises: obtaining second sample data and corresponding data labels, wherein the second sample data comprises sample multimedia information of a plurality of second sample multimedia, associated sample multimedia information associated with publishing object information of the second sample multimedia, and associated sample position information of the associated sample multimedia information; the data label is geographical position information of media content in the second sample multimedia; and the associated sample position information of the associated sample multimedia information is geographical position information of media content of the associated sample multimedia; inputting the sample multimedia information of the second sample multimedia, the associated sample multimedia information, and the associated sample position information into the feature extraction model to perform feature extraction processing, to obtain feature vectors of the plurality of second sample multimedia and feature vectors of the associated sample multimedia; inputting the feature vectors of the plurality of second sample multimedia and the feature vectors of the associated sample multimedia into a preset machine learning model to perform position prediction processing, to obtain position prediction information of the second sample multimedia; training the preset machine learning model based on the position prediction information and the corresponding data labels, to obtain the position prediction model.

7. The method of claim 6, wherein, After the step of inputting the feature vector of the target multimedia and the feature vector of the historical multimedia into the position prediction model to obtain target position information of the target multimedia, the method further comprises: updating sample multimedia information of the second sample multimedia based on the target multimedia information, to obtain updated sample multimedia information of the second sample multimedia; updating the data label based on the target position information, to obtain an updated data label.

8. A processing apparatus of multimedia position information, characterized by, comprises: a target information acquisition module configured to acquire target multimedia information and publishing object information of a target multimedia; a to-be-extracted information determination module configured to determine historical multimedia associated with the publishing object information, historical multimedia information of the historical multimedia, and historical position information of the historical multimedia; the historical position information is geographical position information of media content in the historical multimedia; The target feature extraction module is configured to input the target multimedia information, the historical multimedia information, and position information of the historical multimedia into a feature extraction model, perform feature extraction processing, and obtain a feature vector of the target multimedia and a feature vector of the historical multimedia. The position information prediction module is configured to input the feature vector of the target multimedia and the feature vector of the historical multimedia into a position prediction model to obtain target position information of the target multimedia, the target position information being geographical position information of media content in the target multimedia. The first sample information acquisition module is configured to acquire first sample display position information, first sample feature vectors, and node information corresponding to a plurality of first sample multimedia, the node information being adjacency information of a bipartite graph structure formed by first nodes and second nodes. The first nodes are the first sample display position information, and the second nodes are the first sample multimedia. The sample feature extraction module is configured to input the first sample display position information and the node information into a preset feature extraction model, perform feature extraction, and obtain predicted feature vectors. The first training module is configured to train the preset feature extraction model based on the first sample feature vectors and the predicted feature vectors, and obtain the feature extraction model.

9. The apparatus of claim 8, wherein, The position information prediction module includes: The feature merging unit is configured to perform feature merging processing on the feature vector of the target multimedia and the feature vector of the historical multimedia to obtain semantic features corresponding to the target multimedia and the historical multimedia, respectively. The fitting unit is configured to perform fitting processing on the semantic features of the target multimedia and the historical multimedia in a multi-layer perceptron of the position prediction model to obtain the target position information of the target multimedia.

10. The apparatus of claim 8, wherein, The device further includes: The position mapping table acquisition module is configured to acquire a position mapping table, the position mapping table representing a corresponding relationship between identification information and position text information. The position text information determination module is configured to determine target position text information corresponding to identification information of the target position from the position mapping table.

11. The apparatus of claim 8, wherein, The device further includes: The display position information acquisition module is configured to acquire target display position information of the target multimedia after the target multimedia is displayed, the target display position information being display position information of the target multimedia. The first updating module is configured to update the first sample feature vectors and the first sample display position information based on the feature vector of the target multimedia information and the target display position information, respectively, to obtain updated first sample feature vectors and first sample display position information.

12. The apparatus of claim 8, wherein, The sample information acquisition module includes: The display position information acquisition unit is configured to determine a plurality of display position information corresponding to the plurality of first sample multimedia. The reference position information acquisition unit is configured to determine reference position information corresponding to each first sample multimedia according to the plurality of display position information corresponding to the plurality of first sample multimedia. The sample information acquisition unit is configured to filter out the display position information with a distance less than a preset distance threshold from the reference position information, and obtain first sample display position information corresponding to each sample multimedia.

13. The apparatus of claim 8, wherein, The device further comprises: The second sample information acquisition module is configured to acquire second sample data and a corresponding data label. The second sample data includes sample multimedia information of a plurality of second sample multimedia, associated sample multimedia information associated with publication object information of the second sample multimedia, and associated sample position information of the associated sample multimedia information. The data label is geographic position information of media content in the second sample multimedia. The associated sample position information of the associated sample multimedia information is geographic position information of respective media content of the associated sample multimedia. The sample feature determination module is configured to input the sample multimedia information of the second sample multimedia, the associated sample multimedia information, and the associated sample position information into the feature extraction model to perform feature extraction processing, and obtain feature vectors of the plurality of second sample multimedia and feature vectors of the associated sample multimedia. The prediction information determination module is configured to input the feature vectors of the plurality of second sample multimedia and the feature vectors of the associated sample multimedia into a preset machine learning model to perform position prediction processing, and obtain position prediction information of the second sample multimedia. The second training module is configured to train the preset machine learning model based on the position prediction information and the corresponding data label, and obtain the position prediction model.

14. The apparatus of claim 13, wherein, The device further comprises: The second update module is configured to update the sample multimedia information of the second sample multimedia based on the target multimedia information, and obtain updated sample multimedia information of the second sample multimedia. The third update module is configured to update the data label based on the target position information, and obtain an updated data label.

15. An electronic device, comprising: comprises: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the processing method of multimedia position information according to any one of claims 1 to 7.

16. A computer readable storage medium characterized by: When the instructions in the computer readable storage medium are executed by the processor of the electronic device, the electronic device can execute the processing method of multimedia position information according to any one of claims 1 to 7.

17. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instructions are executed by the processor to implement the processing method of multimedia position information according to any one of claims 1 to 7.