Fingerprint generation method, device, server and storage medium

By obtaining the resource features of text, images and videos in multimedia resources, determining the minimum multimodal fusion distance, and generating multimodal fingerprints, the problem of insufficient accuracy in recommending newly released resources is solved, and accurate recommendation of multimedia resources is achieved.

CN114282023BActive Publication Date: 2025-10-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110969825.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-23
Publication Date
2025-10-10
Estimated Expiration
2041-08-23

Smart Images

  • Figure CN114282023B_ABST
    Figure CN114282023B_ABST
Patent Text Reader

Abstract

The application discloses a fingerprint generation method and device, a server and a storage medium, and belongs to the technical field of networks. Through the technical scheme provided by the embodiment of the application, in the generation process, the corresponding resource set is determined based on the resource characteristics of each different mode, each resource set contains resources matched with the multimedia resource under the corresponding mode, and the multimedia resource with the most similar content is determined based on the comprehensive matching degree of the historical multimedia resources in the resource set and the multimedia resource in multiple modes, so that the determination of the multi-modal fingerprint is performed. The multi-modal fingerprint determined by the above method fuses resource content and is easy to store and calculate, can not only play a role in identifying multimedia resources, but also can learn some correlation information between multimedia resources through the multi-modal fingerprint, thereby improving the accuracy of recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of network, in particular to a fingerprint generation method and device, a server and a storage medium. BACKGROUND

[0002] Modality refers to the source or form of data. When recommending multimedia resources containing multiple modalities, the resource identifier of the multimedia resource, that is, the fingerprint, is usually taken as a factor affecting recommendation.

[0003] In related technologies, the fingerprint of the multimedia resource is usually generated based on the upload time of the multimedia resource. After the multimedia resource is published, the future performance of the multimedia resource can be predicted based on the historical performance data of the multimedia resource to improve the accuracy of recommendation.

[0004] However, since the fingerprint of the multimedia resource is obtained based on time, for some newly published multimedia resources, there is no historical performance data, so the accuracy of recommendation cannot be improved accordingly. Therefore, there is an urgent need for a fingerprint generation method that can improve the accuracy of recommendation. SUMMARY

[0005] The embodiments of the present application provide a fingerprint generation method, device, server and storage medium, which can improve the accuracy of recommendation. The technical scheme is as follows:

[0006] On the one hand, a fingerprint generation method is provided, which comprises:

[0007] Based on the data of at least two modalities in the multimedia resource, resource features of the at least two modalities are obtained, the multimedia resource comprising data of at least two modalities of text, image and video;

[0008] Based on the resource features of the at least two modalities, resource sets of the at least two modalities are obtained, each resource set of the modality comprising historical multimedia resources matching the resource features of the corresponding modality;

[0009] Based on the resource features of the at least two modalities of the multimedia resource, a minimum multi-modality fusion distance is determined, the minimum multi-modality fusion distance corresponding to the historical multimedia resource with the maximum comprehensive matching degree between the multimedia resource on each modality in the resource set of the at least two modalities;

[0010] In the case where the minimum multi-modality fusion distance is greater than the distance threshold, the initial multi-modality fingerprint of the multimedia resource is determined as the multi-modality fingerprint of the multimedia resource, the initial multi-modality fingerprint being determined based on the stored multi-modality fingerprint in the multimedia resource library.

[0011] On the one hand, a fingerprint generation device is provided, which comprises:

[0012] A feature acquisition module, configured to acquire resource features of at least two modalities based on data of at least two modalities in a multimedia resource, wherein the multimedia resource includes data of at least two modalities of text, image, and video;

[0013] A collection acquisition module, configured to acquire resource collections of the at least two modalities based on the resource characteristics of the at least two modalities, wherein the resource collection of each modality includes historical multimedia resources that match the resource characteristics of the corresponding modality;

[0014] a distance determination module, configured to determine a minimum multimodal fusion distance based on resource features of the at least two modalities of the multimedia resource, the minimum multimodal fusion distance corresponding to a historical multimedia resource in the resource set of the at least two modalities having the greatest degree of comprehensive matching with the multimedia resource in each modality;

[0015] The first fingerprint determination module is used to determine the initial multimodal fingerprint of the multimedia resource as the multimodal fingerprint of the multimedia resource when the minimum multimodal fusion distance is greater than the distance threshold, and the initial multimodal fingerprint is determined based on the multimodal fingerprint stored in the multimedia resource library.

[0016] In one possible implementation, the feature acquisition module is used to:

[0017] Based on the data of the at least two modalities, a feature extraction network of the corresponding modality is called to perform feature extraction to obtain resource features of the at least two modalities.

[0018] In one aspect, a server is provided, comprising one or more processors and one or more memories, wherein the one or more memories store at least one computer program, and the one or more processors load and execute the computer program to implement the fingerprint generation method described above.

[0019] In one aspect, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium. The computer program is loaded and executed by a processor to implement the above-mentioned fingerprint generation method.

[0020] On the one hand, a computer program product or computer program is provided, which includes program code, the program code is stored in a computer-readable storage medium, a processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device performs the above-mentioned fingerprint generation method.

[0021] Through the technical solution provided in the embodiments of the present application, during the generation process, corresponding resource sets are determined based on the resource characteristics of each different modality, and each resource set contains resources that match the multimedia resource in the corresponding modality. Therefore, based on the comprehensive matching degree between the historical multimedia resources in these resource sets and the multimedia resource in multiple modalities, the multimedia resource with the most similar content is determined, thereby determining the multimodal fingerprint. The multimodal fingerprint determined by the above method integrates the resource content and is easy to store and calculate. It can not only play the role of identifying multimedia resources, but also learn some correlation information between multimedia resources through the multimodal fingerprint, thereby improving the accuracy of recommendation. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] Figure 1 This is a schematic diagram of an implementation environment of a fingerprint generation method provided in an embodiment of the present application;

[0024] Figure 2 This is a flow chart of a fingerprint generation method provided in an embodiment of the present application;

[0025] Figure 3 This is a flow chart of a feature extraction network training method provided in an embodiment of the present application;

[0026] Figure 4 This is a feature extraction network training architecture diagram provided in an embodiment of the present application;

[0027] Figure 5 This is a flow chart of a fingerprint generation method provided in an embodiment of the present application;

[0028] Figure 6 This is a schematic structural diagram of a fingerprint generation device provided in an embodiment of the present application;

[0029] Figure 7 This is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0031] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on the quantity and execution order.

[0032] In this application, the term "at least one" means one or more, and the meaning of "plurality" means two or more. For example, multiple modes refer to two or more modes.

[0033] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form a resource pool for on-demand, flexible and convenient use. Cloud computing technology will become a key support. The backend services of technical network systems, such as video websites, image websites, and more portals, require extensive computing and storage resources. With the rapid development and application of the internet industry, every item will likely have its own unique identification mark, which will need to be transmitted to the backend system for logical processing. Different levels of data will be processed separately. All kinds of industry data require strong system support, which can only be achieved through cloud computing.

[0034] Cloud computing refers to the delivery and usage model of IT infrastructure, enabling on-demand, scalable access to resources over the internet. Broadly speaking, cloud computing refers to the delivery and usage model of services, enabling on-demand, scalable access to services over the internet. These services can be IT-related, software-related, internet-related, or other services. Cloud computing is the product of the convergence of traditional computer and network technologies, including grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing. Driven by the growth of the internet, real-time data streams, and the diversification of connected devices, as well as the demand for search services, social networking, mobile commerce, and open collaboration, cloud computing has rapidly developed. Unlike previous parallel and distributed computing approaches, the emergence of cloud computing will fundamentally revolutionize the entire internet model and enterprise management model.

[0035] A database, in short, can be thought of as a digital filing cabinet—a place where electronic files are stored, allowing users to add, query, update, and delete data. A database is a collection of data stored in a specific way, shared by multiple users, with minimal redundancy, and independent of applications.

[0036] Big data refers to collections of data that cannot be captured, managed, and processed within a specific timeframe using conventional software tools. These massive, rapidly growing, and diverse information assets require new processing models to enhance decision-making, insight discovery, and process optimization. With the advent of the cloud era, big data has attracted increasing attention. Big data requires specialized technologies to efficiently process large amounts of time-sensitive data. Technologies suitable for big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the internet, and scalable storage systems.

[0037] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning.

[0038] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge sub-models to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.

[0039] Deep learning (DL) is a subcategory of machine learning. Inspired by the workings of the human brain, it utilizes deep neural networks to solve feature expression problems. Deep neural networks themselves are not a new concept and can be understood as neural network structures consisting of multiple hidden layers. To improve the training effectiveness of deep neural networks, adjustments have been made to aspects such as neuronal connection methods and activation functions. The goal is to build and simulate neural networks that analyze and learn like the human brain, mimicking the brain's mechanisms to interpret data such as text, images, and sound.

[0040] Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain is essentially a decentralized database, a series of data blocks generated by cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.

[0041] Figure 1 This is a schematic diagram of the implementation environment of a fingerprint generation method provided in an embodiment of the present application, see Figure 1 , the implementation environment may include a terminal 110 and a server 120.

[0042] Optionally, the terminal 110 is a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. An application supporting multimedia resource sharing is installed and running on the terminal 110.

[0043] Optionally, the terminal 110 can be connected to the server 120 via a wireless network or a wired network. The terminal 110 can send multimedia resources to the server 120, and the server 120 processes the received multimedia resources, for example, analyzes and publishes them.

[0044] Optionally, server 120 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, distribution networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.

[0045] Optionally, the terminal 110 and the server 120 can serve as nodes on the blockchain system to store relevant data involved in the fingerprint generation process, such as multimedia resources and fingerprints.

[0046] It should be noted that the multimedia resources processed by the embodiment of the present application can be uploaded to the server by the terminal or obtained by the server itself, and the embodiment of the present application does not limit this.

[0047] After introducing the implementation environment of the embodiment of the present application, the application scenario of the embodiment of the present application will be introduced in combination with the above implementation environment. It should be noted that in the following description, the terminal is also the above terminal 110, and the server is also the above server 120.

[0048] The fingerprint generation method provided in the embodiment of the present application can be applied to the scenario of multimedia resource recommendation. That is, through the fingerprint generation method provided in the embodiment of the present application, after the terminal uploads the multimedia resource to the server, the server processes the multimedia resource and generates a multimodal fingerprint. The rich relevant information contained in the multimodal fingerprint is used to achieve accurate recommendation of the multimedia resource. In addition, the fingerprint generation method provided in the embodiment of the present application can also be applied to other scenarios that require the generation of multimodal feature identifiers, such as in the scenario of e-commerce product recommendation. Of course, with the development of science and technology, the fingerprint generation method provided in the embodiment of the present application can also be applied to other scenarios that require the generation of multimodal feature identifiers, and the embodiment of the present application does not limit this.

[0049] In the context of multimedia resource recommendation, for example, the multimedia resource is an advertisement. An advertiser uploads the advertisement via terminal 110 in the aforementioned implementation environment. Server 120 receives the advertisement and processes the ad's text, image, and video data to generate and store a corresponding multimodal fingerprint. Based on the ad's multimodal fingerprint, server 120 makes accurate recommendations for the ad.

[0050] In the scenario of e-commerce product recommendation, the product seller can use the terminal 110 to upload the product to the server 120. The server 120 receives the product, processes the product name, picture, price and other data, generates and stores the multimodal fingerprint corresponding to the product, and accordingly, makes accurate recommendations based on the multimodal fingerprint of the product.

[0051] After introducing the implementation environment and application scenarios of the embodiments of the present application, the fingerprint generation method provided by the embodiments of the present application is described below. Figure 2 This is a flow chart of a fingerprint generation method provided by an embodiment of the present application, taking the execution subject as a server as an example, see Figure 2 , methods include:

[0052] 201. The server obtains resource features of at least two modalities based on data of at least two modalities in a multimedia resource, where the multimedia resource includes data of at least two modalities of text, image, and video.

[0053] The multimedia resource is a resource that includes data in at least two modalities. The at least two modalities may include the following combinations: text and image; text and video; image and video; or text, image, and video. This embodiment of the application does not limit the specific form of the sample multimedia resource.

[0054] The server uses the corresponding feature extraction model to extract features based on the data of each modality included in the multimedia resource, thereby obtaining resource features of the at least two modalities. The feature extraction model used will vary depending on the modalities included in the multimedia resource, and will not be described in detail here.

[0055] Taking a multimedia resource containing two modal data, text and image, as an example, when executing step 201, the server calls a text feature extraction network and an image feature extraction network to process the text and image respectively to obtain the text features and image features of the multimedia resource.

[0056] In some embodiments, the server is a node in the blockchain system, and the above-mentioned multimedia resources are stored on the blockchain of the blockchain system. The server obtains the multimedia resources from the blockchain configured by itself and realizes data sharing through blockchain-based multimedia resource storage.

[0057] 202. The server obtains resource sets of at least two modalities based on the resource features of the at least two modalities, where each resource set of the modality includes historical multimedia resources that match the resource features of the corresponding modality.

[0058] Taking a multimedia resource containing two modal data, text and image, as an example, the resource features obtained include text features and image features. Accordingly, when executing step 202, based on the text features and image features, a resource set corresponding to the text features and a resource set corresponding to the image features are obtained respectively.

[0059] 203. The server determines a minimum multimodal fusion distance based on resource characteristics of at least two modalities of the multimedia resource, where the minimum multimodal fusion distance corresponds to a historical multimedia resource in the resource set of the at least two modalities that has the greatest comprehensive matching degree with the multimedia resource in each modality.

[0060] The comprehensive matching degree between two multimedia resources is determined based on the similarity of the two multimedia resources in multiple modalities.

[0061] 204. When the minimum multimodal fusion distance is greater than the distance threshold, the server determines the initial multimodal fingerprint of the multimedia resource as the multimodal fingerprint of the multimedia resource, where the initial multimodal fingerprint is determined based on the multimodal fingerprints stored in the multimedia resource library.

[0062] Among them, the distance threshold is used to determine the minimum multimodal fusion distance. In different application scenarios, the distance threshold can be set according to the accuracy requirements of the multimodal fingerprint, and the embodiments of the present application do not limit this.

[0063] Through the technical solution provided in the embodiments of the present application, during the generation process, corresponding resource sets are determined based on the resource characteristics of each different modality, and each resource set contains resources that match the multimedia resource in the corresponding modality. Therefore, based on the comprehensive matching degree between the historical multimedia resources in these resource sets and the multimedia resource in multiple modalities, the multimedia resource with the most similar content is determined, thereby determining the multimodal fingerprint. The multimodal fingerprint determined by the above method integrates the resource content and is easy to store and calculate. It can not only play the role of identifying multimedia resources, but also learn some correlation information between multimedia resources through the multimodal fingerprint, thereby improving the accuracy of recommendation.

[0064] Next, the training process of the text feature extraction network, image feature extraction network, and video feature extraction network involved in the embodiments of this application is introduced. Figure 3 This is a flow chart of the feature extraction network training method involved in the embodiment of this application, see Figure 3 Taking multimedia resources containing text, images, and videos as an example, the training process includes:

[0065] 301. The server obtains multiple sample multimedia resources, each of which includes data of at least two of the three modalities of text, image, and video.

[0066] The sample multimedia resource is a resource that includes data in at least two modalities. The at least two modalities may include the following combinations: text and pictures; text and video; pictures and videos; text, pictures, and videos. The embodiment of the present application does not limit the specific form of the sample multimedia resource. For example, in an advertising scenario, the sample multimedia resource may be an advertisement that includes advertising text, advertising pictures, and advertising videos. The advertising text may be a description or slogan of the advertising content, the advertising picture may be a picture of the object being advertised, and the advertising video may be a video of the object being advertised.

[0067] In an embodiment of the present application, the server can obtain the multiple sample multimedia resources from the sample resource database. The multiple sample multimedia resources can be resources that have been published on any platform. For example, in an advertising scenario, the multiple sample multimedia resources are multiple published advertisements.

[0068] 302. The server obtains each tag of each sample multimedia resource.

[0069] Tags are used to identify information such as the data type of the sample multimedia resource. For text, tags can be semantic interpretations, such as "happy" or "sad." For images, tags can be specific objects, such as cats, cars, or whether they contain faces. For videos, tags can be actions or scenes, such as eating or being at the beach.

[0070] It should be noted that the label needs to be manually marked in advance, or the label can be obtained by calling the resource classification model through the server.

[0071] 303. The server trains a text feature extraction network, an image feature extraction network, and a video feature extraction network respectively based on the multiple sample multimedia resources and corresponding labels.

[0072] Among them, the text feature extraction network is used to extract features from the text in the input sample multimedia resources to obtain corresponding text features; the image feature extraction network is used to extract features from the images in the input sample multimedia resources to obtain corresponding image features; and the video feature extraction network is used to extract features from the video in the input sample multimedia resources to obtain corresponding video features.

[0073] Taking the training process of the text feature extraction network as an example, the training process includes multiple iterations. In the i-th iteration, i is an integer greater than 1. The text in the sample multimedia resource is input into the text feature extraction network obtained in the i-1-th iteration. The text is processed by the text feature extraction network to obtain the sample text features. Based on the sample text features and labels, the loss function is calculated. When the loss function meets the training end conditions, the training is completed and the text feature extraction network used in this iteration is obtained. When the loss function does not meet the training end conditions, the network parameters of the text feature extraction network are adjusted, and the i+1-th iteration process is performed based on the adjusted text feature extraction network until any subsequent iteration process meets the training end conditions. Among them, the training end conditions are that the loss function reaches the minimum value or the number of iterations reaches the target number or other conditions, and the embodiments of the present application do not limit this.

[0074] In some embodiments, the text feature extraction network can be a Bidirectional Encoder Representations from Transformers (BERT) model.

[0075] Taking the training process of the image feature extraction network as an example, the training process includes multiple iterations. During the i-th iteration, the image in the sample multimedia resource is input into the image feature extraction network obtained by the i-1-th iteration. The image is processed by the image feature extraction network to obtain the sample image features. Based on the sample image features and labels, the loss function is calculated. When the loss function meets the training end conditions, the training is completed and the image feature extraction network used in this iteration is obtained. When the loss function does not meet the training end conditions, the network parameters of the image feature extraction network are adjusted, and the i+1-th iteration process is performed based on the adjusted image feature extraction network until any subsequent iteration process meets the training end conditions. Among them, the training end conditions are that the loss function reaches the minimum value or the number of iterations reaches the target number or other conditions, and the embodiments of the present application do not limit this.

[0076] In some embodiments, the image feature extraction network may be a deep residual network (ResNet).

[0077] Taking the training process of the video feature extraction network as an example, the training process includes multiple iterations. During the i-th iteration, the video in the sample multimedia resource is input into the video feature extraction network obtained by the i-1-th iteration. The video is processed by the video feature extraction network to obtain the sample video features. Based on the sample video features and labels, the loss function is calculated. When the loss function meets the training end conditions, the training is completed and the video feature extraction network used in this iteration is obtained. When the loss function does not meet the training end conditions, the network parameters of the video feature extraction network are adjusted, and the i+1-th iteration process is performed based on the adjusted video feature extraction network until any subsequent iteration process meets the training end conditions. Among them, the training end conditions are that the loss function reaches the minimum value or the number of iterations reaches the target number or other conditions, and the embodiments of the present application do not limit this.

[0078] In some embodiments, the video feature extraction network may be a Vision Transformer (ViT) model.

[0079] The following combination Figure 4 For an explanation of the feature extraction network training process, see Figure 4, input the data from the sample multimedia resources with labels into the feature extraction network, process the data through the feature extraction network to obtain the sample resource features, calculate the loss function based on the sample resource features and labels, and when the loss function meets the training end condition, the training is completed and the feature extraction network used for this iteration is obtained. If the loss function does not meet the training end condition, the network parameters of the feature extraction network are adjusted, and the i+1th iteration process is performed based on the adjusted feature extraction network until any subsequent iteration process meets the training end condition. The training end condition is that the loss function reaches the minimum value or the number of iterations reaches the target number or other conditions.

[0080] The above steps 301 to 303 introduce the training process of the text feature extraction network, the image feature extraction network, and the video feature extraction network used in the multimodal fingerprint generation method provided in the embodiment of the present application. The multimodal fingerprint generation method provided in the embodiment of the present application will be described in detail with reference to some examples.

[0081] Figure 5 This is a flow chart of a fingerprint generation method provided in an embodiment of the present application, see Figure 5 , taking multimedia resources including text, pictures and videos as an example, the methods include:

[0082] 501. A terminal sends multimedia resources to a server. The multimedia resources include text, pictures, and videos.

[0083] In an embodiment of the present application, the multimedia resource includes data in three modalities: text, picture, and video. In some embodiments, the multimedia resource may also include data in any two modalities among the multiple modalities.

[0084] In an embodiment of the present application, the terminal is the advertiser's terminal. The advertiser can upload multimedia resources to the server by performing corresponding operations on the terminal, so that the multimedia resources are stored in the server, and the server performs subsequent processing, such as analysis or publication.

[0085] 502. The server receives the multimedia resource and obtains an initial multimodal fingerprint of the multimedia resource.

[0086] The initial multimodal fingerprint is an integer. For a newly received multimedia resource, the initial multimodal fingerprint set by the server is a global variable that can be used as identification information of the multimedia resource.

[0087] In some embodiments, the process of the server obtaining the initial multimodal fingerprint of the multimedia resource includes: determining the initial multimodal fingerprint of the multimedia resource based on the number of multimodal fingerprints stored in a multimedia resource library; optionally, the server may also determine the initial multimodal fingerprint of the multimedia resource based on the maximum multimodal fingerprint among the fingerprints stored in the multimedia resource library. The multimedia resource library is used to store historical multimedia resources and the multimodal fingerprints corresponding to the historical multimedia resources.

[0088] It should be noted that, in the multimedia resource library, different historical multimedia resources may correspond to the same multimodal fingerprint, that is, the number of historical multimedia resources and the number of multimodal fingerprints are not necessarily the same.

[0089] If the number of multimodal fingerprints stored in the multimedia resource library is 0, the value of the initial multimodal fingerprint of the multimedia resource is set to 0. If the number of multimodal fingerprints stored in the multimedia resource library is not 0, the value of the initial multimodal fingerprint of the multimedia resource is set to the number of stored fingerprints.

[0090] It should be noted that the multimodal fingerprint can be an 8-bit integer. Accordingly, when the value of the initial multimodal fingerprint is 0, it is expressed as 00000000. When the value of the initial multimodal fingerprint is a stored number, for example, the stored number is 1234, it is expressed as 00001234.

[0091] In the embodiment of the present application, a multimedia resource newly uploaded by the terminal is used as an example for explanation. In some embodiments, the server can also obtain multimedia resources for which multiple modal fingerprints have not been determined to perform the initial multimodal fingerprint setting in step 502 and the subsequent fingerprint determination process.

[0092] 503. The server inputs the text, picture and video in the multimedia resource into a text feature extraction network, a picture feature extraction network and a video feature extraction network respectively to obtain text features, picture features and video features of the multimedia resource.

[0093] The training process of the feature extraction network is shown in steps 301-303, which will not be described in detail here.

[0094] It should be noted that, in the case that the multimedia resource includes data of the three modalities described above, the extracted features for one multimedia resource include text features, image features and video features, and in some embodiments, in the case that the multimedia resource includes data of any two modalities described above, the extracted features for one multimedia resource can include the two features described above. For example, in the case that the multimedia resource includes text and image, the extracted features include text features and image features.

[0095] In the embodiments of the present application, in the feature extraction process, the server calls the corresponding feature extraction model to perform feature extraction according to the data of each modality included in the multimedia resource. In the case that the modalities included in the multimedia resource are different, the feature extraction model called by the server will be different, which will not be described here.

[0096] 504、The server performs retrieval in the multimedia resource library based on the text features of the multimedia resource, and obtains a first resource set of the multimedia resource, the first resource set including first multimedia resources matching the text features of the multimedia resource.

[0097] In the embodiments of the present application, in the case that the multimedia resource library stores a plurality of historical multimedia resources and a plurality of multi-modal fingerprints corresponding to the historical multimedia resources, the server acquires any historical multimedia resource from the multimedia resource library, extracts text features of the historical multimedia resource, and compares the text features of the multimedia resource and the historical multimedia resource based on the text features. If the text similarity between the two text features meets the similarity condition, the historical multimedia resource is determined as the first multimedia resource matching the multimedia resource.

[0098] The text similarity between the text features can be represented by the Euclidean distance between the text features, and the similarity condition met by the text similarity can be that the Euclidean distance is ranked within the target number of positions, that is, in the multimedia resource library, the text similarity between the text features of the historical multimedia resource and the text features of the multimedia resource is ranked within a certain number of positions. The target number is used to limit the number of resources in the first resource set. For example, the target number can be 10, that is, the historical multimedia resources with the top 10 similarities are put into the first resource set.

[0099] Of course, the server can acquire a plurality of historical multimedia resources at the same time, and perform parallel comparison based on the text features of the acquired historical multimedia resources and the text features of the multimedia resource to improve processing efficiency.

[0100] 505、The server performs retrieval in the multimedia resource library based on the picture feature of the multimedia resource, and obtains a second resource set of the multimedia resource, the second multimedia resource set including a second multimedia resource whose picture feature matches the multimedia resource.

[0101] In the embodiment of the present application, in the case that the multimedia resource library stores a plurality of historical multimedia resources and a plurality of multi-modal fingerprints corresponding to the historical multimedia resources, the server acquires any historical multimedia resource from the multimedia resource library, extracts a picture feature of the historical multimedia resource, and compares the picture feature of the multimedia resource with the picture feature of the historical multimedia resource. If the picture similarity between the two picture features meets a similarity condition, the historical multimedia resource is determined as a second multimedia resource matching the multimedia resource.

[0102] The picture similarity between the picture features can be represented by the Euclidean distance between the picture features, and the similarity condition met by the picture similarity can be that the Euclidean distance is ranked within a target number of positions. That is, in the multimedia resource library, the picture similarity between the picture feature of the historical multimedia resource and the picture feature of the multimedia resource is ranked within a certain number of positions. The target number is used to limit the number of resources in the second resource set. For example, the target number can be 10, that is, the historical multimedia resource whose similarity is ranked within the top 10 is put into the second resource set.

[0103] Similarly, the server can simultaneously acquire a plurality of historical multimedia resources, and perform parallel comparison based on the picture features of the acquired historical multimedia resources and the picture feature of the multimedia resource, to improve processing efficiency.

[0104] 506、The server performs retrieval in the multimedia resource library based on the picture feature of the multimedia resource, and obtains a second resource set of the multimedia resource, the second multimedia resource set including a second multimedia resource whose picture feature matches the multimedia resource.

[0105] In the embodiment of the present application, in the case that the multimedia resource library stores a plurality of historical multimedia resources and a plurality of multi-modal fingerprints corresponding to the historical multimedia resources, the server acquires any historical multimedia resource from the multimedia resource library, extracts a picture feature of the historical multimedia resource, and compares the picture feature of the multimedia resource with the picture feature of the historical multimedia resource. If the picture similarity between the two picture features meets a similarity condition, the historical multimedia resource is determined as a second multimedia resource matching the multimedia resource.

[0106] The video similarity between video features can be represented by the Euclidean distance between the video features. The similarity condition satisfied by the video similarity can be that the Euclidean distance size ranking is within the top target digits, that is, in the multimedia resource library, the video similarity between the video features of the historical multimedia resource and the video features of the multimedia resource is ranked within the top certain digits. The target digit is used to limit the number of resources in the third resource set. For example, the target digit can be 10, that is, the historical multimedia resources ranked in the top 10 in similarity are placed in the third resource set.

[0107] Likewise, the server may simultaneously obtain multiple historical multimedia resources, and perform parallel comparison based on the video features of the obtained historical multimedia resources and the video features of the multimedia resource, so as to improve processing efficiency.

[0108] The above processes 504 to 506 are all explained by taking the case where the multimedia resource library does not store the various modal features corresponding to the historical multimedia resources as an example. In the case where the multimedia resource library stores the various modal features corresponding to the historical multimedia resources, there is no need to obtain the multimedia resources in the multimedia resource library and then extract their various modal features. The server can directly obtain the historical multimedia resources and their corresponding various modal features for subsequent retrieval and comparison. The specific retrieval process in this case will not be described here.

[0109] The above steps 504 to 506 are the process of searching the three types of features separately. In this process, the retrieval process can be executed separately based on the current order, and the three types of features can also be retrieved simultaneously. Of course, the retrieval can also be performed in any order, and the embodiment of the present application does not limit this.

[0110] 507. The server determines a multimodal fusion distance between the multimedia resource and each multimedia resource in each resource set based on the text features, image features, and video features of the multimedia resource. Each multimodal fusion distance is used to represent the comprehensive matching degree between the multimedia resource and the corresponding multimedia resource in the set in each modality.

[0111] In an embodiment of the present application, taking the first resource set as an example, the server obtains the text features, image features and video features of a first multimedia resource, and respectively calculates the text similarity dt between the first multimedia resource and the multimedia resource, the image similarity di between the first multimedia resource and the multimedia resource, and the video similarity dv between the first multimedia resource and the multimedia resource, and then performs weighted summation on the respectively calculated text similarity dt, image similarity di and video similarity dv to obtain the multimodal fusion distance of the first multimedia resource.

[0112] Among them, the first multimedia resource text features, image features and video features are extracted by the corresponding feature extraction network, and each multimedia resource in each resource set has its corresponding modal features.

[0113] Taking any of the above similarities expressed using Euclidean distance as an example, the calculation formula for the above multimodal fusion distance is shown in formula (1).

[0114] D L1 =ω1d(X t , Y t )+ω2d(X i , Y i )+ω3d(X v , Y v ) (1)

[0115] In formula (1), D L1 is the multimodal fusion distance of the first multimedia resource; X t is the text feature of the multimedia resource, X i is the image feature of the multimedia resource, X v is the video feature of the multimedia resource; Y t is the text feature of the first multimedia resource, Y i is the image feature of the first multimedia resource, Y v is the video feature of the first multimedia resource; ω1, ω2 and ω3 are the distance weights corresponding to the three modalities of text, picture and video respectively; d(x, y) is the distance function.

[0116] In the embodiment of the present application, the case where the distance function d is the Euclidean distance is used as an example for explanation. When determining the similarity, the cosine distance function or the Manhattan distance function can also be used for calculation, and the embodiment of the present application does not limit this.

[0117] In some embodiments, the above distance weights are all 1 / 3. Of course, the distance weights can also be other values. That is, the distance weights corresponding to each modality can be reasonably adjusted according to the specific application scenario. For example, for an illustrated ancient poem with both picture features and text features, the text features have a greater overall impact on it. At this time, the distance weight corresponding to the text features can be adjusted to a larger value.

[0118] 508. The server determines a minimum multimodal fusion distance among the multiple multimodal fusion distances.

[0119] In an embodiment of the present application, the server compares multiple multimodal fusion distances to determine a multimodal fusion distance with a minimum value, and the multimodal fusion distance with the minimum value is the minimum multimodal fusion distance.

[0120] 509. When the minimum multimodal fusion distance is less than or equal to the distance threshold, the server determines the multimodal fingerprint of the historical multimedia resource corresponding to the minimum multimodal fusion distance as the multimodal fingerprint of the multimedia resource.

[0121] Among them, the distance threshold can be set according to the accuracy requirement of the multimodal fingerprint. For example, when the accuracy requirement is high, the distance threshold can be set to a smaller value. When the accuracy requirement is low, the distance threshold can be set to a larger value. The embodiments of the present application are not limited to this.

[0122] For example, the distance threshold may be set to 0.1. Based on the above example, if the minimum multimodal fusion distance is less than or equal to 0.1, and the corresponding multimodal fingerprint of the historical multimedia resource is 00001234, then the fingerprint of the multimedia resource is determined to be 00001234.

[0123] For example, if the multimedia resource library does not store any historical multimedia resources, then the multimedia resource is the first multimedia resource. Therefore, steps 303 to 309 do not need to be performed. Instead, the initial multimodal fingerprint of the multimedia resource can be directly determined as the multimodal fingerprint of the multimedia resource, and the multimedia resource and the corresponding multimodal fingerprint can be stored in the multimedia resource library. Based on the above example, the multimodal fingerprint of the multimedia resource is 00000000.

[0124] 510. When the minimum multimodal fusion distance is greater than the distance threshold, the server determines the initial multimodal fingerprint of the multimedia resource as the multimodal fingerprint of the multimedia resource.

[0125] The above process is described using the example of the initial multimodal fingerprint being the number of fingerprints stored in the multimedia resource library. In some embodiments, the initial multimodal fingerprint can also be determined based on the maximum fingerprint among the fingerprints stored in the multimedia resource library, by adding 1 to the maximum multimodal fingerprint in the multimedia resource library to obtain the initial multimodal fingerprint. For example, when a server processes a new multimedia resource, if the maximum multimodal fingerprint in the multimedia resource library is 99, the initial multimodal fingerprint of the new multimedia resource is 100.

[0126] Accordingly, when it is determined that the minimum multimodal fusion distance is greater than the distance threshold, the server determines the initial multimodal fingerprint of the multimedia resource as the multimodal fingerprint of the multimedia resource.

[0127] The initial multi-modal fingerprint can be the number of stored fingerprints in the multimedia resource library. For example, if 11 multi-modal fingerprints have been stored in the multimedia resource library, the multi-modal fingerprint of the multimedia resource is 00000011. In some embodiments, the initial multi-modal fingerprint can also be the maximum fingerprint in the stored fingerprints in the multimedia resource library plus 1. For example, if the maximum multi-modal fingerprint in the stored fingerprints in the multimedia resource library is 1234, the multi-modal fingerprint of the multimedia resource is 00001235.

[0128] In some embodiments, the multi-modal fingerprint can also be set to more bits according to the scene requirements. For example, the multi-modal fingerprint can be 16 bits, and the multi-modal fingerprint can also include other flag bits and date information, which is not limited in the embodiments of the present application.

[0129] 511. The server stores the multimedia resource, the corresponding multi-modal fingerprint, and the corresponding text feature, picture feature, and video feature into the multimedia resource library.

[0130] In the above technical solution, the features of different modal data of the multimedia resource are screened respectively, and then the comprehensive matching degree of the multimedia resource and the corresponding historical multimedia resource in each mode is determined based on the screened historical multimedia resources, so that the multi-modal fingerprint of the multimedia resource can be determined based on the comprehensive matching degree, so that the multi-modal fingerprint contains information for describing the multimedia resource from multiple modes, greatly improving the representativeness of the multi-modal fingerprint for the multimedia resource, and improving the accuracy of processing when the multi-modal fingerprint is applied to process the multimedia resource.

[0131] Further, since the multi-modal fingerprint is an integer data, the number of integer data bytes is small, so that when it is stored, calculated, or called, the computer storage space can be saved, the process time can be reduced, and the processing efficiency can be effectively improved.

[0132] By the technical solutions provided in the embodiments of the present application, in the generation process, the respective resource sets are determined based on the resource features of different modalities, each resource set contains resources matching the multimedia resource under the corresponding modality, and then the multimedia resource most similar in content is determined based on the comprehensive matching degree between the historical multimedia resources in the resource sets and the multimedia resource in multiple modalities, so as to determine the multi-modal fingerprint. The multi-modal fingerprint determined by the above method fuses resource content and is easy to store and calculate, which not only can play a role in identifying multimedia resources, but also can learn some correlation information between multimedia resources through the multi-modal fingerprint, thereby improving the accuracy of recommendation. Since the features of each modality are applied when generating the multi-modal fingerprint, taking the multimedia resource as an advertisement as an example, the performance of the advertisement in terms of conversion rate and click rate can be improved.

[0133] Figure 6 is a structural schematic diagram of a fingerprint generation device provided by an embodiment of the present application, referring to Figure 6 The device comprises:

[0134] The feature acquisition module 601 is configured to acquire resource features of at least two modalities based on data of the at least two modalities in the multimedia resource, the multimedia resource comprising data of at least two modalities in text, image and video.

[0135] The set acquisition module 602 is configured to acquire resource sets of the at least two modalities based on the resource features of the at least two modalities, each resource set of the modality comprising historical multimedia resources matching the resource features of the corresponding modality.

[0136] The distance determination module 603 is configured to determine a minimum multi-modal fusion distance based on the resource features of the at least two modalities of the multimedia resource, the minimum multi-modal fusion distance corresponding to a historical multimedia resource in the resource sets of the at least two modalities that has the maximum comprehensive matching degree with the multimedia resource in each modality.

[0137] The first fingerprint determination module 604 is configured to, in the case that the minimum multi-modal fusion distance is greater than the distance threshold, determine an initial multi-modal fingerprint of the multimedia resource as the multi-modal fingerprint of the multimedia resource, the initial multi-modal fingerprint being determined based on the stored multi-modal fingerprints in the multimedia resource library.

[0138] In a possible implementation manner, the device further comprises:

[0139] The second fingerprint determination module is configured to, in the case that the minimum multi-modal fusion distance is less than or equal to the distance threshold, determine the multi-modal fingerprint of the historical multimedia resource corresponding to the minimum multi-modal fusion distance as the multi-modal fingerprint of the multimedia resource.

[0140] In one possible implementation, the device further includes:

[0141] The initial fingerprint determination module is used to determine the initial multimodal fingerprint based on the number of multimodal fingerprints stored in the multimedia resource library, or to determine the initial multimodal fingerprint based on the maximum multimodal fingerprint stored in the multimedia resource library.

[0142] In one possible implementation, the feature acquisition module 601 is used to:

[0143] Based on the data of the at least two modalities, a feature extraction network of the corresponding modality is called to perform feature extraction to obtain resource features of the at least two modalities.

[0144] In one possible implementation, the set acquisition module 602 is configured to:

[0145] For the resource features of any one of the at least two modalities, a comparison is performed based on the resource features and the resource features of any historical multimedia resource. If the similarity between the two resource features meets the similarity condition, it is determined that the historical multimedia resource matches the multimedia resource, and the historical multimedia resource is placed in the resource set of the modality.

[0146] In one possible implementation, the distance determination module 603 includes:

[0147] a first determining unit, configured to determine a multimodal fusion distance between the multimedia resource and each historical multimedia resource in each resource set, wherein each multimodal fusion distance is used to represent a comprehensive matching degree between the multimedia resource and the corresponding historical multimedia resource in the set in each modality;

[0148] The second determining unit is configured to determine a minimum multimodal fusion distance among the plurality of multimodal fusion distances.

[0149] In one possible implementation, the first determining unit includes:

[0150] a determination subunit, configured to determine text similarity, image similarity, and video similarity between each of the historical multimedia resources and the multimedia resource;

[0151] The summing subunit is used to perform weighted summation on the text similarity, the image similarity, and the video similarity of each of the historical multimedia resources to obtain a multimodal fusion distance of each of the historical multimedia resources.

[0152] It should be noted that the device for generating a fingerprint provided in the above embodiments only divides the above functions into different functional modules for example when generating a fingerprint, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the device for generating a fingerprint and the method for generating a fingerprint provided in the above embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be repeated here.

[0153] Through the technical solutions provided in the embodiments of the present application, in the generation process, the corresponding resource set is determined based on the resource features of each different modality, each resource set contains resources matching the multimedia resource under the corresponding modality, and the multimedia resource most similar in content is determined based on the comprehensive matching degree of the historical multimedia resources in the resource set and the multimedia resource in multiple modalities, so as to determine the multi-modal fingerprint. The multi-modal fingerprint determined by the above method fuses resource content and is easy to store and calculate, which not only plays a role in identifying multimedia resources, but also can learn some correlation information between multimedia resources through the multi-modal fingerprint, thereby improving the accuracy of recommendation.

[0154] The embodiments of the present application provide a server for executing the above method, and the server herein is the server 120 described above. The structure of the server will be introduced as follows:

[0155] Figure 7 FIG. 7 is a structural schematic diagram of a server provided in the embodiments of the present application. The server 700 can have great differences due to different configurations or performances, and can include one or more processors (Central Processing Units, CPUs) 701 and one or more memories 702. At least one computer program is stored in the one or more memories 702, and the at least one computer program is loaded and executed by the one or more processors 701 to implement the method provided in each method embodiment. Of course, the server 700 can also have a wired or wireless network interface, a keyboard, an input and output interface, and other components for realizing the functions of the device, so as to perform input and output. The server 700 can also include other components for realizing the functions of the device, which will not be described here.

[0156] In the example embodiment, a computer readable storage medium, such as a memory including a computer program executable by a processor to perform the fingerprint generation method in the above embodiment, is also provided. For example, the computer readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0157] In the example embodiment, a computer program product or computer program including program code stored in a computer readable storage medium is also provided, the program code being read by a processor of a computer device from the computer readable storage medium, the processor executing the program code to cause the computer device to perform the fingerprint generation method.

[0158] Those of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by a program instructing relevant hardware, and the program can be stored in a computer readable storage medium, such as a Read-Only Memory (ROM), a magnetic disk or an optical disk, etc.

[0159] The above is only an optional embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A fingerprint generation method, characterized in that: The method comprises: Obtaining resource features of the at least two modalities based on data of at least two modalities in a multimedia resource, wherein the multimedia resource includes data of at least two modalities of text, image, and video; Based on the resource features of the at least two modalities, obtaining resource sets of the at least two modalities, wherein the resource set of each modality includes historical multimedia resources matching the resource features of the corresponding modality; Determining a minimum multimodal fusion distance based on resource features of the at least two modalities of the multimedia resource, the minimum multimodal fusion distance corresponding to a historical multimedia resource in the resource set of the at least two modalities that has the greatest comprehensive matching degree with the multimedia resource in each modality; When the minimum multimodal fusion distance is greater than the distance threshold, the initial multimodal fingerprint of the multimedia resource is determined as the multimodal fingerprint of the multimedia resource, and the initial multimodal fingerprint is determined based on the multimodal fingerprints stored in the multimedia resource library.

2. The method according to claim 1, characterized in that The method further comprises: When the minimum multimodal fusion distance is less than or equal to a distance threshold, the multimodal fingerprint of the historical multimedia resource corresponding to the minimum multimodal fusion distance is determined as the multimodal fingerprint of the multimedia resource.

3. The method according to claim 1, characterized in that The method further comprises: Determining the initial multimodal fingerprint based on the number of multimodal fingerprints stored in the multimedia resource library; or, The initial multimodal fingerprint is determined based on the maximum multimodal fingerprint stored in the multimedia resource library.

4. The method according to claim 1, wherein The acquiring of resource features of at least two modalities based on data of at least two modalities in the multimedia resource includes: Based on the data of the at least two modalities, a feature extraction network of the corresponding modality is called to perform feature extraction to obtain resource features of the at least two modalities.

5. The method according to claim 1, wherein The acquiring of the resource set of the at least two modalities based on the resource features of the at least two modalities includes: For the resource features of any one of the at least two modalities, a comparison is performed based on the resource features and the resource features of any historical multimedia resource. If the similarity between the two resource features meets the similarity condition, it is determined that the historical multimedia resource matches the multimedia resource, and the historical multimedia resource is placed in the resource set of the modality.

6. The method according to claim 1, characterized in that The determining of the minimum multimodal fusion distance based on the resource features of the at least two modalities of the multimedia resource includes: Determining a multimodal fusion distance between the multimedia resource and each historical multimedia resource in each resource set, each multimodal fusion distance being used to represent a comprehensive degree of matching between the multimedia resource and the corresponding historical multimedia resource in the set in each modality; A minimum multimodal fusion distance among the plurality of multimodal fusion distances is determined.

7. The method according to claim 6, characterized in that Determining the multimodal fusion distance between the multimedia resource and each historical multimedia resource in each resource set includes: Determining text similarity, image similarity, and video similarity between each of the historical multimedia resources and the multimedia resource; The text similarity, image similarity, and video similarity of each of the historical multimedia resources are weightedly summed to obtain a multimodal fusion distance of each of the historical multimedia resources.

8. A fingerprint generating device, characterized in that: The device comprises: A feature acquisition module, configured to acquire resource features of at least two modalities based on data of at least two modalities in a multimedia resource, wherein the multimedia resource includes data of at least two modalities of text, image, and video; A collection acquisition module, configured to acquire resource collections of the at least two modalities based on the resource characteristics of the at least two modalities, wherein the resource collection of each modality includes historical multimedia resources that match the resource characteristics of the corresponding modality; a distance determination module, configured to determine a minimum multimodal fusion distance based on resource features of the at least two modalities of the multimedia resource, the minimum multimodal fusion distance corresponding to a historical multimedia resource in the resource set of the at least two modalities that has the greatest degree of comprehensive matching with the multimedia resource in each modality; The first fingerprint determination module is used to determine the initial multimodal fingerprint of the multimedia resource as the multimodal fingerprint of the multimedia resource when the minimum multimodal fusion distance is greater than the distance threshold, and the initial multimodal fingerprint is determined based on the multimodal fingerprint stored in the multimedia resource library.

9. The device according to claim 8, characterized in that The device further comprises: The second fingerprint determination module is configured to determine, when the minimum multimodal fusion distance is less than or equal to a distance threshold, the multimodal fingerprint of the historical multimedia resource corresponding to the minimum multimodal fusion distance as the multimodal fingerprint of the multimedia resource.

10. The device according to claim 8, characterized in that The device further comprises: The initial fingerprint determination module is used to determine the initial multimodal fingerprint based on the number of multimodal fingerprints stored in the multimedia resource library, or to determine the initial multimodal fingerprint based on the maximum multimodal fingerprint stored in the multimedia resource library.

11. The device according to claim 8, characterized in that The collection acquisition module is used to: For the resource features of any one of the at least two modalities, a comparison is performed based on the resource features and the resource features of any historical multimedia resource. If the similarity between the two resource features meets the similarity condition, it is determined that the historical multimedia resource matches the multimedia resource, and the historical multimedia resource is placed in the resource set of the modality.

12. The device according to claim 8, characterized in that The distance determination module includes: a first determining unit, configured to determine a multimodal fusion distance between the multimedia resource and each historical multimedia resource in each resource set, wherein each multimodal fusion distance is used to represent a comprehensive matching degree between the multimedia resource and the corresponding historical multimedia resource in the set in each modality; The second determining unit is configured to determine a minimum multimodal fusion distance among the plurality of multimodal fusion distances.

13. The device according to claim 12, characterized in that The first determining unit includes: a determination subunit, configured to determine text similarity, image similarity, and video similarity between each of the historical multimedia resources and the multimedia resource; The summing subunit is used to perform weighted summation on the text similarity, image similarity and video similarity of each of the historical multimedia resources to obtain a multimodal fusion distance of each of the historical multimedia resources.

14. A server, characterized in that: The server includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the fingerprint generation method according to any one of claims 1 to 7.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the fingerprint generation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Media information processing method, server and storage medium

    CN110149529A

  • Fingerprint Anti-counterfeiting method, and electronic device

    WO2021008551A1